Agreed! I wonder if we could use a little fine tuning to enforce more visual consistency between the key frame generations — combined with depth conditioning it seems to take us most of the way there
Fine-tuning visual consistency in keyframe generation with depth conditioning
By
–