- Extended ControlNet to text-to-video diffusion models using AnimateDiff and Motion LoRA.
- Modified model architecture to prevent training collapse where generated videos ignored depth-map conditioning; stabilized training by adding auxiliary supervision losses.
- Conducted ablation experiments on prompts, depth conditioning, and text guidance.