Stage I —Supervised finetuning
3DGS render
Input video
VAE Encoder
noise
cond. latents
binary mask
LoRA
Video-DiTbase frozen
denoised
latents
VAE Decoder
FixAnything output
Output target video