This paper introduces Moving Alphabet, a procedural testbed that renders moving letters with controlled fonts, colors, sizes, positions, directions, and speeds. By corrupting ground-truth metadata, the authors study how video distribution, duration, and caption quality affect text-to-video training. The reported findings are that diverse and balanced data improves generalization, caption quality strongly affects both model quality and training efficiency, and classifier-free guidance plus high-quality fine-tuning can only partially recover from corrupted captions. The study argues that pre-training data deserves more systematic scientific investigation.
No heat snapshots are available in the last 24 hours.