The paper introduces GenCeption, a feed-forward perception model built from a pretrained text-to-video diffusion backbone and controlled by text instructions. It targets multiple vision tasks, including depth, surface-normal, camera-pose estimation, expression-referring segmentation, and 3D keypoint prediction. According to the abstract, GenCeption matches or surpasses several specialized systems and requires 7 to 500 times less training data to reach comparable performance to D4RT and VGGT-Omega. A model trained only on synthetic human videos reportedly generalizes to real footage and out-of-distribution categories such as animals and robots.
No heat snapshots are available in the last 24 hours.