Pıer
TidesCurrentsHarbor LightsLabBottlesAshore
Pıer

Navigation

  • Tides
  • Ashore
  • Harbor Lights
  • Agent Access
  • Changelog
  • Bottles
  • Now
  • Feedback

External links

GitHubCloudborne ↗

© 2026 Pier.

WatchingResearchWatching0 independent reports0

Video Generation Models are General-Purpose Vision Learners

First seen · 7/13/2026, 12:00 PMLatest activity · 7/13/2026, 12:00 PM

The paper introduces GenCeption, a feed-forward perception model built from a pretrained text-to-video diffusion backbone and controlled by text instructions. It targets multiple vision tasks, including depth, surface-normal, camera-pose estimation, expression-referring segmentation, and 3D keypoint prediction. According to the abstract, GenCeption matches or surpasses several specialized systems and requires 7 to 500 times less training data to reach comparable performance to D4RT and VGGT-Omega. A model trained only on synthetic human videos reportedly generalizes to real footage and out-of-distribution categories such as animals and robots.

Event heat · last 24 hours

No heat snapshots are available in the last 24 hours.

No heat snapshots are available in the last 24 hours.

Reporting Timeline

  1. AggregatorHuggingFace Daily Papers7/13, 12:00 PMnot independentRepresentative
    Video Generation Models are General-Purpose Vision Learners