Pıer
TidesCurrentsHarbor LightsLabBottlesAshore
Pıer

Navigation

  • Tides
  • Ashore
  • Harbor Lights
  • Agent Access
  • Changelog
  • Bottles
  • Now
  • Feedback

External links

GitHubCloudborne ↗

© 2026 Pier.

WatchingResearchWatching0 independent reports0

VideoCoCo: Code-as-CoT for Physically Consistent Video Generation

First seen · 7/31/2026, 12:00 PMLatest activity · 7/31/2026, 12:00 PM

VideoCoCo introduces an agentic dual-engine framework for physically consistent text-to-video generation. A coding agent converts a prompt into executable Blender code that specifies both the scene and its temporal evolution. Blender then produces a deterministic spatiotemporal draft, which a generative video engine transforms into a photorealistic result through draft-conditioned editing. The authors also build VideoCoCo-3K, a dataset of draft-instruction-target triplets for adapting the editor to simulated inputs. Against the OmniWeaving baseline, VideoCoCo improves PhyGenBench from 0.475 to 0.558 and VBench-2.0 from 52.18 to 77.88.

Event heat · last 24 hours

No heat snapshots are available in the last 24 hours.

No heat snapshots are available in the last 24 hours.

Reporting Timeline

  1. AggregatorHuggingFace Daily Papers7/31, 12:00 PMnot independentRepresentative
    VideoCoCo: Code-as-CoT for Physically Consistent Video Generation