Hunyuan3D-Buffalo 1.0 proposes one architecture for 3D understanding, text-to-3D generation, instruction-guided editing, and text-grounded part generation. Its reported 87M-example corpus contains 25M understanding samples, 50M text-to-3D pairs, and 12M editing pairs produced with Nano3D-v2. The framework combines Hunyuan3D-VLM for semantic, structural, and spatial understanding with Hunyuan3D DiT for synthesis. Editing and part generation additionally condition diffusion on a source-object representation to preserve global structure and regions that should remain unchanged.
No heat snapshots are available in the last 24 hours.