Revisiting Avatar-As-Image: High-Fidelity Registration is All You Need
Representing clothed 3D avatars as standardized 2D UV texture and displacement maps has long promised to bridge 3D assets with 2D image models, yet imprecise scan registration has historically bottlenecked output quality. AvaImg resolves this through a multi-stage optimization pipeline that enforces body-inside-clothing constraints via signed winding numbers, reducing runtime tenfold while maintaining a 34.48 dB PSNR across six datasets. Passing these maps through a frozen FLUX VAE adds just 0.76 mm of Chamfer error, showing that 2D generative priors can directly handle high-fidelity 3D human synthesis.
Why it's worth reading
It demonstrates that high-precision registration allows frozen 2D foundation models like FLUX to process 3D clothed avatars with sub-millimeter geometric loss.