Oxygen-TryOn is presented as a fashion-native foundation model for virtual try-on, treating the task as understanding-driven generation from multiple references rather than mask-based inpainting. It accepts clean product images or in-the-wild worn-item photos, supports variable numbers of references and free multi-item composition, and handles full- or half-body subjects. The training recipe combines a dedicated data engine with continued pre-training, supervised fine-tuning, and reinforcement learning using hybrid rewards. The authors report state-of-the-art consistency and realism on public benchmarks and their Oxygen-TryOn Bench, but the abstract provides no numerical results or detailed evaluation protocol.
No heat snapshots are available in the last 24 hours.