Mage-Flow is a compact 4B-scale foundation model family for text-to-image generation and instruction-based image editing. It co-designs Mage-VAE, a lightweight high-fidelity latent tokenizer, with a native-resolution multimodal Diffusion Transformer trained using rectified flow matching. The paper reports more than an order-of-magnitude lower tokenization cost and about 2.5x higher end-to-end training throughput through native-resolution packing and CUDA kernel fusion. Base, RL-aligned, and 4-step Turbo variants are provided. On one NVIDIA A100 at 1024², the authors report 0.59 seconds for Turbo generation and 1.02 seconds for Turbo editing.
No heat snapshots are available in the last 24 hours.