Boogu-Image-0.1 is an open unified multimodal understanding and generation model family with Base, Turbo, Edit, and Edit-Turbo variants. It targets high-quality text-to-image generation, fast inference, instruction-based editing, and Chinese-English text rendering. The authors attribute its performance to a stronger multimodal encoder, agentic prompt rewriting, improved data and training pipelines, and inference-time scaling. They report using 208.62 million unique images, with an estimated theoretical training cost of about $400,000 for the base model. Weights, code, and recipes are released under Apache 2.0.
No heat snapshots are available in the last 24 hours.