Argus-Unified is a compact unified multimodal model for image understanding and generation. It reuses a pretrained vision-language model and introduces hybrid visual tokens: continuous tokens retain information for understanding, while learned discrete tokens support image generation. Training has two stages: learning a quantizer and decoder over a frozen vision encoder, followed by unified multimodal modeling with an LLM initialized from a pretrained VLM. The paper reports using 15.6M data samples and about $2,000 in training cost, claiming state-of-the-art results on GQA, POPE, and VQAv2, with competitive generation quality against Janus and Janus-Pro.
No heat snapshots are available in the last 24 hours.