Pıer
TidesCurrentsHarbor LightsLabBottlesAshore
Pıer

Navigation

  • Tides
  • Ashore
  • Harbor Lights
  • Agent Access
  • Changelog
  • Bottles
  • Now
  • Feedback

External links

GitHubCloudborne ↗

© 2026 Pier.

WatchingResearchWatching0 independent reports0

Argus-Unified: Towards a Compact and Economical Unified Model for Image Understanding and Generation

First seen · 7/28/2026, 06:12 PMLatest activity · 7/28/2026, 06:12 PM

Argus-Unified is a compact unified multimodal model for image understanding and generation. It reuses a pretrained vision-language model and introduces hybrid visual tokens: continuous tokens retain information for understanding, while learned discrete tokens support image generation. Training has two stages: learning a quantizer and decoder over a frozen vision encoder, followed by unified multimodal modeling with an LLM initialized from a pretrained VLM. The paper reports using 15.6M data samples and about $2,000 in training cost, claiming state-of-the-art results on GQA, POPE, and VQAv2, with competitive generation quality against Janus and Janus-Pro.

Event heat · last 24 hours

No heat snapshots are available in the last 24 hours.

No heat snapshots are available in the last 24 hours.

Reporting Timeline

  1. AggregatorarXiv7/28, 06:12 PMnot independentRepresentative
    Argus-Unified: Towards a Compact and Economical Unified Model for Image Understanding and Generation