DataEvolver replaces the conventional crawl-filter-freeze pipeline for text-rich image data with an iterative, feedback-driven multi-agent process. A Retriever gathers candidates, a Verifier scores quality and records rejection causes, a Critic turns round-level failures into semantic feedback, and a Generator synthesizes samples for under-covered regions. With a 0.75M-data budget on PixArt-alpha, the paper reports OCR-F1 improvements of 85.3% on TextScenesHQ and 35.3% on LongTextBench over the strongest baseline. The gains also transfer to Show-o2, suggesting the data construction strategy is not specific to one downstream generator.
No heat snapshots are available in the last 24 hours.