SearchEyes proposes a simulated search world built on a typed knowledge graph to unify training-data construction, search environments, and reward design for multimodal deep-search agents. Its Perception-Knowledge Chains (PKC) sample constrained multi-hop paths over the visual-knowledge intersection of Wikidata5M, preserving hop-level entity metadata. These anchors define both a reproducible environment and step-level reward signals. Hop-Anchored Policy Optimization (HaPO) uses them for credit assignment without a separately trained process reward model. Across six multimodal knowledge-intensive benchmarks, SearchEyes-27B reportedly improves the average score by 6.2 points over the strongest open-source baseline.
No heat snapshots are available in the last 24 hours.