This paper introduces HIEVI-RAG, a hierarchical multimodal RAG framework for closed-domain long-document understanding. It decomposes complex questions into atomic sub-questions, performs coarse visual page retrieval, verifies candidate pages with EVIAGENT, and generates answers through memory-guided iterative reasoning. EVIAGENT is described as a multi-page verifier trained with GRPO for cross-page reasoning over multi-image blocks. The authors report evaluations on four benchmarks, claiming that HIEVI-RAG outperforms existing open-source baselines and exceeds the strongest reported baseline by an average of 8.05% in accuracy.
No heat snapshots are available in the last 24 hours.