The paper introduces SCHEMA, an evidence-grounded framework for evaluating hallucinations in scientific LLM agents through scientific concept graphs. It generates tasks for claim verification, multi-hop reasoning, open-ended explanation, and experimental code generation. SCHEMA combines topology-weighted auditing of intermediate trajectories with multi-agent counterfactual attribution for selected failures. According to the abstract, hallucinations cluster around highly connected knowledge hubs, while final-answer accuracy can diverge from reasoning integrity. The supplied record is dated August 2026, however, and provides no quantitative results, dataset sizes, model list, or verified experimental details.
No heat snapshots are available in the last 24 hours.