The paper introduces MAS-HQ (Multi-Agent System Hallucination Quest), a resource-aware protocol that evaluates factuality together with compute cost in competitive matches. Its Q-Score subtracts normalized cost from factuality. In the reported example, a Best-of-4 agent achieves higher raw factuality than a more efficient system, 0.9169 versus 0.9103, but receives a lower Q-Score, 0.5169 versus 0.5217, while using roughly four times the tokens and latency. The abstract reports small but consistent efficiency gains across summarization and open-domain QA over 100 trials.
No heat snapshots are available in the last 24 hours.