The paper introduces a two-level meta-rubric framework for evaluating factual completeness, complementing precision-focused methods that mainly detect false claims. A structured rubric captures the organization, importance, open-ended sets, ordered processes, and relationships required in a complete answer, then compiles these requirements into machine-gradable binary checks. GAMUT contains 1,813 questions grounded in real wearable imagery across 10 domains, with an additional text-only variant. Evaluation of 14 frontier and open-weight models reports a best score of 58.7% for Gemini 3.1 Pro, while claiming strong discrimination and robustness across judges.
No heat snapshots are available in the last 24 hours.