MIRA-Math evaluates a narrow but important capability: solving mathematical problems when the solver-facing view omits exactly one necessary atomic fact. A model must request that fact in natural language under a strict budget, receive it through a constrained responder, and integrate it into an exact answer. The benchmark contains 2,310 generated instances across 22 mathematical families, including algebra, probability, linear systems, Markov chains, circuits, interpolation, and numerical boundary-value problems. Its design separates successful information requests from downstream mathematical accuracy, exposing failures that ordinary fully specified benchmarks conceal.
No heat snapshots are available in the last 24 hours.