AdvancedMathBench evaluates advanced mathematical proof generation and verification across undergraduate and doctoral qualifying-exam levels. ProverBench includes 296 problems, while VerifierBench contains 888 model-generated proof trajectories with expert ground truth. The authors introduce an automatic verification pipeline trained on large-scale expert annotations to assess both validity and fine-grained proof errors. GPT-5.5-xhigh scores 75.8 on the UGD split and 66.1 on the QE split for proof generation. The best verification result reaches only 65.1 Balanced F1, with generally low true-negative rates.
No heat snapshots are available in the last 24 hours.