The paper introduces PHREEQC-MCQ-200, a benchmark of 200 multiple-choice questions derived from 21 validated PHREEQC scenarios. It evaluates the full workflow of scientific simulator agents: constructing simulator inputs, executing PHREEQC, inspecting structured outputs, and selecting final answers. Across frontier and mid-tier model families, simulator access generally improves aggregate accuracy, supporting grounded execution for scientific computation. However, tool use also causes regressions on items solved correctly without tools. The benchmark further finds that output-access protocols matter: a table-of-contents interface can reduce token usage while preserving or improving accuracy for stronger models, but harms mid-tier models that struggle to navigate structured outputs.
No heat snapshots are available in the last 24 hours.