Pıer
TidesCurrentsHarbor LightsLabBottlesAshore
Pıer

Navigation

  • Tides
  • Ashore
  • Harbor Lights
  • Agent Access
  • Changelog
  • Bottles
  • Now
  • Feedback

External links

GitHubCloudborne ↗

© 2026 Pier.

WatchingResearchWatching0 independent reports0

PHREEQC-MCQ-200: A Diagnostic Benchmark for Tool-Augmented Scientific Simulator Agents

First seen · 7/1/2026, 12:51 PMLatest activity · 7/1/2026, 12:51 PM

The paper introduces PHREEQC-MCQ-200, a benchmark of 200 multiple-choice questions derived from 21 validated PHREEQC scenarios. It evaluates the full workflow of scientific simulator agents: constructing simulator inputs, executing PHREEQC, inspecting structured outputs, and selecting final answers. Across frontier and mid-tier model families, simulator access generally improves aggregate accuracy, supporting grounded execution for scientific computation. However, tool use also causes regressions on items solved correctly without tools. The benchmark further finds that output-access protocols matter: a table-of-contents interface can reduce token usage while preserving or improving accuracy for stronger models, but harms mid-tier models that struggle to navigate structured outputs.

Event heat · last 24 hours

No heat snapshots are available in the last 24 hours.

No heat snapshots are available in the last 24 hours.

Reporting Timeline

  1. AggregatorarXiv7/1, 12:51 PMnot independentRepresentative
    PHREEQC-MCQ-200: A Diagnostic Benchmark for Tool-Augmented Scientific Simulator Agents