Pıer
TidesCurrentsHarbor LightsLabBottlesAshore
Pıer

Navigation

  • Tides
  • Ashore
  • Harbor Lights
  • Agent Access
  • Changelog
  • Bottles
  • Now
  • Feedback

External links

GitHubCloudborne ↗

© 2026 Pier.

WatchingResearchWatching0 independent reports0

JOR-Bench: Japanese Operations Research Benchmarks for Large Language Models

First seen · 7/18/2026, 07:24 PMLatest activity · 7/18/2026, 07:24 PM

JOR-Bench introduces five Japanese-language benchmarks for testing whether large language models can formulate and solve operations research problems. It translates IndustryOR, MAMO Complex LP, NL4OPT, OptiBench, and OptMATH into 1,319 problems covering linear, mixed-integer, nonlinear, and combinatorial optimization. The authors evaluate seven multilingual and Japanese-specialized models in English and Japanese, standardizing execution through the Python interface of OR-Tools. Strong multilingual models show nearly language-neutral formulation performance, with an average English-Japanese accuracy difference of only -0.3 percentage points. However, error analysis identifies pragmatic ambiguity in Japanese prompts, where models sometimes return decision-variable values instead of the requested objective value.

Event heat · last 24 hours

No heat snapshots are available in the last 24 hours.

No heat snapshots are available in the last 24 hours.

Reporting Timeline

  1. AggregatorarXiv7/18, 07:24 PMnot independentRepresentative
    JOR-Bench: Japanese Operations Research Benchmarks for Large Language Models