JOR-Bench introduces five Japanese-language benchmarks for testing whether large language models can formulate and solve operations research problems. It translates IndustryOR, MAMO Complex LP, NL4OPT, OptiBench, and OptMATH into 1,319 problems covering linear, mixed-integer, nonlinear, and combinatorial optimization. The authors evaluate seven multilingual and Japanese-specialized models in English and Japanese, standardizing execution through the Python interface of OR-Tools. Strong multilingual models show nearly language-neutral formulation performance, with an average English-Japanese accuracy difference of only -0.3 percentage points. However, error analysis identifies pragmatic ambiguity in Japanese prompts, where models sometimes return decision-variable values instead of the requested objective value.
No heat snapshots are available in the last 24 hours.