NOLLI: A Difficulty-Calibrated Puzzle Benchmark for Diagnosing the English-Korean Performance Gap
NOLLI is a procedurally generated, seed-regenerable benchmark for diagnosing English-Korean performance differences. It contains 15 puzzle types across 25 tasks and 7,500 items, with unique solutions and deterministic scoring. Across 15 frontier, open-weight, and Korean-developed models, matched English-Korean accuracy was statistically equivalent within a +/-10 percentage-point margin for models above a 3% accuracy floor. Larger gaps appeared in writing-system-intensive tasks: Korean Cipher lagged English by up to 68.7 points, while Cryptarithmetic using the same Hangul jamo showed no systematic penalty. The authors interpret the contrast as evidence about multi-step sub-syllabic execution, not simply language presentation.
Why it's worth reading
It separates language presentation, Hangul writing-system processing, and rule execution, giving multilingual-model evaluators a more actionable way to interpret Korean benchmark gaps.