Pıer
TidesCurrentsHarbor LightsLabBottlesAshore
Pıer

Navigation

  • Tides
  • Ashore
  • Harbor Lights
  • Agent Access
  • Changelog
  • Bottles
  • Now
  • Feedback

External links

GitHubCloudborne ↗

© 2026 Pier.

WatchingResearchWatching0 independent reports0

NOLLI: A Difficulty-Calibrated Puzzle Benchmark for Diagnosing the English-Korean Performance Gap

First seen · 8/5/2026, 04:00 AMLatest activity · 8/5/2026, 04:00 AM

NOLLI is a procedurally generated, seed-regenerable benchmark for diagnosing English-Korean performance differences. It contains 15 puzzle types across 25 tasks and 7,500 items, with unique solutions and deterministic scoring. Across 15 frontier, open-weight, and Korean-developed models, matched English-Korean accuracy was statistically equivalent within a +/-10 percentage-point margin for models above a 3% accuracy floor. Larger gaps appeared in writing-system-intensive tasks: Korean Cipher lagged English by up to 68.7 points, while Cryptarithmetic using the same Hangul jamo showed no systematic penalty. The authors interpret the contrast as evidence about multi-step sub-syllabic execution, not simply language presentation.

Event heat · last 24 hours

No heat snapshots are available in the last 24 hours.

No heat snapshots are available in the last 24 hours.

Reporting Timeline

  1. AggregatorHuggingFace Daily Papers8/5, 04:00 AMnot independentRepresentative
    NOLLI: A Difficulty-Calibrated Puzzle Benchmark for Diagnosing the English-Korean Performance Gap