Pıer
TidesCurrentsHarbor LightsLabBottlesAshore
Pıer

Navigation

  • Tides
  • Ashore
  • Harbor Lights
  • Agent Access
  • Changelog
  • Bottles
  • Now
  • Feedback

External links

GitHubCloudborne ↗

© 2026 Pier.

WatchingResearchWatching0 independent reports0

Codifying the Judge: Scalable Evaluation via Program Distillation

First seen · 7/28/2026, 12:00 PMLatest activity · 7/28/2026, 12:00 PM

The paper introduces PAJAMA, which distills an LLM judge’s decision logic into a committee of inspectable and editable programs. These programs score candidates directly, while low-confidence cases can fall back to an LLM. According to the abstract, programmatic judges match a 13B LLM judge across five datasets and four model families, and improve accuracy and throughput when used as routing signals. On RewardBench, a reward model trained from program verdicts outperforms one trained on proprietary-LLM labels at two orders of magnitude lower API cost.

Event heat · last 24 hours

No heat snapshots are available in the last 24 hours.

No heat snapshots are available in the last 24 hours.

Reporting Timeline

  1. AggregatorHuggingFace Daily Papers7/28, 12:00 PMnot independentRepresentative
    Codifying the Judge: Scalable Evaluation via Program Distillation