Pıer
TidesCurrentsHarbor LightsLabBottlesAshore
Pıer

Navigation

  • Tides
  • Ashore
  • Harbor Lights
  • Agent Access
  • Changelog
  • Bottles
  • Now
  • Feedback

External links

GitHubCloudborne ↗

© 2026 Pier.

WatchingResearchWatching0 independent reports0

OmegaUse-OfficeVal: Benchmarking LLM Agents on Long-Horizon Office-Suite Tasks with Economic Grounding

First seen · 7/30/2026, 12:00 PMLatest activity · 7/30/2026, 12:00 PM

OmegaUse-OfficeVal is a benchmark for evaluating LLM agents on long-horizon office-suite workflows with task-level economic grounding. It contains 100 practitioner-derived tasks adapted through a privacy-preserving process, requiring 2.32 hours of human labor on average. Each task includes human labor time and a task-price proxy, enabling comparisons between human and inference costs. Code-based verifiers are built from fine-grained rubrics. The abstract reports that evaluated frontier LLMs are substantially faster and cheaper than humans, but still below human-level deliverable quality. The dataset and code are open-sourced.

Event heat · last 24 hours

No heat snapshots are available in the last 24 hours.

No heat snapshots are available in the last 24 hours.

Reporting Timeline

  1. AggregatorHuggingFace Daily Papers7/30, 12:00 PMnot independentRepresentative
    OmegaUse-OfficeVal: Benchmarking LLM Agents on Long-Horizon Office-Suite Tasks with Economic Grounding