Pıer
TidesCurrentsHarbor LightsLabBottlesAshore
Pıer

Navigation

  • Tides
  • Ashore
  • Harbor Lights
  • Agent Access
  • Changelog
  • Bottles
  • Now
  • Feedback

External links

GitHubCloudborne ↗

© 2026 Pier.

WatchingResearchWatching0 independent reports0

UniClawBench: A Universal Benchmark for Proactive Agents on Real-World Tasks

First seen · 7/10/2026, 12:00 PMLatest activity · 7/10/2026, 12:00 PM

UniClawBench is a capability-driven benchmark for proactive agents operating on dynamic, real-world tasks. It organizes evaluation around five capabilities: Skill Usage, Exploration, Long-Context Reasoning, Multimodal Understanding, and Cross-Platform Coordination. The benchmark contains 400 bilingual tasks executed in live Docker containers and scored with fine-grained, step-level checkpoints instead of static answers. Its closed-loop setup uses an executor agent, a hidden supervisor agent, and a user agent to simulate multi-turn feedback while hiding grading criteria. The authors also evaluate models across multiple agent frameworks to separate base-model capabilities from framework effects.

Event heat · last 24 hours

No heat snapshots are available in the last 24 hours.

No heat snapshots are available in the last 24 hours.

Reporting Timeline

  1. AggregatorHuggingFace Daily Papers7/10, 12:00 PMnot independentRepresentative
    UniClawBench: A Universal Benchmark for Proactive Agents on Real-World Tasks