Pıer
TidesCurrentsHarbor LightsLabBottlesAshore
Pıer

Navigation

  • Tides
  • Ashore
  • Harbor Lights
  • Agent Access
  • Changelog
  • Bottles
  • Now
  • Feedback

External links

GitHubCloudborne ↗

© 2026 Pier.

WatchingResearchWatching0 independent reports0

Tencent WorkBuddy Bench: A Multi-Domain Coding-Agent Benchmark with Contamination-Resistant Task Construction

First seen · 7/24/2026, 12:00 PMLatest activity · 7/24/2026, 12:00 PM

Tencent introduces WorkBuddy Bench, an open, multi-domain benchmark for coding agents across Code, Web, Office, and Security workflows. Each task is reverse-engineered from a real commit, pull request, or business scenario, then rewritten as a colloquial role-play request so the original prompt is difficult to recover through web search. The release includes task directories, environment images, evaluation harnesses, tests, and reference solutions. It uses a uniform task format and reproducible protocol across CodeBuddy Code and Claude Code, while reporting separate subset scores rather than a misleading suite-wide average.

Event heat · last 24 hours

No heat snapshots are available in the last 24 hours.

No heat snapshots are available in the last 24 hours.

Reporting Timeline

  1. AggregatorHuggingFace Daily Papers7/24, 12:00 PMnot independentRepresentative
    Tencent WorkBuddy Bench: A Multi-Domain Coding-Agent Benchmark with Contamination-Resistant Task Construction