Tencent introduces WorkBuddy Bench, an open, multi-domain benchmark for coding agents across Code, Web, Office, and Security workflows. Each task is reverse-engineered from a real commit, pull request, or business scenario, then rewritten as a colloquial role-play request so the original prompt is difficult to recover through web search. The release includes task directories, environment images, evaluation harnesses, tests, and reference solutions. It uses a uniform task format and reproducible protocol across CodeBuddy Code and Claude Code, while reporting separate subset scores rather than a misleading suite-wide average.
No heat snapshots are available in the last 24 hours.