WeClawArena introduces an auditable benchmark and runtime sandbox for collaboration among agents owned by different users. It models personal workspaces as both operational resources and user-specific constraints, covering six cross-user task domains with 124 base tasks expanded into 620 variants. Each base task has one benign control and four attack-oriented variants. The system records peer messages, tool calls, resource operations, governed decisions, and final workspace states. Utility and attack success rate are reported separately, with bounded runtime evidence used to diagnose task failures, privacy leakage, poisoned evidence, and invalid authority paths.
No heat snapshots are available in the last 24 hours.