This paper begins an empirical study of multi-agent AI control, where several agents coordinate toward a malicious objective. The authors introduce FakeLab, a synthetic AI-lab codebase with 9 services, 86 benign tasks, and 4 attack objectives. Their reported central result is a “fragmentation effect”: as more agents coordinate, per-agent monitoring becomes less likely to catch any attacker. An explicit planner increases attack completion rates by up to sevenfold. Stronger trusted monitors reduce undetected success by more than an order of magnitude, but do not eliminate it. The work is based on a controlled synthetic environment.
No heat snapshots are available in the last 24 hours.