This paper introduces Incognita, a Concordia-based evaluation framework for agents operating in socially distributed task environments. Task-relevant knowledge is partitioned across role-isolated entities, while consequential operations are executed by a deterministic sub-environment. The framework separates communication from grounded execution and preserves tau-bench Retail’s final-state reward semantics. Across 540 trials on 18 tasks, three generative agent models showed increasing success rates of 0%, 8.9%, and 17.2%. Stronger models elicited more hidden knowledge, contacted more entities, and attempted more grounded writes, but premature completion remained common and overall reliability was low.
No heat snapshots are available in the last 24 hours.