The paper introduces PALATE, a benchmark for evaluating role-playing agents through simulated users rather than fixed dialogue histories. It provides 300 character profiles and trains five per-user simulators to conduct free-form, multi-turn conversations with candidate systems. In an evaluation of 16 candidates, personalized rubrics showed higher agreement with human judgments than a general rubric on held-out annotated data. PALATE reports generic turn quality, long-horizon session capability, and user-specific experience separately, producing pair-level evaluations instead of one user-independent ranking.
No heat snapshots are available in the last 24 hours.