The paper introduces an interactive simulation interface for collecting human preferences over intermediate decisions in generative-agent trajectories, including planning, memory retrieval, reflection, and action selection. It releases a dataset containing 57K fine-grained annotations and applies supervised fine-tuning and direct preference optimization to open-weight language models. According to the abstract, both methods consistently improve simulation fidelity, coordination, interaction quality, and socially effective behavior. The work frames step-level preference supervision as a training signal for improving local decisions while also influencing long-horizon agent behavior.
No heat snapshots are available in the last 24 hours.