This paper introduces IBA-Bench, a benchmark for testing whether personalized LLM agents can infer implicit user constraints from noisy, longitudinal interaction histories and honor them during task execution. The setup targets a “knowledge-to-action gap” that profile-based question answering and fixed preference snapshots may overlook. The authors also propose IBA-Agent, which combines broad retrieval with trajectory-level alignment to reconcile conflicting priorities. According to the abstract, it substantially improves behavioral alignment across complex scenarios in nine application domains, while state-of-the-art agents still struggle. Exact models, dataset size, baselines, metrics, and numerical gains are not reported in the supplied abstract.
No heat snapshots are available in the last 24 hours.