TOUR introduces a benchmark for evaluating trajectory-level data deletion in offline reinforcement learning. It combines trajectory partitions, matched non-member controls, retraining references, retained-performance anchors, and audits using multiple attack families. Experiments on D4RL locomotion tasks and an exploratory AntMaze extension suggest that deletion conclusions are environment-dependent. Retraining and fine-tuning can provide stronger retained-utility references than uniform GA+Refit, while TrajDeleter is not consistently superior under the same audit. The authors argue that a single likelihood-based membership score can overstate unlearning quality, making matched controls, retraining-relative calibration, attack diversity, and utility preservation necessary.
No heat snapshots are available in the last 24 hours.