This paper introduces Counterfactual Shapley Credit Assignment, a causal framework for separating an agent’s policy contribution from environmental luck in reinforcement learning. It uses Counterfactual Shapley Values (φ-values) to redistribute rewards across trajectories, derives a consistent and efficient estimator, and proposes φ-PPO with Prioritized Trajectory Replay (PTR). According to the abstract, the method addresses sparse causality, high stochasticity, and delayed rewards while preserving the optimal policy. Experiments reportedly show closer alignment with ground-truth reward causes and better sample efficiency than prior methods in difficult environments where competing approaches fail to converge.
No heat snapshots are available in the last 24 hours.