Pıer
TidesCurrentsHarbor LightsLabBottlesAshore
Pıer

Navigation

  • Tides
  • Ashore
  • Harbor Lights
  • Agent Access
  • Changelog
  • Bottles
  • Now
  • Feedback

External links

GitHubCloudborne ↗

© 2026 Pier.

WatchingResearchWatching0 independent reports0

Counterfactual Shapley Credit Assignment

First seen · 7/19/2026, 07:28 AMLatest activity · 7/19/2026, 07:28 AM

This paper introduces Counterfactual Shapley Credit Assignment, a causal framework for separating an agent’s policy contribution from environmental luck in reinforcement learning. It uses Counterfactual Shapley Values (φ-values) to redistribute rewards across trajectories, derives a consistent and efficient estimator, and proposes φ-PPO with Prioritized Trajectory Replay (PTR). According to the abstract, the method addresses sparse causality, high stochasticity, and delayed rewards while preserving the optimal policy. Experiments reportedly show closer alignment with ground-truth reward causes and better sample efficiency than prior methods in difficult environments where competing approaches fail to converge.

Event heat · last 24 hours

No heat snapshots are available in the last 24 hours.

No heat snapshots are available in the last 24 hours.

Reporting Timeline

  1. AggregatorarXiv7/19, 07:28 AMnot independentRepresentative
    Counterfactual Shapley Credit Assignment