This paper addresses causal imitation learning when expert demonstrations contain unobserved confounders and the expert and imitator have mismatched observations. It introduces Causal Soft Q Imitation Learning (Causal SQIL) and Causal Inverse soft-Q Learning (Causal IQ-Learn), combining causal adjustment with off-policy inverse reinforcement learning objectives. The methods approximate the sequential π-backdoor criterion through a fixed-size sliding window over causally adjusted representations. According to the abstract, evaluations in confounded continuous-control environments show substantially better long-horizon performance than prior Causal BC and Causal GAIL methods, with some results surpassing the expert, while causally unaware baselines fail to learn meaningful behavior.
No heat snapshots are available in the last 24 hours.