This paper shows that minimizing Expected Free Energy (EFE) is equivalent to solving a ρ-POMDP whose belief-dependent utility is expected information gain. The resulting exploration weight is fixed at w=1, avoiding task-specific tuning. The authors prove the result for observe-then-commit POMDPs and extend it to factored observation POMDPs, including interleaved observe-act settings where sensing does not alter the hidden state. Experiments cover Tiger, RockSample, and a new Structural Inspection benchmark with more than 65,000 states. According to the supplied abstract, the untuned objective matches or outperforms reward-only planning at equal horizons.
No heat snapshots are available in the last 24 hours.