Pıer
TidesCurrentsHarbor LightsLabBottlesAshore
Pıer

Navigation

  • Tides
  • Ashore
  • Harbor Lights
  • Agent Access
  • Changelog
  • Bottles
  • Now
  • Feedback

External links

GitHubCloudborne ↗

© 2026 Pier.

WatchingResearchWatching0 independent reports0

Principled Direction-Free Intrinsic Motivation through Model-Free Epistemic Free-Energy Estimators

First seen · 7/18/2026, 11:44 PMLatest activity · 7/18/2026, 11:44 PM

This paper proposes a single preference-free intrinsic reward for unsupervised reinforcement learning in environments containing both reducible and irreducible uncertainty. Its central signal is parameter information gain: it encourages exploration where dynamics remain unresolved, then vanishes as the model explains those dynamics. The method combines a pseudocount for epistemic value, a probe-based penalty for aleatoric variance, and a short-horizon gate for informative successors, without fitting an explicit next-state predictor. Freezing reward-defining objects within windows is used to obtain a stationary Bellman operator, bounded learning targets, and conditional uniform-concentration results under stated assumptions.

Event heat · last 24 hours

No heat snapshots are available in the last 24 hours.

No heat snapshots are available in the last 24 hours.

Reporting Timeline

  1. AggregatorarXiv7/18, 11:44 PMnot independentRepresentative
    Principled Direction-Free Intrinsic Motivation through Model-Free Epistemic Free-Energy Estimators