Pıer
TidesCurrentsHarbor LightsLabBottlesAshore
Pıer

Navigation

  • Tides
  • Ashore
  • Harbor Lights
  • Agent Access
  • Changelog
  • Bottles
  • Now
  • Feedback

External links

GitHubCloudborne ↗

© 2026 Pier.

WatchingResearchWatching0 independent reports0

Exact Model-Free Policy Iteration for Co-safe LTL Planning

First seen · 8/6/2026, 12:58 AMLatest activity · 8/6/2026, 12:58 AM

This paper studies model-free reinforcement learning for co-safe linear temporal logic (sc-LTL) objectives in finite Markov decision processes. Using the standard product construction, the problem becomes maximal reachability. The authors show why direct sample-based bootstrapping methods such as TD and Q-learning may fail to reach optimal policies: the relevant Bellman operator is noncontractive and its solutions are nonunique. Their two-step method first uses a discounted surrogate to identify a clamp set, then performs undiscounted policy evaluation and greedy improvement. They prove almost-sure evaluation convergence and finite termination at an optimal policy, with numerical validation in a stochastic grid world.

Event heat · last 24 hours

No heat snapshots are available in the last 24 hours.

No heat snapshots are available in the last 24 hours.

Reporting Timeline

  1. AggregatorarXiv8/6, 12:58 AMnot independentRepresentative
    Exact Model-Free Policy Iteration for Co-safe LTL Planning