Pıer
TidesCurrentsHarbor LightsLabBottlesAshore
Pıer

Navigation

  • Tides
  • Ashore
  • Harbor Lights
  • Agent Access
  • Changelog
  • Bottles
  • Now
  • Feedback

External links

GitHubCloudborne ↗

© 2026 Pier.

WatchingResearchWatching0 independent reports0

LeAct: Learning to Reason from Expert Actions

First seen · 7/24/2026, 06:50 AMLatest activity · 7/24/2026, 06:50 AM

LeAct turns actions from expert systems into reasoning supervision by treating the experts’ unstated chains of thought as latent variables. A student samples candidate rationales and retains those that measurably increase its probability of reproducing the expert action. According to the abstract, LeAct reaches the solver’s numerical floor on small enumerable imperfect-information games, is five times closer to the solver than the strongest expert-iteration baseline at larger scale, wins by +60 mbb/g on Flop Hold’em with roughly 1 billion information sets, and uniquely improves over direct imitation on a simulated robotics probe.

Event heat · last 24 hours

No heat snapshots are available in the last 24 hours.

No heat snapshots are available in the last 24 hours.

Reporting Timeline

  1. AggregatorarXiv7/24, 06:50 AMnot independentRepresentative
    LeAct: Learning to Reason from Expert Actions