LeAct turns actions from expert systems into reasoning supervision by treating the experts’ unstated chains of thought as latent variables. A student samples candidate rationales and retains those that measurably increase its probability of reproducing the expert action. According to the abstract, LeAct reaches the solver’s numerical floor on small enumerable imperfect-information games, is five times closer to the solver than the strongest expert-iteration baseline at larger scale, wins by +60 mbb/g on Flop Hold’em with roughly 1 billion information sets, and uniquely improves over direct imitation on a simulated robotics probe.
No heat snapshots are available in the last 24 hours.