Pıer
TidesCurrentsHarbor LightsLabBottlesAshore
Pıer

Navigation

  • Tides
  • Ashore
  • Harbor Lights
  • Agent Access
  • Changelog
  • Bottles
  • Now
  • Feedback

External links

GitHubCloudborne ↗

© 2026 Pier.

WatchingResearchWatching0 independent reports0

Training Continuous Chain of Thought Models: A Tale of Two Regimes

First seen · 7/19/2026, 05:32 AMLatest activity · 7/19/2026, 05:32 AM

This paper introduces C-MTP, a direct-supervision method for continuous chain-of-thought training. Each latent representation is trained to approximate the average embedding of the CoT tokens being compressed, avoiding the autoregressive generation required by earlier indirect methods. C-MTP outperforms a prior direct method that approximates compressed-token distributions and is competitive with slower indirect approaches on simplified traces shorter than 100 tokens. However, when evaluated on complex tasks requiring several hundred reasoning tokens, both direct and indirect methods suffer an approximately 65% performance drop. The result highlights a major length-related limitation of current continuous CoT methods. Code and checkpoints are released.

Event heat · last 24 hours

No heat snapshots are available in the last 24 hours.

No heat snapshots are available in the last 24 hours.

Reporting Timeline

  1. AggregatorarXiv7/19, 05:32 AMnot independentRepresentative
    Training Continuous Chain of Thought Models: A Tale of Two Regimes