This paper introduces C-MTP, a direct-supervision method for continuous chain-of-thought training. Each latent representation is trained to approximate the average embedding of the CoT tokens being compressed, avoiding the autoregressive generation required by earlier indirect methods. C-MTP outperforms a prior direct method that approximates compressed-token distributions and is competitive with slower indirect approaches on simplified traces shorter than 100 tokens. However, when evaluated on complex tasks requiring several hundred reasoning tokens, both direct and indirect methods suffer an approximately 65% performance drop. The result highlights a major length-related limitation of current continuous CoT methods. Code and checkpoints are released.
No heat snapshots are available in the last 24 hours.