Pıer
TidesCurrentsHarbor LightsLabBottlesAshore
Pıer

Navigation

  • Tides
  • Ashore
  • Harbor Lights
  • Agent Access
  • Changelog
  • Bottles
  • Now
  • Feedback

External links

GitHubCloudborne ↗

© 2026 Pier.

WatchingResearchWatching0 independent reports0

CausalMix: Data Mixture as Causal Inference for Language Model Training

First seen · 7/2/2026, 12:00 PMLatest activity · 7/2/2026, 12:00 PM

CausalMix frames language-model data-mixture optimization as a causal-inference problem. Statistical characteristics of the data pool are treated as covariates, while domain mixture is the treatment. The authors fit a causal model using 512 training runs of Qwen2.5-0.5B, estimate conditional average treatment effects, and extrapolate a mixture to an 800K-data pool and 7B-model training. They also extend the framework to long chain-of-thought data with Qwen3-4B-Base. According to the paper abstract, CausalMix consistently outperforms RegMix and other baselines across downstream tasks and provides visual interpretation through a CATE Interpreter.

Event heat · last 24 hours

No heat snapshots are available in the last 24 hours.

No heat snapshots are available in the last 24 hours.

Reporting Timeline

  1. AggregatorarXiv7/1, 11:56 PMnot independent
    CausalMix: Data Mixture as Causal Inference for Language Model Training
  2. AggregatorHuggingFace Daily Papers7/2, 12:00 PMnot independentRepresentative
    CausalMix: Data Mixture as Causal Inference for Language Model Training