HERMES proposes a reusable hierarchical labeling substrate for pre-training data mixtures. A Learned Semantic Transform followed by three-stage residual vector quantization assigns each document a coarse-to-fine code, whose prefix length controls granularity across roughly 130,000 cells. On a 1B-parameter model trained on 25B tokens, the hierarchy exposed an interaction that fixed-granularity pipelines could not test: at one prefix length, equal-subbucket coverage combined with top-30% within-bucket quality selection improved a 16-task capability macro-average by +0.0253. At the next finer level, the advantage disappeared as candidate pools became approximately five times smaller.
No heat snapshots are available in the last 24 hours.