This paper introduces Graded Large Language Models (GLLMs), an algebraic framework that assigns grades to a transformer's representation space and propagates the resulting weighted scalar action through embeddings, self-attention, and the training objective. The authors characterize beneficial grades using geometric invariant theory, derive closed-form optimality conditions through coincident moment maps, and claim a minimax separation for level-stratified targets. Because grading is absorbed into learned parameters after training, a GLLM reportedly compiles into a standard transformer with unchanged architecture, asymptotic complexity, and inference cost.
No heat snapshots are available in the last 24 hours.