The paper introduces Kimi K3, a 2.8-trillion-parameter Mixture-of-Experts model with 104 billion activated parameters per token, native vision, and a 1-million-token context window. Its architecture combines Kimi Delta Attention, Attention Residuals, and Stable LatentMoE, which routes each token through 16 of 896 experts. Post-training uses reinforcement learning across general, agentic, and coding tasks with multiple reasoning-effort levels. The authors report roughly 2.5x better overall scaling efficiency than Kimi K2 and claim frontier-level results across coding, agents, knowledge, reasoning, and vision. Full model weights are released, although the abstract says performance still trails leading proprietary systems.
No heat snapshots are available in the last 24 hours.