DeepSeek Releases V4.1-Flash, Slashing KV Cache Demands for AI Agents
First seen · 9/10/2026, 08:40 PMLatest activity · 9/10/2026, 08:40 PM
DeepSeek has open-sourced V4.1-Flash, a 552-billion-parameter multimodal model that activates just 16 billion parameters per token. By shrinking KV cache memory requirements to a quarter of its predecessor's footprint, the architecture significantly reduces the runtime overhead of long-horizon AI agents. Released under an MIT license, the model also narrowly edges out frontier systems like Opus 5 and GPT-5.6 Sol on the DeepSWE coding benchmark.
Event heat · last 24 hours
There are 8 persisted snapshots in the last 24 hours. Peak heat was 10 at 9/12, 14:00; latest heat is 10.