Pıer
TidesCurrentsHarbor LightsLabBottlesAshore
Pıer

Navigation

  • Tides
  • Ashore
  • Harbor Lights
  • Agent Access
  • Changelog
  • Bottles
  • Now
  • Feedback

External links

GitHubCloudborne ↗

© 2026 Pier.

Read original
IT之家·Sep 10, 2026, 11:47 AM

DeepSeek V4.1 Flash Deployed on China's Supercomputing Network with Asymmetric MoE Architecture

Original title:DeepSeek V4.1 Flash 模型上线国家超算互联网

Models84

IT Home, September 10 News — The DeepSeek V4.1 Flash model was launched today on China's National Supercomputing Internet platform. Users can log in to the official National Supercomputing Internet website and access the "Model Services" page from the homepage to quickly open the API call interface for the DeepSeek V4.1 Flash model.

According to IT Home, DeepSeek V4.1 Flash is a 552B-parameter MoE model featuring a brand-new Causal-Encoder-Decoder architecture with asymmetric inputs and outputs—activating only 8B parameters for input and 16B for output—making its cost significantly lower than known models of comparable size.

At the same time, V4.1 Flash adopts a new pre-training approach and has undergone larger-scale reinforcement learning post-training. In benchmark tests, it successfully surpassed the intelligence levels of several flagship models, including DeepSeek V4 Pro.

Furthermore, DeepSeek V4.1 Flash substantially reduces the size of the KV Cache. Compared to the previous-generation model, its HBM requirement has been reduced to 1/4 and its SSD requirement to 1/8. In Agent use cases, cache-hit expenses often represent a high proportion of costs; the compression of the KV Cache drastically lowers the operational costs for Agent-based tasks.

Why it's worth reading

It demonstrates how asymmetric MoE architectures and aggressive KV cache compression can dramatically reduce the operational cost of 500B+ parameter models in agent workflows.

Tags

DeepSeekMoEKV-Cache超算互联网大模型推理Agent模型架构

Score breakdown

  • Novelty82
  • Impact84
  • Practicality88
  • Credibility78
  • Timeliness85