Read original
IT之家Products87

Huawei Announces Day-One Ascend Support for Ant Group’s Ling-3.0-flash and Debuts CANN PyPTO

Original title:华为官宣昇腾 0 Day 支持蚂蚁百灵 Ling-3.0-flash 模型,全新算子编程框架 CANN PyPTO 首秀

IT Home, August 6 -- Ant Group's Bailing officially open-sourced its next-generation native hybrid reasoning model, Ling-3.0-flash, yesterday. The model focuses on extremely high intelligence density, a full-process Agent closed loop, and low-cost large-scale deployment. Huawei officially announced today that Ascend has simultaneously completed 0-Day adaptation of the Bailing Ling-3.0-flash model. During the adaptation process, it also introduced the CANN PyPTO operator programming framework for the first time, which can shorten the delivery cycle for complex fused operators and efficiently complete operator integration and inference framework performance tuning. A complete deployment project has also been made available. This adaptation is fully compatible with the Ascend A2 and A3 series products and comes with an official standardized deployment solution. Developers can quickly launch inference services through the official images. The hardware requirements are eight cards for Ascend A2 series products and four cards for Ascend A3 series products.

The CANN PyPTO programming framework is built on PTO ISA (virtual instruction set) and uses a Tile programming paradigm. Developers precisely define computational logic through Tensor-level APIs, while the framework automatically handles the entire process of tiling and scheduling, data movement, and code generation.

In addition, CANNBot PyPTO Agent has built a multi-agent framework that automatically advances the entire operator development process through seven stages. Each stage is executed collaboratively by specialized sub-agents with clearly defined responsibilities. All scheduling is handled by the unified orchestrator pypto-op-orchestrator. The sub-agents exchange information through a memory mechanism and machine-readable states, ensuring seamless context transfer.

The entire process begins with a natural-language operator requirements input and proceeds through automated generation, verification, and tuning, ultimately producing a board-deployable operator that meets accuracy requirements and has been performance-optimized. No manual intervention is required for handoffs between stages. This workflow not only represents a comprehensive upgrade to the operator programming experience but also significantly shortens the operator delivery cycle. IT Home has included the logic for each stage below:

Related reading: “Ant Group Releases Bailing Native Hybrid Reasoning Model Ling-3.0-flash: 124B Total Parameters, with Capabilities Comparable to Those of the Previous-Generation Flagship Ring-2.6-1T.”

Why it's worth reading

The announcement matters now because Ling-3.0-flash received Ascend support immediately after release, while PyPTO targets the operator-development bottleneck that often determines deployment speed and engineering cost on domestic AI hardware.

Tags

昇腾Ling-3.0-flashCANNPyPTO算子编程多智能体推理部署