According to the supplied Google Developers Blog abstract, engineers serving the 397B-parameter Qwen 3.5 MoE on Ironwood (TPU7x) used a modular JAX/Pallas stack, hybrid data and expert parallelism, hierarchical reduce-scatter, Batched Ragged Page Attention, and a fused Gated DeltaNet block. The post reports up to a 4.7x inference speedup on prefill-heavy workloads and operation near hardware roofline limits. However, the stated publication date is August 6, 2026, and the supplied material includes no benchmark configuration, baseline, code, or independent verification.
No heat snapshots are available in the last 24 hours.