AI News (2026/5/21): ZCube, the Next-Generation Large Model Inference Network Architecture
Executive Summary:
ZCube, a network architecture jointly developed by ZhiPu, Yuxun Network, and Tsinghua University, addresses the congestion issues in PD (Parameter Distribution) separated inference by eliminating the Spine layer and adopting a flat topology with single/multi-track hybrid access. Practical tests with GLM-5.1 coding show that ZCube reduces the cost of switches and optical modules by 33%, increases GPU inference throughput by 15%, and reduces the P99 Time to First Token (TTFT) latency by 40.6%, providing a robust foundation for the next generation of large-scale inference clusters.
ZCube, the Next-Generation Large Model Inference Network Architecture
Detailed Information
ZCube, a network architecture jointly developed by ZhiPu, Yuxun Network, and Tsinghua University, aims to solve the common congestion problems in PD (Parameter Distribution) separated inference. By eliminating the traditional Spine layer and adopting a flat network topology with single/multi-track hybrid access, ZCube significantly improves inference efficiency and reduces costs. This innovative solution has been verified through practical tests with GLM-5.1 coding, providing a solid technical foundation for future large-scale inference clusters.
Key Points
Flat Topology: ZCube eliminates the traditional Spine layer and adopts a flat network topology. This design reduces network layers, lowering data transmission latency and congestion risks.
Single/Multi-Track Hybrid Access: ZCube supports both single-track and multi-track hybrid access methods, allowing for flexible network configuration based on actual needs. This not only enhances network flexibility but also optimizes resource utilization.
Cost Efficiency: In practical tests with GLM-5.1 coding, ZCube reduced the cost of switches and optical modules by 33%, while increasing GPU inference throughput by 15%. Additionally, the P99 Time to First Token (TTFT) latency was reduced by 40.6%, significantly improving overall performance.
AI-ALL In-Depth Commentary
The launch of the ZCube network architecture marks a significant technological breakthrough in the field of large model inference. By eliminating the Spine layer and adopting a flat topology, ZCube effectively addresses congestion issues in PD separated inference, enhancing data transmission efficiency. The introduction of single/multi-track hybrid access further enhances network flexibility and resource utilization. These improvements not only substantially reduce hardware costs but also significantly boost GPU inference throughput and TTFT performance. For AI developers and enthusiasts, ZCube offers reliable technical support for building the next generation of large-scale inference clusters, potentially driving further adoption and development of AI applications.
