AI News (2026/5/22): GLM-5.1 High-Speed Version API Launched, GLM-5.1-highspeed

2026年5月22日 06:01

Executive Summary:

The Zhipu Open Platform has released the GLM-5.1 high-speed version API, GLM-5.1-highspeed, which achieves an output speed of 400 tokens/s, setting a new global speed record for large model APIs. This API, jointly developed by Zhipu and the TileRT team, is the first to achieve flagship-level capabilities and ultra-low latency in domestic large models. It is now available to some enterprise customers.

[Zhipu Launches GLM-5.1 High-Speed Version API, Setting a New Global Speed Record for Large Model APIs]

News Details

The Zhipu Open Platform released the GLM-5.1 high-speed version API, GLM-5.1-highspeed, on May 22. This new version of the API achieves an output speed of 400 tokens/s, significantly enhancing the performance of large models in real-time applications. GLM-5.1-highspeed was jointly developed by Zhipu and the TileRT team, aiming to address the latency issues faced by large models in practical applications. The API is suitable for AI programming, real-time interaction, and real-time voice applications, which require extremely low latency.


Key Points

  • 400 tokens/s: The output speed of GLM-5.1-highspeed reaches 400 tokens/s, setting the highest speed record among global large model APIs. This breakthrough significantly reduces the response time of models in real-time applications.

  • Low Latency: By optimizing the Inference process and RAG technology, GLM-5.1-highspeed achieves low-latency output, ensuring stability and reliability in high-concurrency scenarios. This is particularly important for real-time interaction applications that require quick responses.

  • Available to Some Enterprise Customers: Currently, GLM-5.1-highspeed is available to some enterprise customers. These enterprises can use the API to develop and test applications in AI programming, real-time interaction, and real-time voice processing. The availability will be gradually expanded in the future.


AI-ALL In-Depth Commentary

The launch of GLM-5.1-highspeed marks a significant advancement in performance optimization for domestic large models. An output speed of 400 tokens/s not only enhances user experience but also provides developers with more possibilities, especially in the fields of real-time interaction and voice processing. This technological breakthrough is expected to facilitate the implementation of more high-demand applications, further enriching the AI ecosystem. Additionally, the collaboration between Zhipu and the TileRT team demonstrates the potential for domestic enterprises in collaborative innovation, offering new ideas and directions for the future development of large models.

About AI News

We use AI technology to automatically crawl and filter the latest AI news from around the world, providing you with the most valuable industry updates.