AI News (2026/5/25): BitCPM-CANN, the First 1.58-bit Ternary Large Model Trained End-to-End on Huawei Ascend
Executive Summary:
FaceWall AI, in collaboration with Tsinghua University and the OpenBMB open-source community, has released the BitCPM-CANN model, China's first 1.58-bit ternary large model fully trained end-to-end on the Huawei Ascend domestic computing platform. The model comes in four sizes, ranging from 0.5B to 8B parameters, and uses a quantization-aware training approach. This results in a 6x reduction in memory usage during inference while maintaining a model capability retention rate of 90%–97.2%.
FaceWall AI and Tsinghua University Open-Source BitCPM-CANN Model
News Details
FaceWall AI, in collaboration with Tsinghua University and the OpenBMB open-source community, released the BitCPM-CANN model on May 25. This is China's first 1.58-bit ternary large model fully trained end-to-end on the Huawei Ascend domestic computing platform. The BitCPM-CANN model includes four different sizes, with parameter counts of 0.5B, 2B, 4B, and 8B. The model uses quantization-aware training technology during the training phase to significantly reduce memory usage during inference while maintaining high model performance.
Key Points
Quantization-Aware Training: The BitCPM-CANN model employs quantization-aware training technology during the training phase, allowing it to run with 1.58-bit precision during inference. This technology not only drastically reduces the storage requirements of the model but also significantly enhances inference efficiency.
Memory Optimization: Through quantization-aware training, BitCPM-CANN can achieve a 6x reduction in memory usage during inference. This means that under the same hardware conditions, larger-scale models can be deployed or multiple model instances can run simultaneously, thereby improving resource utilization.
Performance Retention: Despite the low-precision quantization, the BitCPM-CANN model retains its capabilities at a rate of 90%–97.2%. This performance level ensures that the model still performs well in practical applications, especially on resource-constrained edge devices.
Multiple Size Versions: BitCPM-CANN offers four different parameter sizes (0.5B, 2B, 4B, and 8B), catering to various application scenarios. Developers can choose the appropriate model version based on specific application requirements and hardware conditions.
Domestic Computing Platform: The model is fully trained end-to-end on the Huawei Ascend domestic computing platform, demonstrating the competitiveness and development potential of the domestic NPU ecosystem in the AI field. This is significant for promoting the self-reliance and control of AI technology in China.
AI-ALL In-Depth Commentary
The release of BitCPM-CANN marks a significant advancement for China in the development of edge AI models and the application of domestic computing platforms. By adopting quantization-aware training technology, the model not only significantly reduces memory usage but also maintains a high performance level, providing strong support for resource-constrained edge devices. Additionally, the design of multiple size versions allows developers to choose flexibly based on specific needs, further enhancing the model's applicability. Most importantly, the end-to-end training on the Huawei Ascend domestic computing platform showcases the maturity and competitiveness of the domestic NPU ecosystem, which is crucial for promoting the self-reliance and widespread application of AI technology.

