AI News (2026/3/12): NVIDIA's Most Powerful Open-Source Weight AI Model: Nemotron 3 Super

2026年3月12日 10:30

Executive Summary:

NVIDIA has released its latest open-source weight AI model, Nemotron 3 Super, which boasts 120 billion parameters and uses the MoE architecture, activating only 1.2 billion parameters during inference. The model shows significant improvements in throughput and accuracy compared to its predecessor, and it features a 1 million Token ultra-long context window, specifically designed for multi-agent systems to effectively address "context explosion" and "inference tax" issues.

Details

NVIDIA released its latest open-source weight AI model, Nemotron 3 Super, on March 12. The model has 120 billion parameters and uses the MoE (Mixture of Experts) architecture, but only activates 1.2 billion parameters during inference. This design not only enhances the model's efficiency but also significantly boosts its performance. Compared to the previous generation, Nemotron 3 Super's throughput has increased by 5 times, and its accuracy has doubled. Additionally, the model features a 1 million Token ultra-long context window, specifically designed for multi-agent systems.


Key Points

  • 120 Billion Parameters: Nemotron 3 Super has a 120 billion parameter MoE architecture, but only activates 1.2 billion parameters during inference. This design maintains the model's strong expressive power while significantly reducing the demand for computational resources.

  • 5x Throughput Improvement: Compared to the previous generation, Nemotron 3 Super's throughput has increased by 5 times. This means that under the same hardware conditions, the new model can process more data and tasks, significantly improving production efficiency.

  • Doubled Accuracy: Nemotron 3 Super's accuracy is twice that of its predecessor. This improvement is crucial for high-precision applications such as natural language processing and image recognition.

  • 1 Million Token Ultra-Long Context Window: The model features a 1 million Token ultra-long context window, effectively addressing the "context explosion" problem. This is particularly important for handling long texts and complex dialog tasks.

  • Latent MoE Technology: Nemotron 3 Super introduces latent MoE technology, a new hybrid architecture that combines the advantages of Mamba-Transformer and MoE. This technology supports multi-Token prediction, further enhancing the model's performance and flexibility.

  • Designed for Multi-Agent Systems: Nemotron 3 Super is specifically designed for multi-agent systems, better supporting Agentic Workflow. This allows multiple agents to work together, improving the overall system's efficiency and response speed.


AI-ALL In-Depth Analysis

The release of Nemotron 3 Super marks another significant advancement by NVIDIA in the field of open-source AI models. By adopting the MoE architecture and latent MoE technology, the model not only reaches a new level in parameter scale but also achieves significant breakthroughs in inference efficiency and accuracy. The 1 million Token ultra-long context window solves the long-standing "context explosion" issue in multi-agent systems, making the handling of complex tasks more efficient and accurate.

For AI developers and enthusiasts, Nemotron 3 Super provides a powerful toolkit for building more complex multi-agent systems and application scenarios. The improvements in throughput and accuracy will help drive the practical application of AI technology in more fields, especially in scenarios requiring high precision and efficient processing.

Furthermore, the open-source nature of Nemotron 3 Super will further promote technical exchanges and innovation within the AI community. Developers can customize and optimize this model, contributing to the advancement and development of the entire AI ecosystem.

About AI News

We use AI technology to automatically crawl and filter the latest AI news from around the world, providing you with the most valuable industry updates.