AI News (2026/2/2): Step 3.5 Flash Launches! Faster, Stronger, and More Stable Agent Brain, Open Source!

2026年2月2日 10:35

Executive Summary:

Step 3.5 Flash, the new open-source base model from Jieyue Xingchen, is designed specifically for Agent scenarios. The model uses a sparse MoE architecture with 196 billion parameters, activating only about 11 billion parameters per token, and supports a context length of 256K. Key advantages include a maximum inference speed of 350 TPS, performance comparable to closed-source models in Agent tasks and mathematical reasoning, and significant efficiency improvements through MTP-3 technology.

Step 3.5 Flash Launches! Faster, Stronger, and More Stable Agent Brain, Open Source!

News Details

On Monday, February 2, Jieyue Xingchen released the new open-source base model Step 3.5 Flash. This model is specifically designed for Agent scenarios and employs a sparse MoE (Mixture of Experts) architecture with 196 billion parameters. During inference, each token activates only about 11 billion parameters and supports a 256K context length. These features enable Step 3.5 Flash to excel in handling complex tasks while maintaining high inference efficiency.


Key Points

  • Sparse MoE Architecture: Step 3.5 Flash utilizes a sparse MoE architecture, meaning that only a small portion of the expert networks is activated when processing each token. This design not only reduces the consumption of computational resources but also enhances the model's inference speed and stability.

  • High Parameter Count with Low Activation Rate: Despite having 196 billion parameters, the activation parameter count per token is only 11 billion. This combination of high parameter count and low activation rate allows Step 3.5 Flash to maintain strong expressive power while significantly reducing computational costs.

  • Long Context Support: Step 3.5 Flash supports a 256K context length, which is crucial for handling long texts and complex tasks. Long context support enables the model to better understand and generate coherent text content, making it suitable for various Agent scenarios.

  • Inference Speed Up to 350 TPS: Step 3.5 Flash can achieve an inference speed of up to 350 TPS (Tokens Per Second), leading the pack among current open-source base models. The high inference speed ensures that the model can quickly respond to user needs in practical applications.

  • MTP-3 Technology: By introducing MTP-3 technology (Multi-Token Prediction), Step 3.5 Flash can generate multiple tokens in a single prediction. Specifically, each prediction can produce 3 tokens, significantly improving the model's generation efficiency and response speed.


AI-ALL In-Depth Commentary

The launch of Step 3.5 Flash marks a significant step forward in the maturity of open-source AI models for Agent scenarios. The combination of sparse MoE architecture and MTP-3 technology not only enhances the model's performance but also drastically reduces computational costs. This is a major boon for developers and enterprises with limited resources. Additionally, the ability to support long context lengths gives Step 3.5 Flash a significant advantage in handling complex tasks, making it a promising candidate for applications such as dialogue systems and virtual assistants. Active participation from the open-source community will further drive the optimization and expansion of the model's applications, bringing new vitality to the AI ecosystem.

About AI News

We use AI technology to automatically crawl and filter the latest AI news from around the world, providing you with the most valuable industry updates.