Daily curated AI news: breakthroughs, product launches, industry trends and more.
Total 157 articles
Step 3.5 Flash, the new open-source base model from Jieyue Xingchen, is designed specifically for Agent scenarios. The model uses a sparse MoE architecture with 196 billion parameters, activating only about 11 billion parameters per token, and supports a context length of 256K. Key advantages include a maximum inference speed of 350 TPS, performance comparable to closed-source models in Agent tasks and mathematical reasoning, and significant efficiency improvements through MTP-3 technology.
Yushu Technology has announced the open-source release of the visual-language-action large model UnifoLM-VLA-0. Based on the Qwen2.5-VL-7B architecture, the model has been trained with 340 hours of real-world data, integrating 2D/3D spatial perception and dynamics prediction capabilities to overcome the limitations of traditional VLMs in physical interactions.
SenseTime has open-sourced the multimodal autonomous reasoning model SenseNova-MARS, offering 8B and 32B versions. The model outperforms Gemini-3-Pro (69.06 points) and GPT-5.2 (67.64 points) with a comprehensive score of 69.74 points in core benchmarks such as MMSearch and HR-MMSearch, becoming the first open-source Agentic VLM to support dynamic visual reasoning and deep integration of image and text search.
KUNLUN's Skywork AI has announced the open-sourcing of its latest video generation model, SkyReels-V3, which integrates three key functions—reference image to video, video extension, and audio-driven virtual characters—into a single architecture. The image-to-video technology outperforms mainstream models like Vidu and Kling in reference consistency (0.6698) and visual quality (0.8119); the video extension supports scene transitions and narrative expansion; and the virtual character module can generate long videos with synchronized audio and visuals.
OpenAI has launched a free research collaboration platform called Prism, based on GPT-5.2, aimed at addressing the fragmentation of research tools. Prism integrates a cloud-based LaTeX editor, supports real-time collaboration with an unlimited number of co-authors, and provides AI assistance throughout the entire process, from abstract to acknowledgments.
Alibaba Tongyi has open-sourced the 6B parameter non-distilled base model Z-Image, aimed at addressing the issues of单一 and homogenization in AI-generated art. The model supports diverse style generation from realistic to anime and has been optimized for better compatibility with fine-tuning methods like LoRA and ControlNet.
On January 27, the Dark Side of the Moon team released Kimi K2.5, the most intelligent and versatile open-source model to date. The model achieves open-source SOTA levels in various benchmark tests, including Agent tasks, code generation, and visual understanding (images/videos), supporting multimodal input and four working modes. It innovatively introduces "Agent cluster" capabilities, enabling the model to autonomously create up to 100 clones to handle complex tasks in parallel, with efficiency improvements of up to 4.5 times.
Alibaba Cloud has officially released its largest and most capable inference model, Qwen3-Max-Thinking. The model has over 1 trillion parameters and 36T Tokens of pre-training data, setting new records in multiple international benchmark tests. Qwen3-Max-Thinking innovatively adopts a test-time expansion mechanism, enhancing inference performance while being more economical.
Tencent's HunYuan team has officially released the HunYuan Image 3.0 model, which has a total of 80 billion parameters and uses a Mixture of Experts (MoE) architecture. The model supports various visual creation functions such as image editing, multi-image fusion, and more. Users can perform operations like adding, deleting, modifying, style transformation, and old photo restoration through natural language commands. The model can intelligently understand image content and generate detailed editing steps.
Zhipu AI has officially launched the "AI Learning Companion" AI learning assistant, and the first batch of user experience slots are now open for application. The product reshapes the learning experience through intelligent summarization, interactive visual cards with AI tutor Q&A, and precise question generation with a "learn-practice-test" closed loop.
Get curated AI news delivered to your inbox daily. Never miss an update.
We never share your email. Unsubscribe anytime.