Daily curated AI news: breakthroughs, product launches, industry trends and more.
Total 157 articles
The Qwen3-TTS series of speech synthesis models, developed by the Tongyi Qwen team, has been officially open-sourced. The series includes models with 1.7B and 0.6B parameters, supporting voice cloning, voice creation, and human-like speech synthesis. It uses an innovative 12Hz multi-codebook speech encoder and a dual-track modeling architecture to achieve efficient speech compression and high-fidelity restoration, with the first audio packet delay as low as 97 milliseconds. The models cover 10 major languages and dialects, including Chinese, English, Japanese, and Korean, and support precise control of voice, emotion, and prosody through natural language commands.
StepFun has open-sourced the 10B parameter multimodal model Step3-VL-10B, which outperforms mainstream large models with 20 times more parameters in various benchmark tests. The model achieves SOTA performance in core dimensions such as visual perception, math competitions, and logical reasoning, thanks to its innovative Parallel Coordinated Reasoning (PaCoRe) mechanism, which significantly enhances complex task handling capabilities.
Zhipu AI has officially open-sourced and released the GLM-4.7-Flash model, which features a hybrid thinking architecture with a total of 30B parameters and only 3B active parameters, providing a high-performance option for lightweight deployment. The model performs excellently on mainstream benchmarks such as SWE-bench Verified and τ²-Bench, surpassing open-source models of the same size to achieve SOTA levels.
ByteDance's AI agent platform "Kouzi" has officially launched version 2.0, introducing four major functional upgrades: Agent Skills, Agent Plan, Agent Office, and Agent Coding, aimed at enhancing AI's capabilities in professional fields.
The multimodal model GLM-Image, jointly developed by Zhipu AI and Huawei, topped the Hugging Face Trending list within 24 hours of being open-sourced. It is the first SOTA model to be fully trained on domestic Ascend chips, showcasing excellent performance in complex visual text generation and long text rendering, particularly in Chinese character generation.
Google has released the TranslateGemma series of open translation models based on the Gemma 3 architecture, offering 4B, 12B, and 27B parameter sizes. The series supports 55 core languages and multimodal image translation. The 12B version outperforms larger baseline models in translation quality, while the 4B model matches the performance of the 12B baseline and runs smoothly on mobile devices.
The Qwen App has officially launched over 400 AI-powered service functions, marking the transition from "chat conversations" to the "AI service era." The app deeply integrates with Alibaba's ecosystem, including Taobao, Alipay, AutoNavi, and Fliggy, enabling users to order food, shop, book flights and hotels, and complete payments within the chat interface without needing to switch apps. The new "Task Assistant" supports multi-step complex task planning, and a learning aid function has been added.
Zhipu AI and Huawei have jointly open-sourced the next-generation image generation model GLM-Image, which is the first SOTA multimodal model to be fully trained on the domestic Ascend Atlas 800T A2 chip. The model uses an innovative "autoregressive + diffusion decoder" hybrid architecture and has achieved the best open-source model performance on the CVTG-2K and LongText-Bench benchmarks, particularly excelling in Chinese character generation tasks.
Self-Variable, a company specializing in embodied intelligent robots, has recently completed a 1 billion yuan A++ funding round, with investors including ByteDance, Sequoia China, and Shenzhen Venture Capital. The company is the only one in China to receive investments from ByteDance, Meituan, and Alibaba, highlighting the market's high recognition of its technological leadership.
Zhipu Technology Co., Ltd. (short for “Zhipu”) has officially listed on the Hong Kong Stock Exchange with the stock code “02513,” becoming the world's first large model stock. The company is dedicated to the development and exploration of general artificial intelligence (AGI) with the core philosophy that “intelligence is and only is our product.” This listing marks a critical step in its development.
Get curated AI news delivered to your inbox daily. Never miss an update.
We never share your email. Unsubscribe anytime.