Curated reviews of mainstream AI models, agents, and dev tools — capabilities, use cases, and how they compare.
Total 464 articles
ZDTaichu5.0-9B is a new generation general-purpose multimodal large model open-sourced by Taichu-AI, specifically designed for understanding the physical world and embodied intelligence scenarios. The...
Iris is a search agent open-sourced by the Xiaohongshu AllSpark team, featuring two versions: 35B (Iris-mini) and 397B (Iris-pro). It utilizes a MoE architecture, with activated parameters of 3B and 1...
Vidu S2 is a real-time video generation and editing model launched by Shengshu Tech, available for open experience upon release. The model consists of two core components: S2-Editing supports real-tim...
Gemini 3.8 Live is a series of native real-time speech dialogue models launched by Google, which includes two variants: Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking. This series employs an en...
Step Audio 3 is a new generation of voice large model series introduced by StepFun. It derives five specialized models—Realtime, ASR, TTS, Gen, and Music—from a single technical foundation, covering t...
YuE2 is an open-source music generation model jointly developed by the Hong Kong University of Science and Technology and the M-A-P team. Its core design concept introduces an explicit symbolic musica...
HiDream-O1-Video-1.0 (shortened as HD-V1) is a native full-modal video generation model launched by HiDream.ai. It supports multi-modal inputs such as text, images, and videos, and can directly genera...
UnifoLM-WLA-1.0 is a 6B-parameter embodied multimodal large model launched by Unitree Robotics, trained on approximately 2500 hours of real-robot operation data. It unifies visual perception, language...
S2 is a multimodal foundation model open-sourced by the Shanghai Artificial Intelligence Lab in December 2025, featuring up to 397B parameters. It is positioned as a "science-savvy" foundation model. ...
Xiaomi-CocktailASR-1 is Xiaomi's open-source Target Speaker ASR (TS-ASR) large model, designed using an end-to-end LLM architecture. It uses a reference speech as a speaker embedding prompt to accurat...
Get in-depth reviews the moment a new AI model drops. Know its capabilities and use cases.
We never share your email. Unsubscribe anytime.