Curated reviews of mainstream AI models, agents, and dev tools — capabilities, use cases, and how they compare.
23 article(s) found · Clear tag
Gemini 3.8 Live is a series of native real-time speech dialogue models launched by Google, which includes two variants: Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking. This series employs an en...
Xiaomi-CocktailASR-1 is Xiaomi's open-source Target Speaker ASR (TS-ASR) large model, designed using an end-to-end LLM architecture. It uses a reference speech as a speaker embedding prompt to accurat...
Mureka V9.5 is a new generation AI music generation model launched by Kunlunwanwei, built upon its self-developed MusiCoT music reasoning framework. It first constructs a global musical structure befo...
VibeThinker-3B is a 3-billion-parameter dense reasoning model open-sourced by the AI team at Weibo. Built upon the Qwen2.5-Coder-3B base model, it undergoes an enhanced Spectrum-to-Signal post-trainin...
Leanstral 1.5 is an open-source formal verification large model from Mistral AI, deeply optimized for Lean 4 automated theorem proving. The model adopts a sparse mixture of experts (MoE) architecture ...
GeneBench-Pro is a research-grade benchmark developed by OpenAI, specifically designed to evaluate AI models' ability to handle judgment-intensive analysis in computational biology. The benchmark comp...
yuxinlu1 Gemma4-12B is an open-source coding and Agentic model series fine-tuned by individual developer Lu Yuxin based on Google's Gemma 4 12B instruction model, comprising the V1 Code version and V2...
Claude Sonnet 5 is the most capable agent model in Anthropic's Sonnet series. Its performance in benchmarks for agentic coding, terminal operations, browser search, and computer use approaches that of...
Intern-S2-Preview is Shanghai AI Laboratory's open scientific multimodal LLM preview, delivering trillion-class scientific capability at 35B parameters. Its "general–specialist fusion" training pipeli...
General365 is an open general reasoning benchmark from Meituan's LongCat team, designed to evaluate large language models (LLMs) purely on logical reasoning in everyday scenarios. The benchmark includ...
Get in-depth reviews the moment a new AI model drops. Know its capabilities and use cases.
We never share your email. Unsubscribe anytime.