Curated reviews of mainstream AI models, agents, and dev tools — capabilities, use cases, and how they compare.
82 article(s) found · Clear tag
Agora-1 is Odyssey's first multi-agent world model, breaking past the single-user limits of traditional world models by enabling humans and AI to interact in the same real-time generated world simulat...
WBench is Meituan LongCat’s first systematic multi-turn benchmark for interactive video world models—289 test cases, 1,058 interaction rounds, six scene types (nature, city, indoor, workspace, fantasy...
SenseNova-Skills is an open-source modular AI office skills library from the OpenSenseNova (SenseTime) team, designed for Agent runtimes. It breaks down four core office capabilities—infographic gener...
Rodin Gen-2.5 from Hyper3D (影眸科技) is marketed as the first commercial AI 3D tool to generate 10M+ polygons directly, built on SIGGRAPH 2025 Best Paper technology. Text, single-image, or multi-view inp...
PlanningBench is an open framework from Tencent Hunyuan with Renmin University of China Gaoling School of Artificial Intelligence and partners, focused on evaluating and training large language model ...
PawBench is a general agent evaluation benchmark from Tongyi Lab for personal assistant and agent scenarios, evaluating base models and runtime frameworks (Harness) together. PawBench v1.0 includes 15...
MiniCPM5-1B is a 1B-parameter edge text foundation model jointly released by ModelBest, Tsinghua University, and the OpenBMB open-source community. On the AA-Index composite benchmark, it scored 17.9 ...
MAI-Image-2.5 is a flagship text-to-image model from Microsoft Research and the strongest release in the MAI-Image family. On the Arena text-to-image leaderboard it climbed to #3 with 1,254 points—72 ...
Hy-Memory is a professional memory plugin from Tencent's Hunyuan team, designed for long-term collaborative Agents such as OpenClaw. Through a 6-layer memory framework, System1/System2 dual-system pro...
Gemini-SQL2 is Google Research's latest AI model dedicated to Text-to-SQL. Built on Gemini 3.1 Pro with targeted post-training, it tops the BIRD benchmark single-model track at 80.04% execution accura...
Get in-depth reviews the moment a new AI model drops. Know its capabilities and use cases.
We never share your email. Unsubscribe anytime.