AI Model Library: LLMs, Agents & Dev Tools

Curated reviews of mainstream AI models, agents, and dev tools — capabilities, use cases, and how they compare.

Expert analysisUpdated regularlyRSS Feed

Tag: Benchmark

82 article(s) found · Clear tag

3 months ago

Agora-1 – Odyssey's First Multi-Agent World Model

Agora-1 is Odyssey's first multi-agent world model, breaking past the single-user limits of traditional world models by enabling humans and AI to interact in the same real-time generated world simulat...

AI AgentEmbodied AIBenchmark
Jun 21, 2026Read more →
3 months ago

WBench – Meituan’s Interactive Video World Model Multi-Turn Benchmark

WBench is Meituan LongCat’s first systematic multi-turn benchmark for interactive video world models—289 test cases, 1,058 interaction rounds, six scene types (nature, city, indoor, workspace, fantasy...

Video AIModel InferenceBenchmark
Jun 21, 2026Read more →
3 months ago

SenseNova-Skills – SenseTime's Open-Source Modular AI Office Skills Library

SenseNova-Skills is an open-source modular AI office skills library from the OpenSenseNova (SenseTime) team, designed for Agent runtimes. It breaks down four core office capabilities—infographic gener...

AI AgentMultimodalBenchmark
Jun 21, 2026Read more →
3 months ago

Rodin Gen-2.5 – Hyper3D's 10M-Polygon AI 3D Model Generator

Rodin Gen-2.5 from Hyper3D (影眸科技) is marketed as the first commercial AI 3D tool to generate 10M+ polygons directly, built on SIGGRAPH 2025 Best Paper technology. Text, single-image, or multi-view inp...

Image GenerationBenchmark
Jun 21, 2026Read more →
3 months ago

PlanningBench – Open LLM Planning Evaluation Framework by Tencent Hunyuan and Partners

PlanningBench is an open framework from Tencent Hunyuan with Renmin University of China Gaoling School of Artificial Intelligence and partners, focused on evaluating and training large language model ...

AI AgentLLMBenchmark
Jun 21, 2026Read more →
3 months ago

PawBench – Tongyi Lab's General Agent Evaluation Benchmark

PawBench is a general agent evaluation benchmark from Tongyi Lab for personal assistant and agent scenarios, evaluating base models and runtime frameworks (Harness) together. PawBench v1.0 includes 15...

AI AgentModel InferenceBenchmark
Jun 21, 2026Read more →
3 months ago

MiniCPM5-1B – Edge Text Foundation Model Open-Sourced by ModelBest and Tsinghua

MiniCPM5-1B is a 1B-parameter edge text foundation model jointly released by ModelBest, Tsinghua University, and the OpenBMB open-source community. On the AA-Index composite benchmark, it scored 17.9 ...

Model InferenceBenchmarkOpen Source
Jun 21, 2026Read more →
3 months ago

MAI-Image-2.5 – Microsoft's Flagship Text-to-Image Model

MAI-Image-2.5 is a flagship text-to-image model from Microsoft Research and the strongest release in the MAI-Image family. On the Arena text-to-image leaderboard it climbed to #3 with 1,254 points—72 ...

MultimodalImage GenerationBenchmark
Jun 21, 2026Read more →
3 months ago

Hy-Memory – Tencent Hunyuan's Agent Memory Plugin

Hy-Memory is a professional memory plugin from Tencent's Hunyuan team, designed for long-term collaborative Agents such as OpenClaw. Through a 6-layer memory framework, System1/System2 dual-system pro...

AI AgentEmbedding & RAGBenchmark
Jun 21, 2026Read more →
3 months ago

Gemini-SQL2 – Google's Text-to-SQL AI Model

Gemini-SQL2 is Google Research's latest AI model dedicated to Text-to-SQL. Built on Gemini 3.1 Pro with targeted post-training, it tops the BIRD benchmark single-model track at 80.04% execution accura...

LLMBenchmark
Jun 21, 2026Read more →
Page 8 of 9 (82 articles total)

Subscribe to AI Model Reviews

Get in-depth reviews the moment a new AI model drops. Know its capabilities and use cases.

We never share your email. Unsubscribe anytime.