AI Model Library: LLMs, Agents & Dev Tools

Curated reviews of mainstream AI models, agents, and dev tools — capabilities, use cases, and how they compare.

Expert analysisUpdated regularlyRSS Feed

Tag: Benchmark

82 article(s) found · Clear tag

1 months ago

Qwen-CUA – The Native Computer Use Agent Introduced by Alibaba Qwen and Others

Qwen-CUA is a native Computer Use Agent introduced jointly by the Qwen team and XLang Lab. Based on a 397B-A17B Mixture-of-Experts (MoE) architecture, it perceives the interface state solely through s...

AI AgentLLMBenchmark
Aug 5, 2026Read more →
1 months ago

Orchard – Microsoft Research's Open-Source Agentic AI Modeling Framework

Orchard is an open-source Agentic AI modeling framework introduced by Microsoft Research. At its core is the Kubernetes-based Orchard Env environment service, which enables cross-domain reuse of sandb...

AI AgentModel InferenceBenchmark
Aug 4, 2026Read more →
1 months ago

E-Bench – A Benchmark for Evaluating AI Agents Introduced by Tencent HunYuan and Others

E-Bench is a multi-step tool usage evaluation benchmark introduced jointly by the Tencent HunYuan team, along with the AI Research Institute (AIR) at Tsinghua University and Southeast University. It i...

AI AgentTool CallingBenchmark
Aug 4, 2026Read more →
1 months ago

SearchOS – A Multi-Agent Search Collaboration Framework Open-Sourced by Ant Group and Others

SearchOS is a multi-agent search collaboration framework jointly open-sourced by the H瓴 Artificial Intelligence School at Renmin University of China and Ant Group. Drawing inspiration from the design ...

AI AgentEmbedding & RAGBenchmark
Aug 1, 2026Read more →
1 months ago

UniWorld-View – RabbitZoom Intelligence Collaborates with Peking University and Others to Open-Source a World Model

UniWorld-View is an open-source world model jointly developed by RabbitZoom Intelligence, Peking University, and the鹏城实验室 (Pengcheng Laboratory). It has topped the WorldScore world model evaluation le...

Video AIAI CodingBenchmark
Aug 1, 2026Read more →
1 months ago

AlphaEval – A Production-Grade Agent Evaluation Framework Introduced by KuaFuAI and Others

AlphaEval is a production-grade Agent evaluation framework jointly introduced by SII-GAIR, Shanghai Jiao Tong University, and KuaFuAI. This framework covers 94 production tasks across seven real-world...

AI AgentBenchmarkAI Tools
Aug 1, 2026Read more →
1 months ago

MAI-Cyber-1-Flash – Microsoft's First AI Model for Cybersecurity

MAI-Cyber-1-Flash is Microsoft's first AI model built from scratch specifically for cybersecurity tasks, based on the multi-agent system MDASH. It can efficiently handle approximately 90% of cybersecu...

AI AgentAI SafetyBenchmark
Aug 1, 2026Read more →
1 months ago

Midjourney V8.2 – Midjourney's Most Advanced AI Image Generation Model

Midjourney V8.2 is the latest version of the AI image generation model released by Midjourney Inc., achieving comprehensive upgrades in prompt understanding, style reference logic, and image generatio...

Benchmark
Aug 1, 2026Read more →
1 months ago

PerceptionBench – Moonshot AI's Open-Source Visual Perception Diagnostic Benchmark

PerceptionBench is an open-source visual perception diagnostic benchmark launched by Moonshot AI, focusing on fine-grained evaluation of atomic-level visual perception capabilities. Starting from the ...

MultimodalLLMBenchmark
Aug 1, 2026Read more →
1 months ago

FeyNoBg – An Open-Source Automatic Background Removal Model from Feyn Labs

FeyNoBg is an open-source automatic background removal model developed by Feyn Labs, built upon the BiRefNet architecture with deep optimization, and features 263 million parameters. The model achieve...

BenchmarkOpen Source
Aug 1, 2026Read more →
Page 5 of 9 (82 articles total)

Subscribe to AI Model Reviews

Get in-depth reviews the moment a new AI model drops. Know its capabilities and use cases.

We never share your email. Unsubscribe anytime.