AI Model Library: LLMs, Agents & Dev Tools

Curated reviews of mainstream AI models, agents, and dev tools — capabilities, use cases, and how they compare.

Expert analysisUpdated regularlyRSS Feed

Tag: Benchmark

82 article(s) found · Clear tag

1 weeks ago

Step Audio 3 – StepFun's Voice Large Model Series

Step Audio 3 is a new generation of voice large model series introduced by StepFun. It derives five specialized models—Realtime, ASR, TTS, Gen, and Music—from a single technical foundation, covering t...

Speech AIModel InferenceBenchmark
Sep 15, 2026Read more →
1 weeks ago

HiDream-O1-Video-1.0: In-Depth Evaluation of a Native Full-Modal Video Generation Model

HiDream-O1-Video-1.0 (shortened as HD-V1) is a native full-modal video generation model launched by HiDream.ai. It supports multi-modal inputs such as text, images, and videos, and can directly genera...

Speech AIVideo AIBenchmark
Sep 15, 2026Read more →
2 weeks ago

EchoWM – Interactive Audio-Visual World Model Open-Sourced by JD.com

EchoWM is an interactive audio-visual world model open-sourced by the JD.com Exploration Research Institute. It continues the native audio-visual generation capabilities of JoyAI-Echo, allowing users ...

Speech AIEmbedding & RAGBenchmark
Sep 11, 2026Read more →
2 weeks ago

LLaDA-Image – A Unified Image Generation and Editing Model Open-Sourced by Ant Group

LLaDA-Image is a 6B parameter unified image generation and editing model open-sourced by the inclusionAI Lab at Ant Group. This model adopts an innovative training approach, first pre-training purely ...

Image GenerationModel InferenceBenchmark
Sep 8, 2026Read more →
2 weeks ago

WikiSkill – Google's Agent Skill Evolution Framework

WikiSkill is an Agent skill evolution framework introduced by Google Research. It employs a three-tier architecture to separately store raw execution traces, structured knowledge, and executable skill...

AI AgentBenchmarkEdge Deployment
Sep 8, 2026Read more →
2 weeks ago

Claude Commerce Agents: A Comprehensive Review of Anthropic's Open-Source E-commerce Agent

Claude Commerce Agents is a set of e-commerce agent (Agent) solutions officially open-sourced by Anthropic in June 2026. It includes two core components: a consumer-facing shopping agent and a merchan...

AI AgentBenchmarkAI Tools
Sep 7, 2026Read more →
3 weeks ago

In-Depth Review of Claude Fable 5.1 – Anthropic's Flagship Large Model

Claude Fable 5.1 is a flagship large model launched by Anthropic, positioned as one of the high-performance AI systems available to the public. Designed for complex reasoning, long-running Agent tasks...

AI AgentModel InferenceBenchmark
Sep 4, 2026Read more →
3 weeks ago

Muse Voice Transcribe – Meta's First Real-Time Audio Perception Model

Muse Voice Transcribe is Meta's first real-time audio perception model, integrating streaming automatic speech recognition (ASR), speaker segmentation, and voice endpoint detection. The model supports...

Speech AIAI CodingBenchmark
Sep 4, 2026Read more →
3 weeks ago

Fusion Model – Review of the Intelligent Allocation Model from FUMO Lab

Fusion Model is an intelligent allocation model launched by FUMO Lab. Its core value lies in transforming traditional "static intelligent consumption" into "adaptive intelligent allocation." This mode...

AI for ScienceBenchmarkAI Tools
Sep 4, 2026Read more →
3 weeks ago

In-Depth Review of GPT-6 Astra: OpenAI's Flagship Large Model Makes a Leap into Agent Capabilities

GPT-6 Astra is OpenAI's flagship large model, officially positioned as the "most intelligent and human-intent-aligned model globally." This model has broken through the interaction boundaries of tradi...

AI AgentAI SafetyBenchmark
Sep 4, 2026Read more →
Page 2 of 9 (82 articles total)

Subscribe to AI Model Reviews

Get in-depth reviews the moment a new AI model drops. Know its capabilities and use cases.

We never share your email. Unsubscribe anytime.