AI Model Library: LLMs, Agents & Dev Tools

Curated reviews of mainstream AI models, agents, and dev tools — capabilities, use cases, and how they compare.

Expert analysisUpdated regularlyRSS Feed

Tag: Benchmark

82 article(s) found · Clear tag

1 months ago

dots.tts – Xiaohongshu and Shanghai Jiao Tong University Open-Source Base Model for Text-to-Speech

dots.tts is a 20-billion-parameter fully continuous autoregressive text-to-speech base model jointly open-sourced by Xiaohongshu's dots team and the X-LANCE Lab at Shanghai Jiao Tong University. The m...

Speech AIAI CodingBenchmark
Aug 13, 2026Read more →
1 months ago

PixelRAG – Berkeley's Open-Source Vision-Native RAG Framework

PixelRAG is an open-source vision-native RAG framework developed by Berkeley's SkyLab/BAIR. It moves beyond the traditional RAG paradigm of extracting text from web pages or PDFs for retrieval, instea...

Embedding & RAGDocument AIBenchmark
Aug 13, 2026Read more →
1 months ago

PAST-Bench – Princeton's Benchmark for Performance Attribution in Personal Agents

PAST-Bench is a benchmark introduced by Princeton University's team led by Mengdi Wang, specifically designed to evaluate the recursive self-improvement capabilities of personal AI agents. By comparin...

AI AgentModel InferenceBenchmark
Aug 12, 2026Read more →
1 months ago

GPT-5.6-Cyber – OpenAI's AI Cybersecurity Model

GPT-5.6-Cyber is a specialized AI model introduced by OpenAI for the cybersecurity domain. It is built upon the GPT-5.6 Sol architecture and has been specifically enhanced and fine-tuned for advanced ...

AI SafetyBenchmark
Aug 11, 2026Read more →
1 months ago

Alpamayo 2 Super – NVIDIA's Open-Source Autonomous Driving AI Inference Model

Alpamayo 2 Super is an open-source autonomous driving AI inference model developed by NVIDIA based on the Cosmos 3 Super Reasoner. It features 360° omnidirectional environmental perception, advanced d...

Embodied AIModel InferenceBenchmark
Aug 9, 2026Read more →
1 months ago

SALMONN-2 – A General Audio Large Language Model Open-Sourced by Tsinghua University and Others

SALMONN-2 is a general audio large language model open-sourced by Tsinghua University, the Shanghai AI Laboratory, and the University of Cambridge. The model employs the SPEAR unified self-supervised ...

Speech AIAI CodingBenchmark
Aug 7, 2026Read more →
1 months ago

InstructAV2AV – An Open-Source Audio-Visual Joint Editing Model Developed by BAAI and Peking University

InstructAV2AV is an open-source audio-visual joint editing model jointly developed by the Beijing Academy of Artificial Intelligence (BAAI) and Peking University. With just a single natural language i...

Speech AIVideo AIBenchmark
Aug 6, 2026Read more →
1 months ago

SeedRealtime – ByteDance's Native Audio-Video Full-Duplex Large Model

SeedRealtime is a native audio-video full-duplex large model introduced by ByteDance's Seed team. It integrates audio, video, and text within a unified architecture, enabling real-time, full-modal int...

Speech AIVideo AIBenchmark
Aug 5, 2026Read more →
1 months ago

Shieldstral – Mistral AI's Open-Source Multimodal Content Safety Classification Model

Shieldstral is an open-source 3B parameter multimodal content safety classification model launched by Mistral AI, built upon the Ministral-3B foundation. This model redefines traditional fixed-categor...

MultimodalEmbedding & RAGBenchmark
Aug 5, 2026Read more →
1 months ago

Hy ASR 3.0 Preview – The New Generation Speech Recognition Model from Tencent Hunyuan

Hy ASR 3.0 Preview is a new generation speech recognition model launched by Tencent Hunyuan, built upon the Hy3 large language model and employing a Mixture of Experts (MoE) architecture. It integrate...

Speech AILLMBenchmark
Aug 5, 2026Read more →
Page 4 of 9 (82 articles total)

Subscribe to AI Model Reviews

Get in-depth reviews the moment a new AI model drops. Know its capabilities and use cases.

We never share your email. Unsubscribe anytime.