AI Model Library: LLMs, Agents & Dev Tools

Curated reviews of mainstream AI models, agents, and dev tools — capabilities, use cases, and how they compare.

Expert analysisUpdated regularlyRSS Feed

Tag: Speech AI

88 article(s) found · Clear tag

1 months ago

InstructAV2AV – An Open-Source Audio-Visual Joint Editing Model Developed by BAAI and Peking University

InstructAV2AV is an open-source audio-visual joint editing model jointly developed by the Beijing Academy of Artificial Intelligence (BAAI) and Peking University. With just a single natural language i...

Speech AIVideo AIBenchmark
Aug 6, 2026Read more →
1 months ago

SeedRealtime – ByteDance's Native Audio-Video Full-Duplex Large Model

SeedRealtime is a native audio-video full-duplex large model introduced by ByteDance's Seed team. It integrates audio, video, and text within a unified architecture, enabling real-time, full-modal int...

Speech AIVideo AIBenchmark
Aug 5, 2026Read more →
1 months ago

MAGI-2-preview – Sand.ai's Open-Source Multimodal Video Generation Model

MAGI-2-preview is a multimodal video generation model developed and open-sourced by Sand.ai, utilizing a Mixture of Experts (MoE) architecture with a total parameter count of 114B, activating only 6B ...

MultimodalSpeech AIVideo AI
Aug 5, 2026Read more →
1 months ago

Hy ASR 3.0 Preview – The New Generation Speech Recognition Model from Tencent Hunyuan

Hy ASR 3.0 Preview is a new generation speech recognition model launched by Tencent Hunyuan, built upon the Hy3 large language model and employing a Mixture of Experts (MoE) architecture. It integrate...

Speech AILLMBenchmark
Aug 5, 2026Read more →
1 months ago

Qwen-Audio-Agent – Alibaba's Open-Source Real-Time Voice Agent Framework

Qwen-Audio-Agent is an open-source real-time voice Agent framework developed by Alibaba's Voice AI team, built upon the Qwen-Audio-3.0-Realtime model. As a unified entry layer for real-time voice inte...

AI AgentSpeech AIAI Coding
Aug 1, 2026Read more →
1 months ago

Grok Voice Think Fast 2.0 – SpaceXAI's Voice Model

Grok Voice Think Fast 2.0 is a new generation end-to-end speech-to-speech model launched by SpaceXAI (xAI), which employs a native unified architecture. It eliminates the traditional cascaded process ...

Speech AIBenchmark
Aug 1, 2026Read more →
1 months ago

MiniMax H3 – A General-Purpose Multimodal Generation Model from MiniMax

MiniMax H3 is a general-purpose multimodal generation model officially released by MiniMax on July 31, 2026. This model breaks the boundaries between traditional tasks and modalities, achieving unifie...

MultimodalSpeech AIVideo AI
Aug 1, 2026Read more →
1 months ago

Qwen-Audio-3.0-ASR-Flash – A Speech Recognition Large Model from Alibaba Qwen

Qwen-Audio-3.0-ASR-Flash is a large speech recognition model launched by the Qwen team at Alibaba Cloud, offering three API service versions—Flash, Filetrans, and Streaming—via the Alibaba Cloud BaiLi...

Speech AIBenchmarkAI Tools
Aug 1, 2026Read more →
2 months ago

FLUX 3 – A Multimodal Foundation Model Launched by Black Forest Labs

FLUX 3 is a multimodal foundation model launched by Black Forest Labs, which for the first time jointly learns images, videos, and audio under a unified architecture. Based on Self-Flow technology, th...

MultimodalSpeech AIImage Generation
Jul 28, 2026Read more →
2 months ago

MAI-Voice-2-Flash – Microsoft's High-Speed Text-to-Speech Model

MAI-Voice-2-Flash is a high-speed text-to-speech (TTS) model introduced by Microsoft's AI team for high-concurrency, low-latency scenarios. While maintaining natural tone and high audio quality, its i...

Speech AIVideo AIModel Inference
Jul 28, 2026Read more →
Page 5 of 9 (88 articles total)

Subscribe to AI Model Reviews

Get in-depth reviews the moment a new AI model drops. Know its capabilities and use cases.

We never share your email. Unsubscribe anytime.