AI Model Library: LLMs, Agents & Dev Tools

Curated reviews of mainstream AI models, agents, and dev tools — capabilities, use cases, and how they compare.

Expert analysisUpdated regularlyRSS Feed

Tag: Speech AI

88 article(s) found · Clear tag

3 months ago

Stable Audio 3 – Stability AI's Open-Source Audio Generation Model Series

Stable Audio 3 is Stability AI's next-generation open-source audio generation model series, built on flow-matching latent-space diffusion architecture. It supports text-to-music and sound effects, aud...

Speech AIOpen SourceEdge Deployment
Jun 21, 2026Read more →
3 months ago

SpaceMind – iFLYTEK’s Agentic Smart-Space Architecture

SpaceMind is iFLYTEK’s agentic architecture for intelligent spaces. It upgrades traditional smart homes into L2.5 proactive agents with perception, understanding, memory, decision-making, and executio...

AI AgentMultimodalSpeech AI
Jun 21, 2026Read more →
3 months ago

MiMo Code – Xiaomi’s Open-Source Terminal AI Coding Assistant

MiMo Code is Xiaomi’s MIT-licensed terminal AI coding assistant, built on the OpenCode framework. It bundles a limited-time free MiMo-V2.5 model with a persistent memory system, Compose full-flow mode...

AI AgentMultimodalSpeech AI
Jun 21, 2026Read more →
3 months ago

Mega-ASR – NTU, NUS & Shanghai AI Lab Open-Source Speech Recognition Model

Mega-ASR is an open all-scenario robust ASR foundation from Nanyang Technological University (NTU), National University of Singapore (NUS), and Shanghai AI Lab. Built on Qwen3-ASR 1.7B, it uses compos...

Speech AIOpen Source
Jun 21, 2026Read more →
3 months ago

MAI-Voice-2 – Microsoft's Next-Generation Text-to-Speech Model

MAI-Voice-2 is Microsoft's next-generation text-to-speech (TTS) model launched in June 2026, built on a proprietary speech foundation architecture. It delivers a generational leap in fidelity, languag...

Speech AIAI CodingAI Safety
Jun 21, 2026Read more →
3 months ago

Magenta RealTime 2 – Google's Open Real-Time Music Generation Model

Magenta RealTime 2 (MRT2) is Google Magenta's second-generation open local real-time music generation model. Built on frame-by-frame autoregressive architecture, it compresses audio generation latency...

AI AgentMultimodalSpeech AI
Jun 21, 2026Read more →
3 months ago

LongCat-Video-Avatar 1.5 – Meituan's Open-Source Digital Human Video Generation Model

LongCat-Video-Avatar 1.5 is Meituan LongCat's open audio-driven digital human video framework, built on the 13.6B-parameter LongCat-Video base. Upgrading the audio encoder from Wav2Vec2 to Whisper-Lar...

MultimodalSpeech AIVideo AI
Jun 21, 2026Read more →
3 months ago

Khala – Central Conservatory of Music & Tsinghua Open-Source AI Music Model

Khala is an open-source AI music foundation model jointly developed by the Central Conservatory of Music and Tsinghua University, focused on high-fidelity song generation. It uses a 64-layer deep acou...

Speech AIAI CodingAI Safety
Jun 21, 2026Read more →
3 months ago

JoyAI-Echo – JD's Open-Source Long-Form Audio-Video Generation Framework

JoyAI-Echo is an open-source long-form audio-video generation framework from JD Future Academy, designed for minute-scale multi-shot story generation. Through four technical innovations—a paired cross...

AI AgentMultimodalSpeech AI
Jun 21, 2026Read more →
3 months ago

Hojo-ASR-V1 – Hojo’s Open-Source Automatic Speech Recognition Model

Hojo-ASR-V1 is an open-source automatic speech recognition model from startup Hojo. Its four-stage hybrid stack combines Whisper feature extraction, Qwen3-Omni audio encoding, Conformer adaptation, an...

AI AgentMultimodalSpeech AI
Jun 21, 2026Read more →
Page 8 of 9 (88 articles total)

Subscribe to AI Model Reviews

Get in-depth reviews the moment a new AI model drops. Know its capabilities and use cases.

We never share your email. Unsubscribe anytime.