AI Model Library: LLMs, Agents & Dev Tools

Curated reviews of mainstream AI models, agents, and dev tools — capabilities, use cases, and how they compare.

Expert analysisUpdated regularlyRSS Feed

Tag: Speech AI

88 article(s) found · Clear tag

1 weeks ago

HiDream-O1-Video-1.0: In-Depth Evaluation of a Native Full-Modal Video Generation Model

HiDream-O1-Video-1.0 (shortened as HD-V1) is a native full-modal video generation model launched by HiDream.ai. It supports multi-modal inputs such as text, images, and videos, and can directly genera...

Speech AIVideo AIBenchmark
Sep 15, 2026Read more →
1 weeks ago

Xiaomi-CocktailASR-1: In-Depth Evaluation of a Target Speaker ASR Model Based on an End-to-End LLM Architecture

Xiaomi-CocktailASR-1 is Xiaomi's open-source Target Speaker ASR (TS-ASR) large model, designed using an end-to-end LLM architecture. It uses a reference speech as a speaker embedding prompt to accurat...

Speech AIReasoning ModelEmbedding & RAG
Sep 14, 2026Read more →
2 weeks ago

EchoWM – Interactive Audio-Visual World Model Open-Sourced by JD.com

EchoWM is an interactive audio-visual world model open-sourced by the JD.com Exploration Research Institute. It continues the native audio-visual generation capabilities of JoyAI-Echo, allowing users ...

Speech AIEmbedding & RAGBenchmark
Sep 11, 2026Read more →
2 weeks ago

Suno v6 – The Next-Generation Hierarchical Model Series for AI Music Generation

Suno v6 is a new generation of model series introduced by the AI music generation platform Suno. It includes the flagship model v6 (targeted at Pro/Premier paying users), the experimental model v6-wil...

Speech AIAI Tools
Sep 10, 2026Read more →
2 weeks ago

AuK – Tencent HunYuan's Open-Source Foundation Model for Speech Generation and Editing

AuK is an open-source foundation model for speech generation and editing developed by the Tencent HunYuan team, featuring 1.5 billion parameters and utilizing a flow-matching diffusion architecture in...

Speech AIImage GenerationLLM
Sep 10, 2026Read more →
2 weeks ago

MiniCPM5-2B – End-to-End Native Audio Modeling Architecture Open-Sourced by RWKV and Tsinghua University

MiniCPM5-2B is an end-side language foundation model open-sourced jointly by RWKV, the OpenBMB community, and Tsinghua University. With approximately 2.5 billion parameters, it supports a context leng...

AI AgentSpeech AIAI Coding
Sep 8, 2026Read more →
3 weeks ago

MiniMax H3 Max: Live-level Speed and Ecosystem Evolution in Real-time Video Generation

MiniMax H3 Max is a real-time video generation model introduced by MiniMax, based on the open-source H3 model, with post-training and inference optimization. This model supports two input methods: tex...

Speech AIVideo AIModel Inference
Sep 4, 2026Read more →
3 weeks ago

FireRedAudio – In-Depth Review of Xiaohongshu FireRed's General-Purpose Audio Language Model

FireRedAudio is a general-purpose audio language model open-sourced by the FireRed team at Xiaohongshu in August 2026. It is built upon the Qwen3.5 autoregressive large language model with 9B paramete...

Speech AIAI CodingAI Safety
Sep 4, 2026Read more →
3 weeks ago

MAI-Transcribe-2 – Microsoft's Most Advanced AI Speech-to-Text Model

MAI-Transcribe-2 is an AI speech-to-text (STT/ASR) model launched by Microsoft, achieving comprehensive leadership in three core dimensions: accuracy, speed, and cost. The model supports 60 languages,...

Speech AIEmbedding & RAGAI Tools
Sep 4, 2026Read more →
3 weeks ago

Muse Voice Transcribe – Meta's First Real-Time Audio Perception Model

Muse Voice Transcribe is Meta's first real-time audio perception model, integrating streaming automatic speech recognition (ASR), speaker segmentation, and voice endpoint detection. The model supports...

Speech AIAI CodingBenchmark
Sep 4, 2026Read more →
Page 2 of 9 (88 articles total)

Subscribe to AI Model Reviews

Get in-depth reviews the moment a new AI model drops. Know its capabilities and use cases.

We never share your email. Unsubscribe anytime.