AI Model Library: LLMs, Agents & Dev Tools

Curated reviews of mainstream AI models, agents, and dev tools — capabilities, use cases, and how they compare.

Expert analysisUpdated regularlyRSS Feed

Tag: Speech AI

88 article(s) found · Clear tag

3 weeks ago

FireRedAudio – A General Audio Language Model from FireRedTeam

FireRedAudio is a general-purpose audio language model developed by FireRedTeam, with code, weights, and paper all publicly released. It uses a shared 9B autoregressive LLM as the inference backbone. ...

Speech AIAI CodingEmbedding & RAG
Sep 1, 2026Read more →
4 weeks ago

shuohao-skills – Open-Source AI Short Film Production Skill Collection, Featuring a Complete Workflow

shuohao-skills is an open-source collection of AI skills specifically designed for agents like Claude Code and Codex. The tool provides a complete workflow for short film production, ranging from nove...

AI AgentSpeech AIVideo AI
Aug 28, 2026Read more →
1 months ago

In-Depth Review of Gemini 3.5 Transcribe: Technical Analysis and Application Outlook of Google's Next-Generation Speech-to-Text Model

Gemini 3.5 Transcribe is Google's latest Speech-to-Text model, built upon the Gemini 3.5 unified multimodal architecture. It supports two core modes: real-time streaming transcription and pre-recorded...

MultimodalSpeech AIAI Tools
Aug 27, 2026Read more →
1 months ago

In-Depth Review of Breeze TTS 2 – BreezeBlue's Leading Text-to-Speech Model

Breeze TTS 2 is a next-generation text-to-speech (TTS) model launched by BreezeBlue. It supports zero-shot character voice design through natural language and precisely controls performance details su...

AI AgentSpeech AIVideo AI
Aug 26, 2026Read more →
1 months ago

PixVerse R2 – A Real-Time Multimodal World Model from Aise Tech

PixVerse R2 is a real-time multimodal world model launched by Aise Tech, an upgraded version of PixVerse R1. The model supports multimodal inputs such as text, images, audio, and action signals, and c...

MultimodalSpeech AIVideo AI
Aug 26, 2026Read more →
1 months ago

OpenStory: Open Source AI Video Production Platform, Generate Style-Consistent Short Drama Videos with One Click

OpenStory is an open-source AI video production platform that specializes in converting scripts into style-consistent short drama videos with a single click. The platform can automatically split scene...

Speech AIVideo AIDocument AI
Aug 25, 2026Read more →
1 months ago

FireRedTTS3 – Xiaohongshu's Open-Source Unified Speech Generation and Editing Model

FireRedTTS3 is an open-source unified speech generation and editing model developed by Xiaohongshu's FireRed team. Built upon the RedAE semantic-enhanced continuous representation and LLM-DiT architec...

Speech AIAI CodingAI Safety
Aug 25, 2026Read more →
1 months ago

Luna-TTS – A Text-to-Speech Large Model from VUI Labs

Luna-TTS is a text-to-speech large model independently developed by VUI Labs, a startup company affiliated with Shanghai Jiao Tong University. This model abandons the mainstream autoregressive generat...

Speech AIImage GenerationAI Coding
Aug 18, 2026Read more →
1 months ago

Mureka V9.5 – A New Generation AI Music Generation Model from Kunlunwanwei

Mureka V9.5 is a new generation AI music generation model launched by Kunlunwanwei, built upon its self-developed MusiCoT music reasoning framework. It first constructs a global musical structure befo...

Speech AIReasoning ModelBenchmark
Aug 18, 2026Read more →
1 months ago

Qwen3.8-27B – A New Generation of Open-Source AI Large Model from the Qwen Team at Alibaba

Qwen3.8-27B is a new generation of large model open-sourced by the Qwen team at Alibaba. It employs an end-to-end native audio modeling architecture with 270 billion dense parameters, supports a nativ...

AI AgentSpeech AIModel Inference
Aug 15, 2026Read more →
Page 3 of 9 (88 articles total)

Subscribe to AI Model Reviews

Get in-depth reviews the moment a new AI model drops. Know its capabilities and use cases.

We never share your email. Unsubscribe anytime.