AI Model Library: LLMs, Agents & Dev Tools

Curated reviews of mainstream AI models, agents, and dev tools — capabilities, use cases, and how they compare.

Expert analysisUpdated regularlyRSS Feed

Tag: Speech AI

88 article(s) found · Clear tag

2 months ago

OpenAI Presence – Enterprise AI Agent Platform for Trusted Voice and Chat Workflows

OpenAI Presence is OpenAI’s enterprise platform for deploying and operating governed AI agents in production. Announced on July 22, 2026, it targets high-volume, high-stakes workflows such as customer

AI AgentSpeech AIAI Safety
Jul 24, 2026Read more →
2 months ago

Step Edge – The Edge Model Suite from StepFusion

Step Edge is StepFusion's edge model suite, comprising four core components: Base, Audio, GUI, and Gen, designed for mobile phones and automotive terminals. Through an edge-cloud collaborative archite...

AI AgentSpeech AIModel Inference
Jul 20, 2026Read more →
2 months ago

Qwen-Audio-3.0-Realtime – Alibaba's Real-Time Speech Interaction Model

Qwen-Audio-3.0-Realtime is a new generation of real-time speech interaction dialogue model introduced by Alibaba Cloud's Tongyi team, offering Plus and Flash versions. It maintains high inference dept...

AI AgentSpeech AIVideo AI
Jul 20, 2026Read more →
2 months ago

Inkling – A Multimodal Foundation Model from Thinking Machines Lab

Inkling is an open-weight multimodal foundation model developed by Thinking Machines Lab, utilizing a Mixture-of-Experts (MoE) architecture and natively supporting unified reasoning across text, image...

AI AgentMultimodalSpeech AI
Jul 20, 2026Read more →
2 months ago

Wan-Streamer v0.2 – A Full-Modal Understanding and Generation Model from Alibaba Tongyi

Wan-Streamer v0.2 is an end-to-end full-modal understanding and generation model introduced by Alibaba Tongyi Lab, designed for real-time full-duplex interaction. This model processes real-time unders...

Speech AIVideo AILLM
Jul 20, 2026Read more →
2 months ago

MuScriptor – An Open-Source Multi-Instrument Music Transcription Model Jointly Released by Kyutai and Mirelo

MuScriptor is an open-source multi-instrument music transcription model jointly developed by Kyutai and Mirelo. It can automatically transcribe audio of music from various genres in the real world int...

Speech AIAI CodingEmbedding & RAG
Jul 20, 2026Read more →
2 months ago

Tempolor v4.7 – QwenTech's Flagship AI Music Generation Model

Tempolor v4.7 is QwenTech's flagship AI music generation model, built on a fourth-generation hierarchical progressive architecture, supporting high-quality 48kHz stereo output. This model focuses on c...

AI AgentSpeech AIAI Tools
Jul 20, 2026Read more →
2 months ago

SayIt – Open-Source AI Voice Input Tool, Automatically Converts Speech into Written Text

SayIt is an open-source AI voice input tool built with Rust, focused on the Windows desktop platform. Users simply hold down a shortcut key to speak, and their speech is converted in real-time into re...

Speech AIAI CodingOpen Source
Jul 8, 2026Read more →
2 months ago

AudioX-Turbo – A Unified and Efficient Audio Generation Framework Jointly Released by Noiz AI and Tsinghua University

AudioX-Turbo is a unified and efficient audio generation framework jointly developed by Noiz AI, the Hong Kong University of Science and Technology, and Tsinghua University. Based on a multimodal diff...

MultimodalSpeech AIVideo AI
Jul 7, 2026Read more →
2 months ago

Fun-ASR-Realtime – Alibaba Qwen's Streaming Real-Time Speech Recognition Model

Fun-ASR-Realtime is a streaming real-time speech recognition large model launched by Alibaba Qwen, designed for low-latency, high-precision speech-to-text scenarios. The model uses the WebSocket strea...

Speech AIModel InferenceBenchmark
Jul 6, 2026Read more →
Page 6 of 9 (88 articles total)

Subscribe to AI Model Reviews

Get in-depth reviews the moment a new AI model drops. Know its capabilities and use cases.

We never share your email. Unsubscribe anytime.