Curated reviews of mainstream AI models, agents, and dev tools — capabilities, use cases, and how they compare.
88 article(s) found · Clear tag
FireRedAudio is a general-purpose audio language model developed by FireRedTeam, with code, weights, and paper all publicly released. It uses a shared 9B autoregressive LLM as the inference backbone. ...
shuohao-skills is an open-source collection of AI skills specifically designed for agents like Claude Code and Codex. The tool provides a complete workflow for short film production, ranging from nove...
Gemini 3.5 Transcribe is Google's latest Speech-to-Text model, built upon the Gemini 3.5 unified multimodal architecture. It supports two core modes: real-time streaming transcription and pre-recorded...
Breeze TTS 2 is a next-generation text-to-speech (TTS) model launched by BreezeBlue. It supports zero-shot character voice design through natural language and precisely controls performance details su...
PixVerse R2 is a real-time multimodal world model launched by Aise Tech, an upgraded version of PixVerse R1. The model supports multimodal inputs such as text, images, audio, and action signals, and c...
OpenStory is an open-source AI video production platform that specializes in converting scripts into style-consistent short drama videos with a single click. The platform can automatically split scene...
FireRedTTS3 is an open-source unified speech generation and editing model developed by Xiaohongshu's FireRed team. Built upon the RedAE semantic-enhanced continuous representation and LLM-DiT architec...
Luna-TTS is a text-to-speech large model independently developed by VUI Labs, a startup company affiliated with Shanghai Jiao Tong University. This model abandons the mainstream autoregressive generat...
Mureka V9.5 is a new generation AI music generation model launched by Kunlunwanwei, built upon its self-developed MusiCoT music reasoning framework. It first constructs a global musical structure befo...
Qwen3.8-27B is a new generation of large model open-sourced by the Qwen team at Alibaba. It employs an end-to-end native audio modeling architecture with 270 billion dense parameters, supports a nativ...
Get in-depth reviews the moment a new AI model drops. Know its capabilities and use cases.
We never share your email. Unsubscribe anytime.