Curated reviews of mainstream AI models, agents, and dev tools — capabilities, use cases, and how they compare.
88 article(s) found · Clear tag
HiDream-O1-Video-1.0 (shortened as HD-V1) is a native full-modal video generation model launched by HiDream.ai. It supports multi-modal inputs such as text, images, and videos, and can directly genera...
Xiaomi-CocktailASR-1 is Xiaomi's open-source Target Speaker ASR (TS-ASR) large model, designed using an end-to-end LLM architecture. It uses a reference speech as a speaker embedding prompt to accurat...
EchoWM is an interactive audio-visual world model open-sourced by the JD.com Exploration Research Institute. It continues the native audio-visual generation capabilities of JoyAI-Echo, allowing users ...
Suno v6 is a new generation of model series introduced by the AI music generation platform Suno. It includes the flagship model v6 (targeted at Pro/Premier paying users), the experimental model v6-wil...
AuK is an open-source foundation model for speech generation and editing developed by the Tencent HunYuan team, featuring 1.5 billion parameters and utilizing a flow-matching diffusion architecture in...
MiniCPM5-2B is an end-side language foundation model open-sourced jointly by RWKV, the OpenBMB community, and Tsinghua University. With approximately 2.5 billion parameters, it supports a context leng...
MiniMax H3 Max is a real-time video generation model introduced by MiniMax, based on the open-source H3 model, with post-training and inference optimization. This model supports two input methods: tex...
FireRedAudio is a general-purpose audio language model open-sourced by the FireRed team at Xiaohongshu in August 2026. It is built upon the Qwen3.5 autoregressive large language model with 9B paramete...
MAI-Transcribe-2 is an AI speech-to-text (STT/ASR) model launched by Microsoft, achieving comprehensive leadership in three core dimensions: accuracy, speed, and cost. The model supports 60 languages,...
Muse Voice Transcribe is Meta's first real-time audio perception model, integrating streaming automatic speech recognition (ASR), speaker segmentation, and voice endpoint detection. The model supports...
Get in-depth reviews the moment a new AI model drops. Know its capabilities and use cases.
We never share your email. Unsubscribe anytime.