Curated reviews of mainstream AI models, agents, and dev tools — capabilities, use cases, and how they compare.
82 article(s) found · Clear tag
MiniMax M3 is MiniMax's next-gen model with MSA sparse attention and MoE—leading programming, agent, and long-context workloads. 196B total parameters, ~11B active per forward pass, up to 1M token con...
MaineCoon is the world’s first real-time audio-visual autoregressive world model optimized for social interaction scenarios. With 22 billion parameters, it delivers 47.5 FPS streaming generation on a ...
LongCat-Video-Avatar 1.5 is Meituan LongCat's open audio-driven digital human video framework, built on the 13.6B-parameter LongCat-Video base. Upgrading the audio encoder from Wav2Vec2 to Whisper-Lar...
Keye-VL-2.0-30B-A3B is Kuaishou's fully open multimodal large language model—30B total parameters (MoE, ~3B active)—and the first to bring DeepSeek Sparse Attention (DSA) into multimodal understanding...
html-video is an open-source HTML CapCut from Open Design (nexu.io), built on the hyperframes framework. Agents write HTML to produce product promos, explainers, and similar videos at world-class qual...
Grok Imagine Video 1.5 is xAI's next-generation image-to-video model built on the in-house Aurora autoregressive engine. Upload a single static image and a natural-language prompt to generate a short ...
Gamma-World is a multi-agent world model from NVIDIA Research, designed so multiple agents interact equally in one simulated world with global consistency. It uses Simplex Rotation Encoding for identi...
Dubbing v2 is ElevenLabs' latest offering in AI dubbing—an end-to-end multilingual dubbing platform that integrates speech recognition, neural machine translation, voice cloning, and synthesis. It can...
Cosmos 3 is NVIDIA's first fully open-source, full-modality physical AI foundation model. Built on an innovative hybrid Transformer architecture, it natively fuses visual reasoning, world generation, ...
ControlFoley is an open-source controllable video sound effect generation model from Xiaomi Research, designed to solve the long-standing controllability challenge in video-to-audio (V2A). The model u...
Get in-depth reviews the moment a new AI model drops. Know its capabilities and use cases.
We never share your email. Unsubscribe anytime.