AI Model Library: LLMs, Agents & Dev Tools

Curated reviews of mainstream AI models, agents, and dev tools — capabilities, use cases, and how they compare.

Expert analysisUpdated regularlyRSS Feed

Tag: Video AI

82 article(s) found · Clear tag

3 months ago

MiniMax M3 – MiniMax's Next-Generation AI Model

MiniMax M3 is MiniMax's next-gen model with MSA sparse attention and MoE—leading programming, agent, and long-context workloads. 196B total parameters, ~11B active per forward pass, up to 1M token con...

AI AgentMultimodalVideo AI
Jun 21, 2026Read more →
3 months ago

MaineCoon – Real-Time Audio-Visual World Model Built for Social Interaction

MaineCoon is the world’s first real-time audio-visual autoregressive world model optimized for social interaction scenarios. With 22 billion parameters, it delivers 47.5 FPS streaming generation on a ...

MultimodalVideo AIAI Safety
Jun 21, 2026Read more →
3 months ago

LongCat-Video-Avatar 1.5 – Meituan's Open-Source Digital Human Video Generation Model

LongCat-Video-Avatar 1.5 is Meituan LongCat's open audio-driven digital human video framework, built on the 13.6B-parameter LongCat-Video base. Upgrading the audio encoder from Wav2Vec2 to Whisper-Lar...

MultimodalSpeech AIVideo AI
Jun 21, 2026Read more →
3 months ago

Keye-VL-2.0-30B-A3B – Kuaishou's Open Multimodal Large Model

Keye-VL-2.0-30B-A3B is Kuaishou's fully open multimodal large language model—30B total parameters (MoE, ~3B active)—and the first to bring DeepSeek Sparse Attention (DSA) into multimodal understanding...

AI AgentMultimodalVideo AI
Jun 21, 2026Read more →
3 months ago

html-video – Open Design Team's Open-Source HTML CapCut

html-video is an open-source HTML CapCut from Open Design (nexu.io), built on the hyperframes framework. Agents write HTML to produce product promos, explainers, and similar videos at world-class qual...

AI AgentVideo AIAI Coding
Jun 21, 2026Read more →
3 months ago

Grok Imagine Video 1.5 – xAI's Image-to-Video Model

Grok Imagine Video 1.5 is xAI's next-generation image-to-video model built on the in-house Aurora autoregressive engine. Upload a single static image and a natural-language prompt to generate a short ...

MultimodalSpeech AIVideo AI
Jun 21, 2026Read more →
3 months ago

Gamma-World – NVIDIA's Multi-Agent World Model

Gamma-World is a multi-agent world model from NVIDIA Research, designed so multiple agents interact equally in one simulated world with global consistency. It uses Simplex Rotation Encoding for identi...

AI AgentVideo AIAI Coding
Jun 21, 2026Read more →
3 months ago

Dubbing v2 – ElevenLabs' AI Dubbing Model

Dubbing v2 is ElevenLabs' latest offering in AI dubbing—an end-to-end multilingual dubbing platform that integrates speech recognition, neural machine translation, voice cloning, and synthesis. It can...

MultimodalSpeech AIVideo AI
Jun 21, 2026Read more →
3 months ago

Cosmos 3 – NVIDIA's Open-Source Full-Modality Physical AI Foundation Model

Cosmos 3 is NVIDIA's first fully open-source, full-modality physical AI foundation model. Built on an innovative hybrid Transformer architecture, it natively fuses visual reasoning, world generation, ...

MultimodalVideo AIEmbodied AI
Jun 21, 2026Read more →
3 months ago

ControlFoley – Xiaomi's Open-Source Controllable Video Sound Effect Generation Model

ControlFoley is an open-source controllable video sound effect generation model from Xiaomi Research, designed to solve the long-standing controllability challenge in video-to-audio (V2A). The model u...

MultimodalSpeech AIVideo AI
Jun 21, 2026Read more →
Page 8 of 9 (82 articles total)

Subscribe to AI Model Reviews

Get in-depth reviews the moment a new AI model drops. Know its capabilities and use cases.

We never share your email. Unsubscribe anytime.