AI Model Library: LLMs, Agents & Dev Tools

Curated reviews of mainstream AI models, agents, and dev tools — capabilities, use cases, and how they compare.

Expert analysisUpdated regularlyRSS Feed

Tag: Multimodal

111 article(s) found · Clear tag

1 months ago

SenseNova U1.5 Lite – A Lightweight Native Unified Multimodal Large Model Open-Sourced by SenseTime

SenseNova U1.5 Lite is a lightweight native unified multimodal large model open-sourced by SenseTime, specifically designed for real-world visual creation workflows. This model achieves a native multi...

MultimodalDocument AITool Calling
Aug 21, 2026Read more →
1 months ago

In-depth Review of DeepSeek-V4-Flash-Vision-Exp: A Multimodal Vision Agent with Zero Text Capability Loss

DeepSeek-V4-Flash-Vision-Exp is an experimental multimodal vision understanding model introduced by DeepSeek. This model adds a visual input channel to the full text capabilities of DeepSeek-V4-Flash,...

AI AgentMultimodal
Aug 21, 2026Read more →
1 months ago

Mage – Microsoft's Open-Source 4B Parameter Multimodal Model Family

Mage is a family of open-source 4B parameter multimodal models developed by Microsoft, consisting of the streaming video/image understanding model Mage-VL and the image generation and editing model Ma...

MultimodalVideo AIAI Coding
Aug 19, 2026Read more →
1 months ago

dots3-note preview – In-Depth Review of Xiaohongshu's Open-Source Multimodal MoE Model

dots3-note preview is an open-source multimodal MoE model developed by Xiaohongshu's dots model lab. As the first version of the dots3 series, it shares the same technical lineage as the IMO 2026 full...

AI AgentMultimodalSpeech AI
Aug 14, 2026Read more →
1 months ago

HOMIE – An Open-Source Digital Human Video Generation Framework from The Hong Kong University of Science and Technology

HOMIE is an open-source digital human video generation framework developed by the Department of Computer Science and Engineering at The Hong Kong University of Science and Technology. It is built upon...

MultimodalVideo AIDocument AI
Aug 12, 2026Read more →
1 months ago

Wan3.0 – Alibaba's WanXiang Latest Video Generation Large Model

Wan3.0 is the latest video generation large model launched by Alibaba Cloud's WanXiang team, achieving comprehensive upgrades in generation duration, multimodal input, consistency maintenance, and rea...

MultimodalVideo AIDocument AI
Aug 7, 2026Read more →
1 months ago

Hunyuan3D-Buffalo 1.0 – Tencent Hunyuan's Unified 3D Multimodal Framework

Hunyuan3D-Buffalo 1.0 is a unified 3D multimodal framework introduced by the Tencent Hunyuan team. It integrates multiple tasks—such as 3D question answering, spatial localization, text-to-3D generati...

MultimodalTool CallingEdge Deployment
Aug 6, 2026Read more →
1 months ago

MAGI-2-preview – Sand.ai's Open-Source Multimodal Video Generation Model

MAGI-2-preview is a multimodal video generation model developed and open-sourced by Sand.ai, utilizing a Mixture of Experts (MoE) architecture with a total parameter count of 114B, activating only 6B ...

MultimodalSpeech AIVideo AI
Aug 5, 2026Read more →
1 months ago

Shieldstral – Mistral AI's Open-Source Multimodal Content Safety Classification Model

Shieldstral is an open-source 3B parameter multimodal content safety classification model launched by Mistral AI, built upon the Ministral-3B foundation. This model redefines traditional fixed-categor...

MultimodalEmbedding & RAGBenchmark
Aug 5, 2026Read more →
1 months ago

SenseNova U1.5-Lite-Preview – A Lightweight Multimodal Model Open-Sourced by SenseTime

SenseNova U1.5-Lite-Preview is the preview version of a lightweight, native unified multimodal model open-sourced by SenseTime, developed iteratively based on the NEO-Unify architecture. With only 8B-...

MultimodalDocument AIEdge Deployment
Aug 4, 2026Read more →
Page 3 of 12 (111 articles total)

Subscribe to AI Model Reviews

Get in-depth reviews the moment a new AI model drops. Know its capabilities and use cases.

We never share your email. Unsubscribe anytime.