Curated reviews of mainstream AI models, agents, and dev tools — capabilities, use cases, and how they compare.
111 article(s) found · Clear tag
SenseNova U1.5 Lite is a lightweight native unified multimodal large model open-sourced by SenseTime, specifically designed for real-world visual creation workflows. This model achieves a native multi...
DeepSeek-V4-Flash-Vision-Exp is an experimental multimodal vision understanding model introduced by DeepSeek. This model adds a visual input channel to the full text capabilities of DeepSeek-V4-Flash,...
Mage is a family of open-source 4B parameter multimodal models developed by Microsoft, consisting of the streaming video/image understanding model Mage-VL and the image generation and editing model Ma...
dots3-note preview is an open-source multimodal MoE model developed by Xiaohongshu's dots model lab. As the first version of the dots3 series, it shares the same technical lineage as the IMO 2026 full...
HOMIE is an open-source digital human video generation framework developed by the Department of Computer Science and Engineering at The Hong Kong University of Science and Technology. It is built upon...
Wan3.0 is the latest video generation large model launched by Alibaba Cloud's WanXiang team, achieving comprehensive upgrades in generation duration, multimodal input, consistency maintenance, and rea...
Hunyuan3D-Buffalo 1.0 is a unified 3D multimodal framework introduced by the Tencent Hunyuan team. It integrates multiple tasks—such as 3D question answering, spatial localization, text-to-3D generati...
MAGI-2-preview is a multimodal video generation model developed and open-sourced by Sand.ai, utilizing a Mixture of Experts (MoE) architecture with a total parameter count of 114B, activating only 6B ...
Shieldstral is an open-source 3B parameter multimodal content safety classification model launched by Mistral AI, built upon the Ministral-3B foundation. This model redefines traditional fixed-categor...
SenseNova U1.5-Lite-Preview is the preview version of a lightweight, native unified multimodal model open-sourced by SenseTime, developed iteratively based on the NEO-Unify architecture. With only 8B-...
Get in-depth reviews the moment a new AI model drops. Know its capabilities and use cases.
We never share your email. Unsubscribe anytime.