AI Model Library: LLMs, Agents & Dev Tools

Curated reviews of mainstream AI models, agents, and dev tools — capabilities, use cases, and how they compare.

Expert analysisUpdated regularlyRSS Feed

Tag: Multimodal

111 article(s) found · Clear tag

1 months ago

PerceptionBench – Moonshot AI's Open-Source Visual Perception Diagnostic Benchmark

PerceptionBench is an open-source visual perception diagnostic benchmark launched by Moonshot AI, focusing on fine-grained evaluation of atomic-level visual perception capabilities. Starting from the ...

MultimodalLLMBenchmark
Aug 1, 2026Read more →
1 months ago

MiniMax H3 – A General-Purpose Multimodal Generation Model from MiniMax

MiniMax H3 is a general-purpose multimodal generation model officially released by MiniMax on July 31, 2026. This model breaks the boundaries between traditional tasks and modalities, achieving unifie...

MultimodalSpeech AIVideo AI
Aug 1, 2026Read more →
2 months ago

FLUX 3 – A Multimodal Foundation Model Launched by Black Forest Labs

FLUX 3 is a multimodal foundation model launched by Black Forest Labs, which for the first time jointly learns images, videos, and audio under a unified architecture. Based on Self-Flow technology, th...

MultimodalSpeech AIImage Generation
Jul 28, 2026Read more →
2 months ago

BigMac – Xiaohongshu's Open-Source Framework for Pipeline-Parallel Training of Multimodal Large Models

BigMac is an open-source pipeline-parallel training framework for multimodal large models developed by the dots infra team at Xiaohongshu. It aims to address the fundamental contradiction between comp...

MultimodalAI CodingLLM
Jul 28, 2026Read more →
2 months ago

Cosmos 3 Edge – NVIDIA's 4B Parameter Open-Source World Model

Cosmos 3 Edge is a 4B parameter open-source world model introduced by NVIDIA, specifically designed for real-time inference in robotics and edge AI. Based on the Nemotron architecture, the model integ...

MultimodalImage GenerationEmbodied AI
Jul 28, 2026Read more →
2 months ago

SciReasoner – A Multimodal Foundation Model for Scientific Reasoning Developed by the Shanghai AI Lab and Others

SciReasoner is a new-generation multimodal foundation model for scientific reasoning, jointly developed by the Shanghai Artificial Intelligence Laboratory with teams from the Chinese University of Hon...

MultimodalAI for ScienceLLM
Jul 24, 2026Read more →
2 months ago

Mage-Flow – Microsoft's Open-Source Lightweight Multimodal Image Generation Model Series

Mage-Flow is a lightweight multimodal image generation model series open-sourced by Microsoft, based on a flow-matching architecture and featuring only 4B parameters. The model includes four variants:...

MultimodalImage GenerationEdge Deployment
Jul 24, 2026Read more →
2 months ago

SenseNova-Vision – SenseTime's Open-Source Unified Vision Large Model for Understanding and Generation

SenseNova-Vision is an open-source unified vision large model developed by SenseTime, capable of both understanding and generating visual content. Its core innovation lies in natively integrating clas...

Multimodal
Jul 20, 2026Read more →
2 months ago

Inkling – A Multimodal Foundation Model from Thinking Machines Lab

Inkling is an open-weight multimodal foundation model developed by Thinking Machines Lab, utilizing a Mixture-of-Experts (MoE) architecture and natively supporting unified reasoning across text, image...

AI AgentMultimodalSpeech AI
Jul 20, 2026Read more →
2 months ago

Muse Image – AI Image Generation Model by Meta Super Intelligence Lab

Muse Image is Meta Super Intelligence Lab's first self-developed AI image generation model, adopting an Agent architecture that collaborates with Muse Spark for multi-step planning and self-refinement...

AI AgentMultimodalImage Generation
Jul 8, 2026Read more →
Page 4 of 12 (111 articles total)

Subscribe to AI Model Reviews

Get in-depth reviews the moment a new AI model drops. Know its capabilities and use cases.

We never share your email. Unsubscribe anytime.