AI Model Library: LLMs, Agents & Dev Tools

Curated reviews of mainstream AI models, agents, and dev tools — capabilities, use cases, and how they compare.

Expert analysisUpdated regularlyRSS Feed

Tag: Multimodal

112 article(s) found · Clear tag

3 months ago

PaddleOCR-VL-1.6 – Baidu's Document Parsing Vision-Language Model

PaddleOCR-VL-1.6 is a vision-language model (VLM) for document parsing from Baidu's PaddlePaddle team. With only 0.9B parameters, it achieves a new SOTA of 96.33% on the OmniDocBench v1.6 benchmark, s...

MultimodalDocument AILLM
Jun 21, 2026Read more →
3 months ago

Odysseus – Open-Source Local Self-Hosted AI Workspace

Odysseus is an open-source, locally self-hostable AI workspace that integrates more than ten capabilities—chat, agents, deep research, notes, tasks, calendar, and email—into a private AI hub. Through ...

AI AgentMultimodalEmbedding & RAG
Jun 21, 2026Read more →
3 months ago

MiniMax M3 – MiniMax's Next-Generation AI Model

MiniMax M3 is MiniMax's next-gen model with MSA sparse attention and MoE—leading programming, agent, and long-context workloads. 196B total parameters, ~11B active per forward pass, up to 1M token con...

AI AgentMultimodalVideo AI
Jun 21, 2026Read more →
3 months ago

MiMo Code – Xiaomi’s Open-Source Terminal AI Coding Assistant

MiMo Code is Xiaomi’s MIT-licensed terminal AI coding assistant, built on the OpenCode framework. It bundles a limited-time free MiMo-V2.5 model with a persistent memory system, Compose full-flow mode...

AI AgentMultimodalSpeech AI
Jun 21, 2026Read more →
3 months ago

MaineCoon – Real-Time Audio-Visual World Model Built for Social Interaction

MaineCoon is the world’s first real-time audio-visual autoregressive world model optimized for social interaction scenarios. With 22 billion parameters, it delivers 47.5 FPS streaming generation on a ...

MultimodalVideo AIAI Safety
Jun 21, 2026Read more →
3 months ago

MAI-Image-2.5 – Microsoft's Flagship Text-to-Image Model

MAI-Image-2.5 is a flagship text-to-image model from Microsoft Research and the strongest release in the MAI-Image family. On the Arena text-to-image leaderboard it climbed to #3 with 1,254 points—72 ...

MultimodalImage GenerationBenchmark
Jun 21, 2026Read more →
3 months ago

Magenta RealTime 2 – Google's Open Real-Time Music Generation Model

Magenta RealTime 2 (MRT2) is Google Magenta's second-generation open local real-time music generation model. Built on frame-by-frame autoregressive architecture, it compresses audio generation latency...

AI AgentMultimodalSpeech AI
Jun 21, 2026Read more →
3 months ago

LongCat-Video-Avatar 1.5 – Meituan's Open-Source Digital Human Video Generation Model

LongCat-Video-Avatar 1.5 is Meituan LongCat's open audio-driven digital human video framework, built on the 13.6B-parameter LongCat-Video base. Upgrading the audio encoder from Wav2Vec2 to Whisper-Lar...

MultimodalSpeech AIVideo AI
Jun 21, 2026Read more →
3 months ago

LOGOS – Alibaba's First Open Unified Scientific Foundation Model

LOGOS (Language Of Generative Objects in Science) is the first open unified scientific grammar multi-domain generative foundation model from Alibaba ATH-Token Foundry and Renmin University of China Ga...

MultimodalAI for ScienceOpen Source
Jun 21, 2026Read more →
3 months ago

Kimi K2.7 Code – Moonshot AI’s Open-Source Coding Model

Kimi K2.7 Code is Moonshot AI’s open-source next-generation large language model built specifically for programming. It is deeply optimized on top of the K2 series architecture. The model delivers a m...

AI AgentMultimodalAI Coding
Jun 21, 2026Read more →
Page 9 of 12 (112 articles total)

Subscribe to AI Model Reviews

Get in-depth reviews the moment a new AI model drops. Know its capabilities and use cases.

We never share your email. Unsubscribe anytime.