AI Model Library: LLMs, Agents & Dev Tools

Curated reviews of mainstream AI models, agents, and dev tools — capabilities, use cases, and how they compare.

Expert analysisUpdated regularlyRSS Feed

Tag: Multimodal

112 article(s) found · Clear tag

3 months ago

StepAudio 2.5 Realtime – StepFun's Real-Time Speech LLM

StepAudio 2.5 Realtime from StepFun (阶跃星辰) is an end-to-end real-time speech LLM aimed at human-grade voice conversation. It goes beyond mechanical TTS with deep content understanding and nuanced emot...

MultimodalSpeech AI
Jun 21, 2026Read more →
3 months ago

Higgs Avatar v1 – Real-Time AI Digital Human for Voice Agents

Higgs Avatar v1 from BosonAI (founded by Mu Li) is a real-time AI digital human model for voice agents. From a single static photo it produces lip-synced faces with expressions and head motion for liv...

AI AgentMultimodalSpeech AI
Jun 21, 2026Read more →
3 months ago

Qwen3.7 Preview – Alibaba Tongyi's Next-Generation Flagship LLM Preview

Qwen3.7 Preview is Alibaba Tongyi Qwen's next-generation flagship preview, with Qwen3.7-Max-Preview and Qwen3.7-Plus-Preview targeting extreme complex reasoning and balanced experience respectively. I...

AI AgentMultimodalAI Coding
Jun 21, 2026Read more →
3 months ago

Chronicles-OCR – Cross-Temporal Visual Perception Benchmark for Chinese Script Evolution

Chronicles-OCR is the first industry benchmark to cover the full evolutionary trajectory of Chinese script "seven-style transformation" (七体之变)—jointly released by Tencent Hunyuan, the Institute of Inf...

MultimodalDocument AIAI for Science
Jun 21, 2026Read more →
3 months ago

Qwen3.5-LiveTranslate – Alibaba Tongyi's Real-Time Simultaneous Interpretation Model

Qwen3.5-LiveTranslate is Alibaba Tongyi's next-generation real-time simultaneous interpretation model. It breaks the latency ceiling of traditional cascaded translation pipelines. Through an innovativ...

MultimodalSpeech AIModel Inference
Jun 21, 2026Read more →
3 months ago

Gemini Omni Flash – Google's Multimodal Video Generation Model

Gemini Omni Flash is Google's unified multimodal world generation model unveiled at I/O, positioned to break traditional generative model modality barriers and enable any-input to any-output full-pipe...

MultimodalSpeech AIImage Generation
Jun 21, 2026Read more →
3 months ago

Lance – ByteDance's Lightweight Native Unified Multimodal Model

Lance is a lightweight native unified multimodal model open-sourced by ByteDance's Intelligent Creation team. With only 3B active parameters, it supports the full pipeline of image and video understan...

AI AgentMultimodalVideo AI
Jun 21, 2026Read more →
3 months ago

HiDream-O1-Image-Pro – HiDream.ai's Flagship Image Model

HiDream-O1-Image-Pro is HiDream.ai's flagship image generation model built on the native full-modal UiT (Unified Transformer) architecture with over 200B parameters. It sets new SOTA across text-to-im...

MultimodalImage GenerationLLM
Jun 21, 2026Read more →
3 months ago

Gemini 3.5 Flash – Google's Next-Generation AI Foundation Model

Gemini 3.5 Flash is Google's next-generation AI foundation model, positioned around "frontier intelligence + action capability," marking a major breakthrough in efficient inference and Agent capabilit...

AI AgentMultimodalModel Inference
Jun 21, 2026Read more →
3 months ago

HyperEyes – Xiaohongshu and Cambridge's Parallel Multimodal Search Agent

HyperEyes is a parallel multimodal search agent jointly launched by Xiaohongshu and the University of Cambridge. It introduces the UGS (Unified Grounded Search) paradigm, fusing visual grounding and r...

AI AgentMultimodalTool Calling
Jun 21, 2026Read more →
Page 7 of 12 (112 articles total)

Subscribe to AI Model Reviews

Get in-depth reviews the moment a new AI model drops. Know its capabilities and use cases.

We never share your email. Unsubscribe anytime.