AI Model Library: LLMs, Agents & Dev Tools

Curated reviews of mainstream AI models, agents, and dev tools — capabilities, use cases, and how they compare.

Expert analysisUpdated regularlyRSS Feed

Tag: Multimodal

112 article(s) found · Clear tag

3 months ago

Keye-VL-2.0-30B-A3B – Kuaishou's Open Multimodal Large Model

Keye-VL-2.0-30B-A3B is Kuaishou's fully open multimodal large language model—30B total parameters (MoE, ~3B active)—and the first to bring DeepSeek Sparse Attention (DSA) into multimodal understanding...

AI AgentMultimodalVideo AI
Jun 21, 2026Read more →
3 months ago

JoyAI-Echo – JD's Open-Source Long-Form Audio-Video Generation Framework

JoyAI-Echo is an open-source long-form audio-video generation framework from JD Future Academy, designed for minute-scale multi-shot story generation. Through four technical innovations—a paired cross...

AI AgentMultimodalSpeech AI
Jun 21, 2026Read more →
3 months ago

HPC-Ops – Tencent Hunyuan’s Open-Source Industrial LLM Inference Operator Library

HPC-Ops is an industrial-grade high-performance large-model inference operator library open sourced by Tencent’s Hunyuan AI Infra team. It targets production-scale, high-concurrency inference—not isol...

MultimodalModel InferenceLLM
Jun 21, 2026Read more →
3 months ago

Hojo-ASR-V1 – Hojo’s Open-Source Automatic Speech Recognition Model

Hojo-ASR-V1 is an open-source automatic speech recognition model from startup Hojo. Its four-stage hybrid stack combines Whisper feature extraction, Qwen3-Omni audio encoding, Conformer adaptation, an...

AI AgentMultimodalSpeech AI
Jun 21, 2026Read more →
3 months ago

Hermes Desktop – Nous Research's Hermes Desktop Client

Hermes Desktop is the official desktop client from Nous Research, designed to bring powerful Hermes AI Agent capabilities into a local graphical environment. It supports macOS, Windows, and Linux, wit...

AI AgentMultimodalAI Coding
Jun 21, 2026Read more →
3 months ago

Grok Imagine Video 1.5 – xAI's Image-to-Video Model

Grok Imagine Video 1.5 is xAI's next-generation image-to-video model built on the in-house Aurora autoregressive engine. Upload a single static image and a natural-language prompt to generate a short ...

MultimodalSpeech AIVideo AI
Jun 21, 2026Read more →
3 months ago

Gemini 3.5 Live Translate – Google's Most Real-Time Translation Model

Gemini 3.5 Live Translate is Google's latest real-time speech translation model, built on an end-to-end streaming architecture for near-real-time speech-to-speech translation across 70+ languages. It ...

MultimodalSpeech AILLM
Jun 21, 2026Read more →
3 months ago

Fara1.5 – Microsoft's Browser-Native AI Agent Model Family

Fara1.5 is a browser-native AI agent (CUA, Computer Use Agent) model family from Microsoft Research AI Frontiers, available in 4B, 9B, and 27B parameter sizes. Fine-tuned on Qwen3.5, it uses pure pixe...

AI AgentMultimodalAI Safety
Jun 21, 2026Read more →
3 months ago

EvoQuality – ByteDance's Open-Source Self-Evolving Image Quality Assessment Model

EvoQuality is a self-evolving vision-language model framework jointly developed by ByteDance and City University of Hong Kong, focused on no-reference image quality assessment (NR-IQA). Built on Qwen2...

MultimodalImage GenerationAI Safety
Jun 21, 2026Read more →
3 months ago

EchoBird – Open-Source AI Agent Desktop Management Tool

EchoBird is an open-source AI Agent desktop management tool built with a Tauri + Rust architecture. It unifies installation, configuration, and execution of CLI-driven AI coding agents—Claude Code, Co...

AI AgentMultimodalAI Coding
Jun 21, 2026Read more →
Page 10 of 12 (112 articles total)

Subscribe to AI Model Reviews

Get in-depth reviews the moment a new AI model drops. Know its capabilities and use cases.

We never share your email. Unsubscribe anytime.