AI Model Library: LLMs, Agents & Dev Tools

Curated reviews of mainstream AI models, agents, and dev tools — capabilities, use cases, and how they compare.

Expert analysisUpdated regularlyRSS Feed

Tag: Multimodal

112 article(s) found · Clear tag

3 months ago

Wall-OSS-0.5 – X Square Robot’s Open Embodied Intelligence Model

Wall-OSS-0.5 is X Square Robot’s open vision-language-action (VLA) model—a 4B-parameter stack on a 3B Qwen2.5-VL backbone achieving zero-shot real-robot deployment without per-task fine-tuning. Gradie...

MultimodalEmbodied AIOpen Source
Jun 21, 2026Read more →
3 months ago

Step 3.7 Flash – StepFun's Open-Source Flash Model Generation

Step 3.7 Flash is StepFun's next-generation open-source Flash model for production-grade Agent deployment. Built on a sparse MoE architecture, it maintains high inference speed (up to 400 tokens/s) wh...

AI AgentMultimodalAI Coding
Jun 21, 2026Read more →
3 months ago

SpaceMind – iFLYTEK’s Agentic Smart-Space Architecture

SpaceMind is iFLYTEK’s agentic architecture for intelligent spaces. It upgrades traditional smart homes into L2.5 proactive agents with perception, understanding, memory, decision-making, and executio...

AI AgentMultimodalSpeech AI
Jun 21, 2026Read more →
3 months ago

SenseNova-U1-8B-MoT-Infographic – SenseTime's Open-Source Infographic-Enhanced Model

SenseNova-U1-8B-MoT-Infographic is SenseTime's open-source infographic-enhanced model built on the unified SenseNova-U1-8B-MoT architecture at 8B parameters. Through targeted data training and reinfor...

MultimodalAI CodingLLM
Jun 21, 2026Read more →
3 months ago

SenseNova-Skills – SenseTime's Open-Source Modular AI Office Skills Library

SenseNova-Skills is an open-source modular AI office skills library from the OpenSenseNova (SenseTime) team, designed for Agent runtimes. It breaks down four core office capabilities—infographic gener...

AI AgentMultimodalBenchmark
Jun 21, 2026Read more →
3 months ago

Seedance 2.0 Mini – ByteDance's Lightweight Video Generation Model

Seedance 2.0 Mini is a cost-efficient lightweight video generation model from ByteDance Volcano Engine for high-frequency short-video production, marketing asset iteration, and early-stage drafts. It ...

MultimodalVideo AIModel Inference
Jun 21, 2026Read more →
3 months ago

Qwen3.7-Plus – Alibaba Tongyi's Agentic Multimodal Large Model

Qwen3.7-Plus is Alibaba Cloud Tongyi Qwen team's next-generation multimodal large model positioned as an "agent foundation," integrating visual perception, language understanding, code generation, and...

AI AgentMultimodalAI Coding
Jun 21, 2026Read more →
3 months ago

Qwen-VLA – Alibaba Tongyi's General Vision-Language-Action Model

Qwen-VLA is Tongyi Lab's general vision-language-action model: Qwen3.5-4B VLM backbone plus 1.15B DiT action decoder. A unified action trajectory prediction framework merges manipulation, navigation, ...

MultimodalEmbodied AILLM
Jun 21, 2026Read more →
3 months ago

Qwen-Robot Suite – Alibaba Tongyi's Physical-World Foundation Model Suite

Qwen-Robot Suite is Alibaba Tongyi Lab's foundation model suite for physical-world intelligence, comprising Qwen-RobotNav (navigation), Qwen-RobotManip (manipulation), and Qwen-RobotWorld (world model...

MultimodalEmbodied AILLM
Jun 21, 2026Read more →
3 months ago

Qwen-Image-Bench – Qwen Team's Text-to-Image Model Evaluation Benchmark

Qwen-Image-Bench is a standardized benchmark dataset from Alibaba's Qwen team (Tongyi Qianwen) for evaluating text-to-image models—1,000 carefully designed test samples covering Chinese and English pr...

MultimodalImage GenerationLLM
Jun 21, 2026Read more →
Page 8 of 12 (112 articles total)

Subscribe to AI Model Reviews

Get in-depth reviews the moment a new AI model drops. Know its capabilities and use cases.

We never share your email. Unsubscribe anytime.