AI Model Library: LLMs, Agents & Dev Tools

Curated reviews of mainstream AI models, agents, and dev tools — capabilities, use cases, and how they compare.

Expert analysisUpdated regularlyRSS Feed

Tag: Multimodal

111 article(s) found · Clear tag

2 months ago

MemGUI-Agent – A Long-Horizon Mobile GUI Agent Jointly Developed by Zhejiang University and Kuaishou

MemGUI-Agent is a long-horizon mobile GUI agent jointly developed by Zhejiang University and Kuaishou, specifically designed for cross-app, multi-step, and long-chain mobile automation tasks. Traditio...

AI AgentMultimodalAI for Science
Jul 7, 2026Read more →
2 months ago

AudioX-Turbo – A Unified and Efficient Audio Generation Framework Jointly Released by Noiz AI and Tsinghua University

AudioX-Turbo is a unified and efficient audio generation framework jointly developed by Noiz AI, the Hong Kong University of Science and Technology, and Tsinghua University. Based on a multimodal diff...

MultimodalSpeech AIVideo AI
Jul 7, 2026Read more →
2 months ago

GenEvolve – Self-Evolving Image Generation Agent by Meituan and Others

GenEvolve is a self-evolving image generation agent jointly developed by the Hong Kong University of Science and Technology (Guangzhou), Meituan, and the National University of Singapore. It formalize...

AI AgentMultimodalImage Generation
Jul 7, 2026Read more →
2 months ago

Leanstral 1.5 – Mistral AI's Open-Source Formal Verification Large Model

Leanstral 1.5 is an open-source formal verification large model from Mistral AI, deeply optimized for Lean 4 automated theorem proving. The model adopts a sparse mixture of experts (MoE) architecture ...

MultimodalAI CodingReasoning Model
Jul 6, 2026Read more →
2 months ago

Vidu S1 – Real-Time Interactive Video Foundation Model by Shengshu Technology

Vidu S1 is a globally leading real-time interactive video foundation model launched by Shengshu Technology, marking the transition of AI video generation from offline batch rendering to real-time bidi...

MultimodalSpeech AIImage Generation
Jul 4, 2026Read more →
2 months ago

WorldCupVoice – Open-Source AI Real-Time Sports Commentary System

WorldCupVoice is an AI real-time sports commentary system built on open-source principles. By integrating with Agora RTC live streams, it uses vision models to analyze match footage in real time, gene...

MultimodalSpeech AIOpen Source
Jul 2, 2026Read more →
2 months ago

LiveWorld – Generative Video World Model from University of Adelaide and Others

LiveWorld is a generative video world model jointly developed by the University of Adelaide, the Australian National University, and other institutions. Its core focus is solving the problem of out-of...

MultimodalVideo AIEmbodied AI
Jul 1, 2026Read more →
2 months ago

Nano Banana 2 Lite – Google's Lightweight AI Image Generation Model

Nano Banana 2 Lite is Google's self-developed lightweight AI image generation model, positioned as a speed-first ultra-fast version capable of generating a single image in 4 seconds, with a cost of on...

AI AgentMultimodalImage Generation
Jul 1, 2026Read more →
2 months ago

LocateAnything – NVIDIA's Visual Language Grounding Model

LocateAnything is a visual language grounding model developed by NVIDIA, based on Parallel Box Decoding (PBD) technology. Users can input natural language to precisely select targets in images. With 3...

MultimodalDocument AIEmbodied AI
Jul 1, 2026Read more →
2 months ago

Wan-Streamer – Alibaba's Open-Source Real-Time Full-Duplex Multimodal Foundation Model

Wan-Streamer is an end-to-end real-time full-duplex multimodal foundation model open-sourced by Alibaba DAMO Academy. It unifies text, audio, and video input/output tokens into a single causal sequenc...

MultimodalSpeech AIVideo AI
Jun 30, 2026Read more →
Page 5 of 12 (111 articles total)

Subscribe to AI Model Reviews

Get in-depth reviews the moment a new AI model drops. Know its capabilities and use cases.

We never share your email. Unsubscribe anytime.