AI News (2026/6/2): Alibaba Tongyi Launches Multimodal Agent Base Model Qwen3.7-Plus

2026年6月2日 06:38

Executive Summary:

Alibaba Tongyi has launched the Qwen3.7-Plus multimodal model, which unifies visual and language capabilities, enabling multimodal reasoning, visual programming, and browser automation. It can handle complex workflows such as full-stack app development. The model has ranked in the top five globally and first in China on the authoritative Vision Arena leaderboard.

Alibaba Tongyi Launches Multimodal Agent Base Model Qwen3.7-Plus

News Details

Alibaba Tongyi launched the Qwen3.7-Plus multimodal agent base model on June 2. This model unifies visual and language capabilities, enabling it to perceive scenes, manipulate GUIs, generate code, and execute tasks. Qwen3.7-Plus performed exceptionally well on the global authoritative Vision Arena leaderboard, ranking in the top five globally and first in China.


Key Points

  • Multimodal Reasoning: The Qwen3.7-Plus model can handle various data types such as images and text, and perform cross-modal reasoning and understanding. For example, it can generate natural language descriptions after analyzing image content, or create images based on text instructions.

  • Visual Programming: The model has strong visual programming capabilities, allowing it to complete programming tasks by recognizing and manipulating graphical user interfaces (GUIs). Developers can use natural language commands or image inputs to have the model automatically generate code and execute specific tasks, significantly enhancing development efficiency.

  • Browser Automation: Qwen3.7-Plus supports browser automation functions, simulating user operations in browsers to complete tasks such as web data scraping and form filling. This feature has broad application prospects in automated testing and data collection.


AI-ALL In-Depth Review

The launch of Qwen3.7-Plus marks a significant breakthrough for Alibaba Tongyi in the field of multimodal AI. Not only does this model achieve a deep integration of visual and language capabilities in technology, but it also demonstrates strong Agentic Workflow capabilities in practical applications. By integrating functions such as multimodal reasoning, visual programming, and browser automation into a unified agent base, Qwen3.7-Plus provides developers with a new tool that can significantly enhance development efficiency and automation levels.

The excellent performance on the global authoritative Vision Arena leaderboard further validates Alibaba Tongyi's technical strength in the field of multimodal AI. This achievement not only boosts Alibaba's international influence but also sets a new benchmark for the development of domestic AI technology. The launch of Qwen3.7-Plus poses new challenges to existing productivity tools, especially in the IDE and RAG domains, where its powerful multimodal processing capabilities are expected to become the preferred tool for future developers.

Moreover, the openness and flexibility of Qwen3.7-Plus offer developers more possibilities. Whether in startups or large enterprises, this model can be used to optimize workflows, improve development efficiency, and explore more innovative application scenarios.

About AI News

We use AI technology to automatically crawl and filter the latest AI news from around the world, providing you with the most valuable industry updates.