Curated reviews of mainstream AI models, agents, and dev tools — capabilities, use cases, and how they compare.
111 article(s) found · Clear tag
MemGUI-Agent is a long-horizon mobile GUI agent jointly developed by Zhejiang University and Kuaishou, specifically designed for cross-app, multi-step, and long-chain mobile automation tasks. Traditio...
AudioX-Turbo is a unified and efficient audio generation framework jointly developed by Noiz AI, the Hong Kong University of Science and Technology, and Tsinghua University. Based on a multimodal diff...
GenEvolve is a self-evolving image generation agent jointly developed by the Hong Kong University of Science and Technology (Guangzhou), Meituan, and the National University of Singapore. It formalize...
Leanstral 1.5 is an open-source formal verification large model from Mistral AI, deeply optimized for Lean 4 automated theorem proving. The model adopts a sparse mixture of experts (MoE) architecture ...
Vidu S1 is a globally leading real-time interactive video foundation model launched by Shengshu Technology, marking the transition of AI video generation from offline batch rendering to real-time bidi...
WorldCupVoice is an AI real-time sports commentary system built on open-source principles. By integrating with Agora RTC live streams, it uses vision models to analyze match footage in real time, gene...
LiveWorld is a generative video world model jointly developed by the University of Adelaide, the Australian National University, and other institutions. Its core focus is solving the problem of out-of...
Nano Banana 2 Lite is Google's self-developed lightweight AI image generation model, positioned as a speed-first ultra-fast version capable of generating a single image in 4 seconds, with a cost of on...
LocateAnything is a visual language grounding model developed by NVIDIA, based on Parallel Box Decoding (PBD) technology. Users can input natural language to precisely select targets in images. With 3...
Wan-Streamer is an end-to-end real-time full-duplex multimodal foundation model open-sourced by Alibaba DAMO Academy. It unifies text, audio, and video input/output tokens into a single causal sequenc...
Get in-depth reviews the moment a new AI model drops. Know its capabilities and use cases.
We never share your email. Unsubscribe anytime.