Curated reviews of mainstream AI models, agents, and dev tools — capabilities, use cases, and how they compare.
32 article(s) found · Clear tag
OpenWorker is an open-source AI desktop agent released by Andrew Ng, fully free under the MIT license. It is not a traditional chatbot, but rather a local-first AI colleague oriented toward "deliverin...
img2threejs is an open-source AI Skill developed by developer hoainho, capable of automatically converting a single reference image into an interactive Three.js 3D model. This tool deeply analyzes the...
SeFi-Image is an open-source text-to-image generation model based on a semantic-priority diffusion architecture, offering three parameter configurations: 1B, 2B, and 5B. This model separates high-leve...
HyOCR-1.5 is a lightweight end-to-end Optical Character Recognition (OCR) expert large model introduced by the Tencent Hunyuan team. With only 1B parameters, it integrates full-stack capabilities incl...
OvisOCR2 is an end-to-end document parsing model developed and fully open-sourced by the Alibaba ATH-MaaS team. It is trained based on the Qwen3.5-0.8B base model and has a parameter scale of only 0.8...
LocateAnything is a visual language grounding model developed by NVIDIA, based on Parallel Box Decoding (PBD) technology. Users can input natural language to precisely select targets in images. With 3...
MiniCPM-V 4.6 is an edge multimodal large language model from ModelBest (OpenBMB) with a 1.3B-parameter LLM backbone, deeply optimized for on-device deployment on mobile hardware. Built on the llama.c...
Chronicles-OCR is the first industry benchmark to cover the full evolutionary trajectory of Chinese script "seven-style transformation" (七体之变)—jointly released by Tencent Hunyuan, the Institute of Inf...
PP-OCRv6 is Baidu PaddlePaddle PaddleOCR team's sixth-generation open OCR model, offered in three tiers—Tiny 1.5M, Small 7.7M, Medium 34.5M—covering browser to server compute. vs. the previous generat...
PaddleOCR-VL-1.6 is a vision-language model (VLM) for document parsing from Baidu's PaddlePaddle team. With only 0.9B parameters, it achieves a new SOTA of 96.33% on the OmniDocBench v1.6 benchmark, s...
Get in-depth reviews the moment a new AI model drops. Know its capabilities and use cases.
We never share your email. Unsubscribe anytime.