AI Model Library: LLMs, Agents & Dev Tools

Curated reviews of mainstream AI models, agents, and dev tools — capabilities, use cases, and how they compare.

Expert analysisUpdated regularlyRSS Feed

Tag: Document AI

32 article(s) found · Clear tag

2 months ago

OpenWorker – Andrew Ng's Open-Source, Free, Local-First AI Desktop Agent

OpenWorker is an open-source AI desktop agent released by Andrew Ng, fully free under the MIT license. It is not a traditional chatbot, but rather a local-first AI colleague oriented toward "deliverin...

AI AgentDocument AIOpen Source
Jul 28, 2026Read more →
2 months ago

img2threejs – Open-Source AI Skill for Converting Single Images into Three.js 3D Models

img2threejs is an open-source AI Skill developed by developer hoainho, capable of automatically converting a single reference image into an interactive Three.js 3D model. This tool deeply analyzes the...

AI CodingDocument AIAI Tools
Jul 24, 2026Read more →
2 months ago

SeFi-Image – Open-Source Text-to-Image Model Based on Semantic-Priority Diffusion

SeFi-Image is an open-source text-to-image generation model based on a semantic-priority diffusion architecture, offering three parameter configurations: 1B, 2B, and 5B. This model separates high-leve...

Image GenerationDocument AIBenchmark
Jul 20, 2026Read more →
2 months ago

HyOCR-1.5 – Tencent Hunyuan's Lightweight End-to-End OCR Expert Model Open Sourced

HyOCR-1.5 is a lightweight end-to-end Optical Character Recognition (OCR) expert large model introduced by the Tencent Hunyuan team. With only 1B parameters, it integrates full-stack capabilities incl...

Video AIAI CodingDocument AI
Jul 20, 2026Read more →
2 months ago

OvisOCR2 – End-to-End Document Parsing Model Developed by Alibaba ATH-MaaS Team

OvisOCR2 is an end-to-end document parsing model developed and fully open-sourced by the Alibaba ATH-MaaS team. It is trained based on the Qwen3.5-0.8B base model and has a parameter scale of only 0.8...

Document AIBenchmarkEdge Deployment
Jul 20, 2026Read more →
2 months ago

LocateAnything – NVIDIA's Visual Language Grounding Model

LocateAnything is a visual language grounding model developed by NVIDIA, based on Parallel Box Decoding (PBD) technology. Users can input natural language to precisely select targets in images. With 3...

MultimodalDocument AIEmbodied AI
Jul 1, 2026Read more →
3 months ago

MiniCPM-V 4.6 – OpenBMB's Open-Source Edge Multimodal LLM

MiniCPM-V 4.6 is an edge multimodal large language model from ModelBest (OpenBMB) with a 1.3B-parameter LLM backbone, deeply optimized for on-device deployment on mobile hardware. Built on the llama.c...

MultimodalVideo AIDocument AI
Jun 21, 2026Read more →
3 months ago

Chronicles-OCR – Cross-Temporal Visual Perception Benchmark for Chinese Script Evolution

Chronicles-OCR is the first industry benchmark to cover the full evolutionary trajectory of Chinese script "seven-style transformation" (七体之变)—jointly released by Tencent Hunyuan, the Institute of Inf...

MultimodalDocument AIAI for Science
Jun 21, 2026Read more →
3 months ago

PP-OCRv6 – PaddleOCR's Sixth-Generation Open OCR Model

PP-OCRv6 is Baidu PaddlePaddle PaddleOCR team's sixth-generation open OCR model, offered in three tiers—Tiny 1.5M, Small 7.7M, Medium 34.5M—covering browser to server compute. vs. the previous generat...

Document AIOpen Source
Jun 21, 2026Read more →
3 months ago

PaddleOCR-VL-1.6 – Baidu's Document Parsing Vision-Language Model

PaddleOCR-VL-1.6 is a vision-language model (VLM) for document parsing from Baidu's PaddlePaddle team. With only 0.9B parameters, it achieves a new SOTA of 96.33% on the OmniDocBench v1.6 benchmark, s...

MultimodalDocument AILLM
Jun 21, 2026Read more →
Page 3 of 4 (32 articles total)

Subscribe to AI Model Reviews

Get in-depth reviews the moment a new AI model drops. Know its capabilities and use cases.

We never share your email. Unsubscribe anytime.