
Shanghai AI Lab's open-source LLM...
8 capabilities, 29 tasks for comprehensive evaluation. CompassRank: LLM and MMBench multimodal leaderboards. Subjective + objective. Best for LLM developers, multimodal researchers.
OpenCompass is Shanghai AI Lab's open-source LLM evaluation platform. CompassKit, CompassHub, CompassRank. 29 core tasks, 100+ datasets. Evaluates 8 capabilities: language, knowledge, logic, creativity, math, code, long text, agents.
Difficulty: Intermediate
Eight Capability Evaluation
Language, knowledge, logic, creativity, math, code, long text, agents
100+ Datasets
29 tasks, 100+ datasets, subjective and objective
MMBench Multimodal
Official multimodal leaderboard for VLMs
Full LLM Evaluation
Developers evaluate models across 8 dimensions
Multimodal Ranking
Rank VLMs via MMBench and related benchmarks
MMBench is OpenCompass multimodal subset for VLM evaluation.
Use CompassKit, configure model and datasets for local or cloud run.
Real reviews and feedback from users