
Meaningful human-level task evaluation
Real human exams reflect true human-level capability. GPT: SAT Math, Chinese English. Four dimensions: understanding, knowledge, reasoning, calculation. Best for human-level capability assessment.
AGI-Eval is a human-centric benchmark by Microsoft using real standardized exams. Covers college entrance exams, LSAT, math competitions, bar exams, civil service, GMAT, GRE. Provides meaningful human-level task evaluation.
Difficulty: Intermediate
Real Standardized Exams
Uses real human exams: Gaokao, LSAT, GMAT, not synthetic
Four Dimensions
Understanding, knowledge, reasoning, calculation
Multi-domain
Covers law, math, language, civil service exams
Human-Level Assessment
Researchers assess model performance on human exams
Education and Application
Model selection for EdTech and exam assistance
Gaokao, LSAT, math competitions, bar exam, civil service, GMAT, GRE.
AGI-Eval uses real exams; MMLU uses designed MC. Former closer to human-level assessment.
Real reviews and feedback from users