Contents
AGI-Eval

AGI-Eval

Meaningful human-level task evaluation

4.0| Editor Rating
China

Editor Review

Real human exams reflect true human-level capability. GPT: SAT Math, Chinese English. Four dimensions: understanding, knowledge, reasoning, calculation. Best for human-level capability assessment.

AI Tools Navigator Editorial TeamUpdated: 2026-02-11

What is AGI-Eval

AGI-Eval is a human-centric benchmark by Microsoft using real standardized exams. Covers college entrance exams, LSAT, math competitions, bar exams, civil service, GMAT, GRE. Provides meaningful human-level task evaluation.

Basic Info

Category:
Country:China

Best For

Researchers

Difficulty: Intermediate

AGI-Eval Key Features

  • Real Standardized Exams

    Uses real human exams: Gaokao, LSAT, GMAT, not synthetic

  • Four Dimensions

    Understanding, knowledge, reasoning, calculation

  • Multi-domain

    Covers law, math, language, civil service exams

AGI-Eval Key Advantages

  • Real human exams for meaningful evaluation
  • GPT exceeds human average on SAT, Gaokao
  • Four dimensions for analysis

AGI-Eval Use Cases

  • Human-Level Assessment

    Researchers assess model performance on human exams

  • Education and Application

    Model selection for EdTech and exam assistance

Frequently Asked Questions

What exams does AGI-Eval include?▼

Gaokao, LSAT, math competitions, bar exam, civil service, GMAT, GRE.

AGI-Eval vs MMLU?▼

AGI-Eval uses real exams; MMLU uses designed MC. Former closer to human-level assessment.

User Reviews

Real reviews and feedback from users

Write a Review

At least 10 characters

0/500

Please sign in to write a review