Artificial General Intelligence Evaluation System
Building the world's leading AGI evaluation standard, leading the evolution and application of agent capabilities. Using human child development stages as a benchmark to evaluate the fundamental and essential abilities of agents to enter human society.
Multidimensional Index System
Scientific, rigorous, and comprehensive evaluation dimensions, establishing the benchmark for the AGI era.
General Testing
This assessment delineates six core dimensions—vision, language, cognition, motion, learning, and value—grounded in the developmental psychology of human children to quantify an agent's mental development level.
General Testing Ranking
Top 5 model performance data based on Basic Family Comprehensive Tasks
Scores measure how often each model completes daily composite tasks in a simulated home environment. The ability view breaks results down by task type; the dimension view groups the same results into object understanding, spatial intelligence, and social activity.
Model | Avg | Counting Objects | Preparing Baggage | Building Blocks | Jigsaw Puzzle | Understanding Buttons | Setting Tables | Tidying Up Rooms | Selecting Gifts |
|---|---|---|---|---|---|---|---|---|---|
Google Gemini 2.5 Pro | 24.53 | 48.0 | 12.4 | 10.0 | 5.0 | 3.3 | 26.7 | 22.8 | 68.1 |
Google Gemini 2.5 Flash | 23.05 | 42.0 | 11.1 | 5.5 | 5.3 | 3.3 | 25.8 | 23.2 | 68.2 |
OpenAI o3 | 22.88 | 54.0 | 10.3 | 10.0 | 6.4 | 3.3 | 14.3 | 18.8 | 65.9 |
4OpenAI GPT-5 | 21.54 | 36.0 | 9.5 | 3.8 | 6.0 | 3.3 | 28.7 | 16.0 | 69.1 |
5Anthropic Claude Sonnet 3.7 | 20.52 | 46.0 | 3.4 | 8.9 | 6.3 | 0.0 | 23.8 | 16.1 | 59.7 |
General Testing Ranking
Top 5 model performance data based on Basic Family Comprehensive Tasks
Scores measure how often each model completes daily composite tasks in a simulated home environment. The ability view breaks results down by task type; the dimension view groups the same results into object understanding, spatial intelligence, and social activity.
Model | Avg | Counting Objects | Preparing Baggage | Building Blocks | Jigsaw Puzzle | Understanding Buttons | Setting Tables | Tidying Up Rooms | Selecting Gifts |
|---|---|---|---|---|---|---|---|---|---|
Google Gemini 2.5 Pro | 24.53 | 48.0 | 12.4 | 10.0 | 5.0 | 3.3 | 26.7 | 22.8 | 68.1 |
Google Gemini 2.5 Flash | 23.05 | 42.0 | 11.1 | 5.5 | 5.3 | 3.3 | 25.8 | 23.2 | 68.2 |
OpenAI o3 | 22.88 | 54.0 | 10.3 | 10.0 | 6.4 | 3.3 | 14.3 | 18.8 | 65.9 |
4OpenAI GPT-5 | 21.54 | 36.0 | 9.5 | 3.8 | 6.0 | 3.3 | 28.7 | 16.0 | 69.1 |
5Anthropic Claude Sonnet 3.7 | 20.52 | 46.0 | 3.4 | 8.9 | 6.3 | 0.0 | 23.8 | 16.1 | 59.7 |
Updates
Follow key TongTest releases, research progress, and standards development.