Are You Smarter Than AI?
The best AI models now score 133 on a private IQ test they have never seen, higher than about 98.6% of people. Enter your IQ to see which of 17 models you beat.
Compare on
On the private test AI has never seen
You beat 4 of 17 AI models
An IQ of 120 is higher than about 91% of people. The top model, GPT 6 Astra Ultra, scores 133 on the private test.
GPT 6 Astra Ultra
OpenAI · vision
133Claude-5.5 Opus MAX
Anthropic · vision
133GPT 5.6 Terra Ultra
OpenAI
132Muse Glimmer
Meta
123GLM 5.2
Z.ai
123You
Rank 14 of 18
120Claude-5 Sonnet
Anthropic
115Bing Copilot
Microsoft
87
The private test was written by a Mensa member and has never been online, so no AI could have memorised it. Its scale tops out at about 136. Faded rows are models your score beats.
Already know your score from another test? See its percentile.
AI IQ leaderboard
Every current model, ranked by the private test. Scores as of 1 October 2026.
| Model | Maker | Private test | Public test | Last tested |
|---|---|---|---|---|
| GPT 6 Astra Ultravision | OpenAI | 133 | 148 | 30 September |
| Claude-5.5 Opus MAXvision | Anthropic | 133 | 144 | 28 September |
| GPT 5.6 Terra Ultra | OpenAI | 132 | 143 | 30 September |
| GPT 6.1 Sol Ultravision | OpenAI | 131 | 136 | 29 September |
| Claude-5.1 Fablevision | Anthropic | 130 | 145 | 3 September |
| GPT 6 Luna Maxvision | OpenAI | 130 | 139 | 1 October |
| Kimi K3 | Moonshot AI | 130 | – | 29 September |
| Grok 4.7 High | xAI | 129 | 142 | 24 September |
| Gemini 3.8 Flash | 129 | 142 | 1 September | |
| Claude-5.5 Opus | Anthropic | 128 | 144 | 1 October |
| Qwen 3.8 Max | Alibaba | 126 | 143 | 10 September |
| Muse Glimmer | Meta | 123 | 130 | 18 September |
| GLM 5.2 | Z.ai | 123 | – | 10 September |
| Claude-5 Sonnet | Anthropic | 115 | 136 | 30 September |
| Bing Copilot | Microsoft | 87 | 101 | 10 September |
| DeepSeek V4 Pro | DeepSeek | 84 | 103 | 10 September |
| Manus | Butterfly Effect | 83 | 118 | 10 September |
How fast AI got smarter
The best single score on each test, every quarter since TrackingAI began testing in 2024.
Why there are two scores
The public test is Mensa Norway's online test: 35 visual matrix puzzles that anyone can take. Because it is online, AI models may have met the puzzles during training. Its scale tops out at 151, and the best models are now close to it.
The private test was written by a Mensa member and has never been published, so no model could have memorised it. Most models score 10 to 20 points lower on it, which is the fairer measure of how well they reason about new problems. Its scale tops out at about 136.
That is why we rank by the private test. If your IQ is above a model's private score, you solved fresh matrix puzzles better than it did.
What AI IQ scores do and don't mean
Matrix puzzles are a fair test of pattern reasoning, and the progress is real. In early 2024 the best model scored around 100. Two years later the best models beat almost everyone. But keep four things in mind:
- No time limit. People take these tests against the clock. Models do not get tired, nervous or rushed.
- Words versus pictures. Text-only models get each puzzle described in words, which is a different task from seeing it.
- A simple conversion. Each correct answer is worth about 3 points, and random guessing gives about 63.5. That is a useful mapping, not a test normed on thousands of people.
- One narrow skill. A high score on matrix puzzles does not mean a model understands the world as you do. They still make mistakes no person would.
Other projects estimate "AI IQ" by mapping coding and maths benchmarks onto an IQ scale. Those numbers use a different method and are not comparable with the scores here.
How the scores are measured
The scores come from TrackingAI.org, run by Maxim Lott, which gives every major model the same two tests each week. We use TrackingAI's own conversions:
- Public test: IQ = 63.5 + 3 × (raw score − 5.833), 35 items; maximum 151.
- Private test: IQ = 63.5 + 3 × 0.8823 × ((valid score − 3) × 35 / 14); maximum about 136.
A model's score is the average of its latest runs (up to 7), because a single run can swing by several points. Human IQ uses a mean of 100 and a standard deviation of 15, so the "higher than X% of people" figures above use that scale. Want to see the puzzle type for yourself? Try today's daily IQ puzzle.
Frequently asked questions
Which AI has the highest IQ?
What is ChatGPT's IQ?
What is Claude's IQ?
What are Gemini's and Grok's IQ?
Are you smarter than ChatGPT?
Why is the private test score lower than the public one?
How is an AI IQ measured?
Is an AI IQ the same as a human IQ?
How fast is AI getting smarter?
Could you beat the best AI?
Take 25 visual matrix puzzles, the same kind the models are tested on. About 12 minutes, no sign-up to start. Our test and the AI tests are different tests, so treat any comparison as a rough guide.
25 questions · about 12 minutes