'AI IQ' has emerged, which evaluates AI with a single score, calculating the score based on the results of various benchmarks.



When comparing the performance of AI models, it's often necessary to read ranking tables that list scores from multiple benchmarks, making it difficult to understand 'which model is actually smarter and to what extent.' On May 12, 2026, engineer and entrepreneur Ryan Hsieh announced 'AI IQ,' a project that converts and displays the performance of cutting-edge AI models onto a human IQ scale.

AI IQ — Intelligently Measuring AI Intelligence

https://www.aiiq.org/




The goal of AI IQ is to visualize, instead of looking at detailed benchmark scores for each model, 'where an AI model is located on the IQ bell curve,' 'how the intelligence scores of cutting-edge AI models change over time,' 'how IQ and emotional intelligence (EQ) look side by side,' and 'what the cost per unit of intelligence is in actual use.' Ms. Shay compiled models such as GPT-5.5, Claude Opus 4.7, Gemini 3.1, Grok 4.3, Kimi K2.6, Qwen 3.6, DeepSeek V4, and Muse Spark into an AI IQ diagram.

The AI IQ chart looks like this. At the time of writing, 'gpt-5.5' was the highest score, located on the far right, followed by 'gpt-5.4', 'gemini-3.1-pro', and 'opus-4.7'.



It should be noted that the 'IQ' in AI IQ is not the result of simply having the AI model take a human IQ test, but rather the result of converting scores from publicly available benchmarks in four areas—abstract reasoning, mathematical reasoning, programming reasoning, and academic reasoning—into 'estimated IQs,' and then calculating an overall IQ from the average of these four scores.

Hovering your cursor over each model on the graph will show you the breakdown of how the 'IQ' score was calculated. A total of 12 benchmarks are used for the calculation: ARC-AGI-1, ARC-AGI-2, FrontierMath T1-3, FrontierMath T4, AIME, ProofBench, Terminal-Bench 2.0, SWE-bench, SciCode, Humanity's Last Exam, CritPt, and GPQA Diamond. For benchmarks where it is easy to get high scores through memorization or inclusion in training data, the score upper limit is compressed to prevent the overall IQ from being unnaturally inflated by using only certain benchmarks. Even when data is missing, a mechanism is employed to conservatively fill in the gaps.



You can also filter the models displayed in the section below. For example, filtering by 'xAI' will show only the Grok series models developed by xAI, allowing you to see the evolution of the Grok series at a glance.



There is also a graph called 'IQ vs Effective Cost' that compares the costs of AI. This 'effective cost' is an indicator that multiplies the token price assuming a task of 2 million input tokens and 1 million output tokens by the token usage efficiency of each model. It is said to be a figure that is close to the actual usage cost, including not only the token unit price on the price list, but also how many tokens the model uses to perform the same task. Even with equivalent IQ scores, Gemini's low cost stands out compared to GPT-based and Opus-based models.



There was also a graph that showed how IQ scores had changed over time.



Looking at the models from OpenAI, Anthropic, and Google, it looks like this. You can see they're in a fierce competition to get the highest score.



On the other hand, the method of summarizing the complex capabilities of AI models into a single IQ score has also drawn criticism. VentureBeat reported that criticism has emerged on X (a Japanese online forum) arguing that 'AI capabilities vary greatly in their strengths and weaknesses across different fields, and summarizing them into a single score may only make them appear more precise, potentially misleading the reality.' While AI IQ is an attempt to make benchmark tables easier to read, some argue that the estimated IQ value should not be taken as 'the intelligence of the AI model itself,' but rather as a conversion value to facilitate comparison of multiple benchmarks.

Mr. Shea stated that cutting-edge AI models are becoming difficult to understand based solely on disparate benchmark tables and initial marketing claims, and that AI IQ aims to make it easier to determine 'which AI models are actually worth using.'

in AI, Posted by log1d_ts