Skip to content

What Is AGI? How Close Are We to Artificial General Intelligence?

AGI, or artificial general intelligence, is the idea of an AI system with broad human-level ability across many different kinds of thinking, rather than strong performance on one narrow job. It does not exist. There is also no universally accepted definition of it and no agreed test for it, which is why serious researchers can look at exactly the same systems and disagree about how close we are.

AI, narrow AI, generative AI and AGI are not the same thing

Three of those four names describe software that already exists. AGI describes software that does not. The argument is about which of the categories today's chatbots belong to.

Artificial intelligence is the umbrella term for software doing things that normally require human perception, language, reasoning or judgment, and it covers spam filters and self-driving cars alike. Narrow AI is any of that built and evaluated for one specific job, however superhuman it is at that job. Generative AI is a capability rather than a level of intelligence: systems that produce new text, images, audio or code. Breadth of output is not breadth of understanding, a distinction worked through in how generative AI sits inside AI as a whole, with the method underneath both handled separately.

AGI is the proposed future category: one system with human-level or better ability across many cognitive tasks, including ones it was not prepared for. Stanford HAI defines it as general, human-level or beyond ability to learn, reason and apply knowledge across a wide range of tasks and domains. Every other term here can be checked against a product. This one cannot, because by definition it is hypothetical.

What Is AI? Real Examples You Already Use covers the narrow-versus-general contrast with everyday examples. The rest of this piece deals with what that comparison cannot settle.

Nobody agrees what AGI means, and that is the actual problem

Stanford HAI states the difficulty inside its own definition: the term is controversial partly because different people mean different things by human-level intelligence, and there is no universally accepted test, so claims are hard to verify.

Three serious attempts show how far apart the targets sit.

OpenAI's Charter, the company's own statement of its mission, defines AGI as highly autonomous systems that outperform humans at most economically valuable work. That is an economic definition, and it is a company's working position rather than a neutral standard.

Google DeepMind researchers led by Meredith Ringel Morris proposed something different in 2023, in work later published at the 41st International Conference on Machine Learning. Rather than a single threshold, they grade systems on two axes, depth of performance and breadth of generality, giving levels that run from emerging up to superhuman. On that framing AGI is a scale, not a finish line.

In October 2025 a group of 33 researchers including Dan Hendrycks, Yoshua Bengio, Gary Marcus and Erik Brynjolfsson posted "A Definition of AGI", a preprint revised again in December, which grounds the question in Cattell-Horn-Carroll theory, an empirically validated model of human cognition, and splits intelligence into ten domains such as reasoning and memory. Without a concrete definition, they argue, the gap between specialised AI and human cognition stays hidden.

These three disagree about what evidence would settle the question. An economic definition could be met by a system that is cognitively lopsided but commercially transformative. A cognitive one could be met by a system nobody deploys. Until a definition is agreed, "we have reached AGI" is not a claim anyone can check.

What current systems actually do, as measured

The measurements are public, and they are strange.

Stanford's 2026 AI Index records large single-year jumps on hard tests. Scores on Humanity's Last Exam, a set of expert-level questions, rose about 30 percentage points in a year, from under 10% to 38.3%. On OSWorld, which tests agents operating a computer, performance went from roughly 12% to 66.3% against a human baseline of 72%.

The same report documents the opposite pattern on tasks most people find trivial: a system that can take a gold medal at the International Mathematical Olympiad may still fail to read an analogue clock, a mismatch the AI Index calls jagged intelligence. The International AI Safety Report 2026, an expert panel assessment chaired by Yoshua Bengio, describes the same shape: leading systems "may excel at some difficult tasks while failing at other, simpler ones", naming counting objects in an image and recovering from basic errors in longer workflows.

Scored against the Hendrycks ten-domain framework, GPT-4 reaches 27% of the human profile it models and GPT-5 reaches 57%. Those numbers depend on accepting that framework, so read them as one group's measurement against one proposed definition. The paper's qualitative finding is more durable: current systems are strong on knowledge-heavy work and weak on foundational machinery, particularly storing new long-term memories.

ARC-AGI-2, released by the ARC Prize Foundation in March 2025, is built from puzzles that at least two human participants solved within two attempts in a controlled study. At launch, ordinary large language models scored zero on it and public reasoning systems scored in single digits. By February 2026 the AI Index recorded leading systems between 66 and 85%. Its successor, ARC-AGI-3, released in 2026, has frontier systems back under 1%. The AI Index adds a caution worth keeping: despite the name, the benchmark tests a specific form of abstraction and pattern inference rather than general intelligence in a broader sense.

Why no authoritative source calls today's chatbots AGI

ChatGPT, Claude and Gemini are general-purpose in the everyday sense: you can ask any of them about tax, poetry or plumbing, which is why the question comes up. That is breadth of subject matter, and it is genuinely new.

That is not general intelligence at the level any of these definitions describes. On the economic definition, these systems do not autonomously perform most economically valuable work. On the cognitive one, they score well short of the human profile and fail specifically on memory and error recovery. They also fail on ordinary tasks no definition thought to include, and the measured limits turn out to be consistently stranger than the intuitive ones. On novelty benchmarks the pattern is that each new test separates systems for roughly a year before it is largely solved, which tells you about the test as much as the systems.

No government or standards body operates any AGI classification scheme at all, the peer-reviewed levels framework from Google DeepMind researchers placed the frontier models of September 2023 at its lowest AGI level and has not been reapplied to current systems, and the companies building them call AGI a goal rather than a shipped product.

Today's systems are broad and unreliable at once, an unfamiliar combination. Why AI makes things up and what an LLM is actually doing explain that better than any AGI headline will.

How close are we? What the evidence supports, and what it does not

Nobody knows, and that is the position the most careful sources take. The International AI Safety Report 2026 is direct about it: the trajectory through 2030 is uncertain, though current trends are consistent with continued improvement, and progress could plausibly slow or plateau, continue at current rates, or accelerate dramatically. It adds that methods for estimating how and when new capabilities emerge remain unreliable. That is an open question, stated as one, in a report written with guidance from over 100 independent experts, including nominees from more than 30 countries and international organizations.

Expert forecasts exist and they move. A survey of 2,778 researchers published at top AI venues, run in October 2023 and reported in the Journal of Artificial Intelligence Research, gave an aggregate 50% chance of high-level machine intelligence by 2047. A smaller and narrower survey the year before, of 738 researchers drawn from two conferences rather than six, put that milestone 13 years later. A forecast shifting 13 years in a little over a year tells you about the volatility of expert opinion, not about a schedule.

Company positions are a third kind of statement. Google DeepMind wrote in April 2025 that AGI, which it described as AI at least as capable as humans at most cognitive tasks, could be here within the coming years. That is the stated view of an organization with a commercial stake in the answer.

None of the three is a date. An expert assessment that keeps the range open is not a countdown, a survey is an average of opinions, and a company statement is a position. Anyone converting them into an arrival year has added certainty the evidence does not contain.

How to read an AGI headline

One reason the question resists answers is that the instruments are poor. François Chollet argued in 2019 that measuring intelligence by skill at specific tasks is the wrong approach, because skill is heavily shaped by prior knowledge and training data: with enough of either you can buy arbitrary levels of skill in a way that hides how well a system generalizes. The 2026 AI Index reports that capability is now outpacing the benchmarks built to measure it. A benchmark win is therefore strong evidence about that benchmark and weak evidence about general ability.

When a claim crosses your feed, four questions do most of the work.

Which definition is being used, and who wrote it? An economic definition and a cognitive one can reach opposite verdicts on the same system.

Separate the result from the forecast. A benchmark score is a measurement. "By 2030" is a projection, and the two should not share a sentence without a label.

Who benefits? A company announcing progress towards a goal it has raised money against is making a company claim. Not automatically false, but not neutral evidence.

Look at what was tested, against what human baseline, and whether the task was novel to the system. Performance on problems resembling training data tells you far less than performance on problems that do not.

The questions that will affect most people soon do not depend on this being settled. Whether AI agents can be trusted to act for you, and which jobs actually change, are answerable now. AGI is a real research problem and also, at present, a word doing a lot of promotional work. Knowing which definition someone is using is usually enough to tell the two apart.

  • AGI: a hypothetical AI system with broad, human-level or greater ability across many cognitive tasks.
  • Narrow AI: AI designed for a specific task or limited set of tasks.
  • Artificial superintelligence: a hypothetical system substantially exceeding human ability across most cognitive domains.
  • Benchmark: a standardised test used to compare systems, and the main evidence behind most AGI claims.
  • Reasoning model: a model that spends more computation working through a problem before answering.
  • Intelligence explosion: a hypothetical rapid rise in ability as AI improvements help produce further improvements.

Frequently Asked Questions

How is AGI different from AI?

Artificial intelligence is the whole field, from spam filters to image generators, and almost all of it is narrow, meaning built and measured for a particular job. AGI describes a proposed future system that handles a wide range of cognitive tasks at human level or better, including tasks it was not prepared for. AI is software that exists today. AGI is a threshold that has not been reached and has not even been agreed on.

Does AGI already exist?

No. The more useful part of the question is who would be entitled to say otherwise, and at the moment nobody is. There is no certifying body, no agreed definition and no accepted test, so an announcement that AGI had arrived would be an organization grading itself against a standard it chose for itself. The organizations most likely to make that announcement are the ones building and selling the systems. Until a definition is settled somewhere outside them, treat any such claim as a position rather than a finding.

Is ChatGPT an AGI?

No government or standards body operates an AGI classification scheme at all, so there is no official answer to point at. ChatGPT and comparable assistants are general-purpose in the sense that you can ask them about any subject, which is why the question comes up, but that is breadth of topic rather than breadth of ability. They fall short of the economic definition, which asks for autonomous performance of most economically valuable work, and of the cognitive one, which asks for a match to the human profile across domains such as memory and reasoning. On the levels-based framework from Google DeepMind researchers, which treats AGI as a scale rather than a threshold, the authors placed the frontier assistants of September 2023 at the lowest AGI rung, the one they call emerging. They have not published a reassessment of later systems.

How far away are we from AGI?

Nobody knows, and the careful sources say so rather than guessing. The International AI Safety Report 2026 leaves the range open: progress could plateau, continue at current rates or accelerate sharply, and it notes that the methods for predicting when new capabilities appear are not reliable. Expert surveys do produce dates, and those dates have moved sharply between one survey and another taken a short time later, which measures how fast opinion moves rather than how fast the technology does. The other half of the answer is definitional. Until people agree what would count as arriving, there is no distance to measure.

Is there a test that would prove a system is AGI?

Not at present. Stanford HAI notes that there is no universally accepted test for AGI, which is why claims about it are hard to verify. Serious proposals exist, including a levels-based framework from Google DeepMind researchers, a cognitive-domain score proposed in a 2025 preprint by a large group of researchers, and novelty-based benchmarks such as ARC-AGI-2. None is a standard, and they would not all return the same verdict on the same system. Choosing a test means choosing a definition, which is the part the field has not settled.

Does AGI mean a conscious or self-aware machine?

No, at least not in any of the definitions researchers work from. The economic version asks what a system can do across valuable work, the levels version asks how deep and how broad its performance is, and the cognitive version scores it against domains such as reasoning and memory. None of them mentions inner experience, and none proposes a way to test for it. Whether a machine could be conscious is a separate philosophical question with no agreed method behind it, and treating the two as one question is a common way for AGI coverage to go wrong.

What is the difference between AGI and superintelligence?

AGI describes broad ability at roughly human level across many cognitive tasks. Artificial superintelligence describes ability substantially beyond humans across most cognitive domains. Both are hypothetical and they are separate ideas, so a claim about one is not a claim about the other. Some researchers argue the gap could close quickly if AI systems start accelerating AI research itself, a scenario called an intelligence explosion, but that is a hypothesis, not an observed process.

Sources

  1. Stanford HAI, "What is AGI (Artificial General Intelligence)?" AI Definitions. https://hai.stanford.edu/ai-definitions/what-is-agi-artificial-general-intelligence
  2. International AI Safety Report 2026, Executive Summary, chaired by Yoshua Bengio. https://internationalaisafetyreport.org/publication/2026-report-executive-summary
  3. International AI Safety Report 2026, Extended Summary for Policymakers. https://internationalaisafetyreport.org/publication/2026-report-extended-summary-policymakers
  4. Stanford HAI, "The 2026 AI Index Report," Chapter 2: Technical Performance, 2026. https://hai.stanford.edu/assets/files/ai_index_report_2026_chapter_2_technical.pdf
  5. Meredith Ringel Morris et al., "Levels of AGI for Operationalizing Progress on the Path to AGI," arXiv, November 2023. https://arxiv.org/abs/2311.02462
  6. Meredith Ringel Morris et al., "Position: Levels of AGI for Operationalizing Progress on the Path to AGI," Proceedings of the 41st International Conference on Machine Learning, 2024. https://proceedings.mlr.press/v235/morris24b.html
  7. Dan Hendrycks et al., "A Definition of AGI," arXiv preprint, v1 October 2025, v3 December 2025. https://arxiv.org/abs/2510.18212
  8. OpenAI, "OpenAI Charter." https://openai.com/charter/
  9. Google DeepMind, "Taking a responsible path to AGI," April 2025. https://deepmind.google/discover/blog/taking-a-responsible-path-to-agi/
  10. Katja Grace et al., "Thousands of AI Authors on the Future of AI," Journal of Artificial Intelligence Research, 2025. https://www.jair.org/index.php/jair/article/view/19087
  11. ARC Prize Foundation, "Announcing ARC-AGI-2 and ARC Prize 2025," March 2025. https://arcprize.org/blog/announcing-arc-agi-2-and-arc-prize-2025
  12. ARC Prize Foundation, "Announcing ARC-AGI-3," 25 March 2026. https://arcprize.org/blog/arc-agi-3-launch
  13. François Chollet, "On the Measure of Intelligence," arXiv, November 2019. https://arxiv.org/abs/1911.01547