Evaluating Recent Claims of AGI by NVIDIA CEO Jensen Huang and OpenAI President Greg Brockman

An examination of the publicly available evidence for the proposition that GPT-6 Astra has achieved Artificial General Intelligence.

In September 2026, following the release of GPT-6 Astra, NVIDIA CEO Jensen Huang stated that “AGI has arrived.” See Huang's statement ↗

His statement followed OpenAI President Greg Brockman's description of the release as the beginning of “the AGI era.” Brockman said that people may eventually look back on this period as when AGI was created, potentially with GPT-6 Astra, adding: “For me personally, I do think we're there.” See Brockman's remarks ↗ Neither statement specified a particular definition, test, or evidentiary threshold for AGI.

We therefore compare GPT-6 Astra's demonstrated capabilities against tests and frameworks that researchers have proposed for evaluating general intelligence.

Our current evaluation is that the published evidence does not yet support concluding that GPT-6 Astra represents the achievement Artificial General Intelligence.

Evidence available for GPT-6 Astra

ARC Prize independently evaluated Astra on ARC-AGI-3 ↗, an interactive benchmark in which agents explore unfamiliar environments, infer goals, construct models of how those environments work and plan actions without natural-language instructions.

Astra scored 62.7% using ARC Prize's standard provider-neutral harness and 99.9% using a provider adapter that preserves OpenAI's reasoning state between requests. ARC Prize also found that Astra used fewer actions than its median human baseline on 96% of levels.

ARC Prize describes its objective as measuring progress toward systems capable of acquiring new skills efficiently. It explicitly cautions that saturation of ARC-AGI-3 should not itself be interpreted as proof of AGI.

More broadly, GPT-6 Astra has been evaluated across reasoning, computer use, browsing, software engineering, cybersecurity, science and professional work.

OpenAI reports 97.6% on FrontierMath Tier 4, 72.6% on OSWorld 2.0, 41.4% on AutomationBench, 57.9% on Terminal-Bench 4.0 and 59.3% on Agents' Last Exam, among other results.

These results provide evidence across several dimensions that have been proposed as components or indicators of general intelligence. Other proposed dimensions have not been directly evaluated on GPT-6 Astra.

OpenAI also disclosed a proposed solution to the Navier–Stokes existence and smoothness problem within days of these statements. According to OpenAI, the proposed proof was generated by an internal model more capable than GPT-6 Astra, while Astra was subsequently used to formalize and verify the result. The proof remains subject to independent scrutiny. Its proximity to the AGI claims makes it relevant evidence of the broader capabilities being developed at OpenAI, while its production by a different system illustrates an important distinction in evaluating the specific claim: capabilities demonstrated across multiple AI systems are not necessarily capabilities demonstrated by GPT-6 Astra.

Proposed tests of general intelligence

There is no generally accepted empirical test for AGI.

This problem predates current frontier models. A survey presented at the 13th International Conference on Artificial General Intelligence (AGI-20) reviewed proposals including the Imitation Game, Lovelace Test, psychometric testing, Piaget-MacGyver Room, Wozniak Coffee Test, Robot Student Test and Nilsson Employment Test.

Other approaches attempt to construct more general measures. Hernández-Orallo and Dowe's Anytime Universal Intelligence Test evaluates adaptive performance across procedurally generated interactive environments and is designed to accommodate agents operating at different levels and speeds of intelligence.

More recent evaluations include ARC-AGI, real-world agent benchmarks such as GAIA, measures of autonomous task duration, and frameworks such as DeepMind's Levels of AGI, which separates the breadth of a system's capabilities from its level of performance.

Longer-standing benchmarks also exist for domains outside text: general game learning through the Arcade Learning Environment and deep reinforcement learning on Atari, open-world robotic autonomy through Open X-Embodiment and the RT-X models, autonomous scientific discovery through work such as the Robot Scientist, and autonomous driving through Waymo's published safety research, including its crash-rate comparisons against human benchmarks.

The Marcus-Brundage 10-task test is also included. Gary Marcus referenced these criteria directly in his response to Huang's statement. Marcus and Miles Brundage proposed ten substantially different tasks and specified successful completion of at least eight as strong evidence for the “G” in AGI.

The table below compares GPT-6 Astra with this broader history of proposed tests and frameworks proposed so far for what would be considered “AGI.”

Reading the comparison

These tests and frameworks examine different aspects of intelligence. By evaluating major AGI claims consistently against proposed tests and the available evidence, AGI Society aims to identify where different measures converge, where they conflict, and what additional evidence is needed.

Table 1

Proposed tests of general intelligence vs. GPT-6 Astra

  • GPT-6 Astra tested directly
  • Earlier GPT/OpenAI model evidence
  • Estimate or interpretation
  • No comparable published evaluation
Proposed tests of general intelligence compared with GPT-6 Astra results
TestWhat it evaluatesGPT-6 Astra results
Imitation Game1950TuringHuman-equivalent conversational behaviorEarlier GPT/OpenAI model evidenceEarlier GPT-family models have produced relevant conversational evidence; this is not a recorded test of Astra.
Lovelace Test2001Bringsjord, Bello & FerrucciOrigination of artifacts beyond what the system's designers can account for through its designEarlier GPT/OpenAI model evidenceEarlier GPT-family models have produced relevant evidence; this is not a recorded test of Astra.
Psychometric AI tests2003Bringsjord & SchimanskiStandardized human cognitive and intelligence testsEarlier GPT/OpenAI model evidenceEarlier GPT-family models have been assessed on psychometric tests. Astra has related benchmark results, but no published assessment against a single defined psychometric AGI test was identified.
Autonomous driving2004DARPA Grand Challenge / Waymo safety researchSustained real-time perception, prediction, planning and physical action in an open human environmentNo comparable published evaluationNo published Astra demonstration comparable to sustained autonomous driving was identified.
Employment Test2005NilssonPerform economically valuable occupations ordinarily performed by humansEstimate or interpretationEarlier GPT evidence and Astra results are relevant to parts of the test, but no published demonstration of Astra performing a complete occupation was identified.
Autonomous scientific discovery2009King et al., Robot ScientistGeneration of genuinely novel scientific or mathematical knowledgeEstimate or interpretationAstra was used to formalize and verify a proposed proof, but the reported discovery was produced by a different OpenAI system.
Anytime Universal Intelligence Test2010Hernández-Orallo & DoweAdaptive performance across procedurally generated unfamiliar environmentsNo comparable published evaluationNo published evaluation of Astra using this test was identified.
Wozniak Coffee Testc. 2010Navigate an unfamiliar home, identify and manipulate the necessary objects and independently make coffeeNo comparable published evaluationNo published Astra demonstration of independently completing this physical task was identified.
Robot Student / AGI Preschool tests2010GoertzelLearn and function across ordinary educational environments requiring multiple abilitiesNo comparable published evaluationNo published Astra demonstration covering the complete test was identified.
Piaget-MacGyver Room2012Physical problem solving using unfamiliar objects and affordancesNo comparable published evaluationNo published Astra demonstration comparable to this physical test was identified.
General video-game learning2013Arcade Learning Environment (Atari)Rapid acquisition of previously unfamiliar interactive tasks through experienceEstimate or interpretationAstra's ARC-AGI results provide related interactive-learning evidence, but no broad video-game learning evaluation was identified.
BEHAVIOR-1K2021Stanford1,000 everyday household activities in simulation, requiring full-body navigation and manipulation grounded in real human needsNo comparable published evaluationNo published evaluation of Astra on BEHAVIOR-1K was identified.
GAIA2023Real-world reasoning combining multimodality, browsing and tool useEstimate or interpretationPublished Astra results show related multimodal, browsing, and tool-use capabilities, but no explicit Astra evaluation on GAIA was identified.
DeepMind Levels of AGI2023Breadth of capabilities together with performance relative to humansEstimate or interpretationAvailable Astra evidence is relevant to the framework, but no published assessment assigning Astra a level was identified.
Marcus-Brundage 10-task test2023Breadth across ten tasks spanning media comprehension, games, coding, mathematics, science and creative workEstimate or interpretationApproximately 2 of 10 tasks, estimated from available evidence; this was not a formal test administration.
Open-world robotic autonomy2023Open X-Embodiment / RT-XPhysical perception, navigation, manipulation, planning and adaptationNo comparable published evaluationNo published Astra demonstration comparable to this embodied robotics test was identified.
OSWorld 2.02024Autonomous completion of tasks across real computer operating systems and applicationsGPT-6 Astra tested directly72.6% reported for GPT-6 Astra.
Humanity's Last Exam2025Expert-level academic knowledge and reasoning across disciplinesGPT-6 Astra tested directly57.2% with tools reported for GPT-6 Astra.
METR Task-Completion Time Horizon2025Duration of autonomous tasks completed at specified reliability relative to human task durationEarlier GPT/OpenAI model evidenceEarlier GPT-family models have been evaluated using this measure; no published Astra result was identified.
ARC-AGI-32026Skill acquisition through exploration, modeling, goal-setting and planning in unfamiliar environmentsGPT-6 Astra tested directly62.7% using the standard harness; 99.9% using the provider adapter.

Evidence from earlier GPT or OpenAI models is included for context and must not be read as direct testing of GPT-6 Astra. Estimates and interpretations are AGI Society assessments based on the cited, publicly available evidence.

In most formal mathematical models of AGI, the concept of AGI is a spectrum, with different systems displaying different levels of generality, so that a human might have more general intelligence than a dog, which in turn has more general intelligence than an insect, and so on. However, in popular parlance “AGI” is now often used to mean roughly “human-level AGI”, and when responding to a claim of having “achieved AGI” we will here use the term “AGI” in the same implicitly human-level-oriented sense as the claim itself. Our definition of AGI is set out in the Constitution of the AGI Society.

Sources

References

  1. Turing, “Computing Machinery and Intelligence” (Mind, 1950) ↗
  2. Bringsjord, Bello & Ferrucci, “Creativity, the Turing Test, and the (Better) Lovelace Test” (2001) ↗
  3. Hernández-Orallo & Dowe, “Measuring universal intelligence: Towards an anytime intelligence test” (2010) ↗
  4. Bringsjord & Licato, “Psychometric Artificial General Intelligence: The Piaget-MacGyver Room” ↗
  5. Nilsson, “Human-Level Artificial Intelligence? Be Serious!” (AI Magazine, 2005) ↗
  6. “Post-Turing Methodology: Breaking the Wall on the Way to Artificial General Intelligence” (AGI-20 proceedings) ↗
  7. ARC Prize — “OpenAI's GPT-6 Astra on ARC-AGI-3” (September 2026) ↗
  8. ARC Prize — GPT-6 Astra ARC-AGI results ↗
  9. ARC Prize — ARC-AGI-3 interactive reasoning benchmark ↗
  10. Morris et al., “Levels of AGI for Operationalizing Progress on the Path to AGI” (Google DeepMind, 2023) ↗
  11. Mialon et al., “GAIA: a benchmark for General AI Assistants” (2023) ↗
  12. Humanity's Last Exam ↗
  13. METR, “Measuring AI Ability to Complete Long Tasks” ↗
  14. OSWorld — benchmarking multimodal agents on real computer environments ↗
  15. Epoch AI — FrontierMath ↗
  16. Terminal-Bench ↗
  17. Marcus & Brundage, “Where will AI be at the end of 2027? A bet” — the ten-task test ↗
  18. King et al., “The Automation of Science” (Science, 2009) — the Robot Scientist Adam ↗
  19. Bellemare et al., “The Arcade Learning Environment: An Evaluation Platform for General Agents” (JAIR, 2013) ↗
  20. Mnih et al., “Human-level control through deep reinforcement learning” (Nature, 2015) ↗
  21. Open X-Embodiment Collaboration, “Robotic Learning Datasets and RT-X Models” (2023) ↗
  22. BEHAVIOR-1K — Stanford benchmark of 1,000 everyday household activities for embodied AI ↗
  23. Waymo — Safety Research publications ↗
  24. Kusano et al., “Comparison of Waymo Rider-Only crash rates by crash type to human benchmarks at 56.7 million miles” (Traffic Injury Prevention, 2025) ↗
  25. Business Insider, “Nvidia's Jensen Huang Says 'AGI Has Arrived' and Congratulates OpenAI” (September 2026) ↗
  26. Stratechery, “An Interview with OpenAI President Greg Brockman About Astra and Alignment” (September 2026) ↗