Evaluating Recent Claims of AGI by NVIDIA CEO Jensen Huang and OpenAI President Greg Brockman

An examination of the publicly available evidence for the proposition that GPT-6 Astra has achieved Artificial General Intelligence.

In September 2026, following the release of GPT-6 Astra, NVIDIA CEO Jensen Huang stated that “AGI has arrived.” See Huang's statement ↗

His statement followed OpenAI President Greg Brockman's description of the release as the beginning of “the AGI era.” Brockman said that people may eventually look back on this period as when AGI was created, potentially with GPT-6 Astra, adding: “For me personally, I do think we're there.” See Brockman's remarks ↗

Neither statement specified a particular definition, test, or evidentiary threshold for AGI.

We therefore compare GPT-6 Astra's demonstrated capabilities against tests and frameworks that researchers have proposed for evaluating general intelligence. This report examines the publicly available evidence for the proposition that GPT-6 Astra has achieved Artificial General Intelligence (AGI).

Evidence available for GPT-6 Astra

ARC Prize independently evaluated Astra on ARC-AGI-3 ↗, an interactive benchmark in which agents explore unfamiliar environments, infer goals, construct models of how those environments work and plan actions without natural-language instructions.

Astra scored 62.7% using ARC Prize's standard provider-neutral harness and 99.9% using a provider adapter that preserves OpenAI's reasoning state between requests. ARC Prize also found that Astra used fewer actions than its median human baseline on 96% of levels.

ARC Prize describes its objective as measuring progress toward systems capable of acquiring new skills efficiently. It explicitly cautions that saturation of ARC-AGI-3 should not itself be interpreted as proof of AGI.

More broadly, GPT-6 Astra has been evaluated across reasoning, computer use, browsing, software engineering, cybersecurity, science and professional work.

OpenAI reports 97.6% on FrontierMath Tier 4, 72.6% on OSWorld 2.0, 41.4% on AutomationBench, 57.9% on Terminal-Bench 4.0 and 59.3% on Agents' Last Exam, among other results.

These results provide evidence across several dimensions that have been proposed as components or indicators of general intelligence. Other proposed dimensions have not been directly evaluated on GPT-6 Astra.

OpenAI also disclosed a proposed solution to the Navier–Stokes existence and smoothness problem within days of these statements. According to OpenAI, the proposed proof was generated by an internal model more capable than GPT-6 Astra, while Astra was subsequently used to formalize and verify the result. The proof remains subject to independent scrutiny. Its proximity to the AGI claims makes it relevant evidence of the broader capabilities being developed at OpenAI, while its production by a different system illustrates an important distinction in evaluating the specific claim: capabilities demonstrated across multiple AI systems are not necessarily capabilities demonstrated by GPT-6 Astra.

Proposed tests of general intelligence

There is no generally accepted empirical test for AGI.

This problem predates current frontier models. A survey presented at the 13th International Conference on Artificial General Intelligence (AGI-20) reviewed proposals including the Imitation Game, Lovelace Test, psychometric testing, Piaget-MacGyver Room, Wozniak Coffee Test, Robot Student Test and Nilsson Employment Test.

Other approaches attempt to construct more general measures. Hernández-Orallo and Dowe's Anytime Universal Intelligence Test evaluates adaptive performance across procedurally generated interactive environments and is designed to accommodate agents operating at different levels and speeds of intelligence.

More recent evaluations include ARC-AGI, real-world agent benchmarks such as GAIA, measures of autonomous task duration, and frameworks such as DeepMind's Levels of AGI, which separates the breadth of a system's capabilities from its level of performance.

Longer-standing benchmarks also exist for domains outside text: general game learning through the Arcade Learning Environment and deep reinforcement learning on Atari, open-world robotic autonomy through Open X-Embodiment and the RT-X models, autonomous scientific discovery through work such as the Robot Scientist, and autonomous driving through Waymo's published safety research, including its crash-rate comparisons against human benchmarks.

The Marcus-Brundage 10-task test is also included. Gary Marcus referenced these criteria directly in his response to Huang's statement. Marcus and Miles Brundage proposed ten substantially different tasks and specified successful completion of at least eight as strong evidence for the “G” in AGI.

The table below compares GPT-6 Astra with this broader history of proposed tests and frameworks proposed so far for what would be considered “AGI.”

Table 1

Proposed tests of general intelligence vs. GPT-6 Astra

Proposed tests of general intelligence compared with GPT-6 Astra results
TestWhat it evaluatesGPT-6 Astra results
Imitation GameTuringHuman-equivalent conversational behaviorPrior GPT evidence*
Lovelace TestBringsjord, Bello & FerrucciOrigination of artifacts beyond what the system's designers can account for through its designPrior GPT evidence*
Psychometric AI testsStandardized human cognitive and intelligence testsPrior GPT evidence; GPT-6 benchmark evidence*
Anytime Universal Intelligence TestHernández-Orallo & DoweAdaptive performance across procedurally generated unfamiliar environmentsNot specifically evaluated
Piaget-MacGyver RoomPhysical problem solving using unfamiliar objects and affordancesNo comparable demonstration identified
Wozniak Coffee TestNavigate an unfamiliar home, identify and manipulate the necessary objects and independently make coffeeNo comparable demonstration identified
Robot Student / AGI Preschool testsGoertzelLearn and function across ordinary educational environments requiring multiple abilitiesNo complete demonstration identified
Employment TestNilssonPerform economically valuable occupations ordinarily performed by humansRelevant prior and GPT-6 evidence; complete test not demonstrated*
ARC-AGI-3Skill acquisition through exploration, modeling, goal-setting and planning in unfamiliar environments62.7% standard harness / 99.9% provider adapter
GAIAReal-world reasoning combining multimodality, browsing and tool useRelevant evidence*
Marcus-Brundage 10-task testBreadth across ten tasks spanning media comprehension, games, coding, mathematics, science and creative work~2/10 estimated*
DeepMind Levels of AGIBreadth of capabilities together with performance relative to humansRelevant evidence; not specifically assessed
METR Task-Completion Time HorizonDuration of autonomous tasks completed at specified reliability relative to human task durationPrior GPT evidence*
OSWorld 2.0Autonomous completion of tasks across real computer operating systems and applications72.6%
Humanity's Last ExamExpert-level academic knowledge and reasoning across disciplines57.2% with tools
Autonomous scientific discoveryKing et al., Robot ScientistGeneration of genuinely novel scientific or mathematical knowledgeFormalization and verification evidence; discovery performed by another system
General video-game learningArcade Learning Environment (Atari)Rapid acquisition of previously unfamiliar interactive tasks through experienceRelated ARC-AGI evidence; no broad evaluation identified
BEHAVIOR-1KStanford1,000 everyday household activities in simulation, requiring full-body navigation and manipulation grounded in real human needsNot evaluated
Open-world robotic autonomyOpen X-Embodiment / RT-XPhysical perception, navigation, manipulation, planning and adaptationNo comparable demonstration identified
Autonomous drivingWaymo safety researchSustained real-time perception, prediction, planning and physical action in an open human environmentNo comparable demonstration identified

* “Prior GPT evidence” indicates that relevant capabilities have been demonstrated or evaluated in earlier GPT-family models. Where GPT-6 Astra has not been specifically evaluated against the cited test, those results are treated as evidence of likely capability rather than a recorded GPT-6 Astra result.

Sources

References

  1. Turing, “Computing Machinery and Intelligence” (Mind, 1950)
  2. Bringsjord, Bello & Ferrucci, “Creativity, the Turing Test, and the (Better) Lovelace Test” (2001)
  3. Hernández-Orallo & Dowe, “Measuring universal intelligence: Towards an anytime intelligence test” (2010)
  4. Bringsjord & Licato, “Psychometric Artificial General Intelligence: The Piaget-MacGyver Room”
  5. Nilsson, “Human-Level Artificial Intelligence? Be Serious!” (AI Magazine, 2005)
  6. “Post-Turing Methodology: Breaking the Wall on the Way to Artificial General Intelligence” (AGI-20 proceedings)
  7. ARC Prize — “OpenAI's GPT-6 Astra on ARC-AGI-3” (September 2026)
  8. ARC Prize — GPT-6 Astra ARC-AGI results
  9. ARC Prize — ARC-AGI-3 interactive reasoning benchmark
  10. Morris et al., “Levels of AGI for Operationalizing Progress on the Path to AGI” (Google DeepMind, 2023)
  11. Mialon et al., “GAIA: a benchmark for General AI Assistants” (2023)
  12. Humanity's Last Exam
  13. METR, “Measuring AI Ability to Complete Long Tasks”
  14. OSWorld — benchmarking multimodal agents on real computer environments
  15. Epoch AI — FrontierMath
  16. Terminal-Bench
  17. Marcus & Brundage, “Where will AI be at the end of 2027? A bet” — the ten-task test
  18. King et al., “The Automation of Science” (Science, 2009) — the Robot Scientist Adam
  19. Bellemare et al., “The Arcade Learning Environment: An Evaluation Platform for General Agents” (JAIR, 2013)
  20. Mnih et al., “Human-level control through deep reinforcement learning” (Nature, 2015)
  21. Open X-Embodiment Collaboration, “Robotic Learning Datasets and RT-X Models” (2023)
  22. BEHAVIOR-1K — Stanford benchmark of 1,000 everyday household activities for embodied AI
  23. Waymo — Safety Research publications
  24. Kusano et al., “Comparison of Waymo Rider-Only crash rates by crash type to human benchmarks at 56.7 million miles” (Traffic Injury Prevention, 2025)
  25. Business Insider, “Nvidia's Jensen Huang Says 'AGI Has Arrived' and Congratulates OpenAI” (September 2026)
  26. Stratechery, “An Interview with OpenAI President Greg Brockman About Astra and Alignment” (September 2026)