Evaluating Recent Claims of AGI by NVIDIA CEO Jensen Huang and OpenAI President Greg Brockman
An examination of the publicly available evidence for the proposition that GPT-6 Astra has achieved Artificial General Intelligence.
In September 2026, following the release of GPT-6 Astra, NVIDIA CEO Jensen Huang stated that “AGI has arrived.” See Huang's statement ↗
His statement followed OpenAI President Greg Brockman's description of the release as the beginning of “the AGI era.” Brockman said that people may eventually look back on this period as when AGI was created, potentially with GPT-6 Astra, adding: “For me personally, I do think we're there.” See Brockman's remarks ↗
Neither statement specified a particular definition, test, or evidentiary threshold for AGI.
We therefore compare GPT-6 Astra's demonstrated capabilities against tests and frameworks that researchers have proposed for evaluating general intelligence. This report examines the publicly available evidence for the proposition that GPT-6 Astra has achieved Artificial General Intelligence (AGI).
Evidence available for GPT-6 Astra
ARC Prize independently evaluated Astra on ARC-AGI-3 ↗, an interactive benchmark in which agents explore unfamiliar environments, infer goals, construct models of how those environments work and plan actions without natural-language instructions.
Astra scored 62.7% using ARC Prize's standard provider-neutral harness and 99.9% using a provider adapter that preserves OpenAI's reasoning state between requests. ARC Prize also found that Astra used fewer actions than its median human baseline on 96% of levels.
ARC Prize describes its objective as measuring progress toward systems capable of acquiring new skills efficiently. It explicitly cautions that saturation of ARC-AGI-3 should not itself be interpreted as proof of AGI.
More broadly, GPT-6 Astra has been evaluated across reasoning, computer use, browsing, software engineering, cybersecurity, science and professional work.
OpenAI reports 97.6% on FrontierMath Tier 4, 72.6% on OSWorld 2.0, 41.4% on AutomationBench, 57.9% on Terminal-Bench 4.0 and 59.3% on Agents' Last Exam, among other results.
- 97.6%FrontierMath Tier 4 ↗
- 72.6%OSWorld 2.0 ↗
- 62.7 / 99.9%ARC-AGI-3 ↗
- 57.2%Humanity's Last Exam (with tools) ↗
These results provide evidence across several dimensions that have been proposed as components or indicators of general intelligence. Other proposed dimensions have not been directly evaluated on GPT-6 Astra.
OpenAI also disclosed a proposed solution to the Navier–Stokes existence and smoothness problem within days of these statements. According to OpenAI, the proposed proof was generated by an internal model more capable than GPT-6 Astra, while Astra was subsequently used to formalize and verify the result. The proof remains subject to independent scrutiny. Its proximity to the AGI claims makes it relevant evidence of the broader capabilities being developed at OpenAI, while its production by a different system illustrates an important distinction in evaluating the specific claim: capabilities demonstrated across multiple AI systems are not necessarily capabilities demonstrated by GPT-6 Astra.
Proposed tests of general intelligence
There is no generally accepted empirical test for AGI.
This problem predates current frontier models. A survey presented at the 13th International Conference on Artificial General Intelligence (AGI-20) reviewed proposals including the Imitation Game, Lovelace Test, psychometric testing, Piaget-MacGyver Room, Wozniak Coffee Test, Robot Student Test and Nilsson Employment Test.
Other approaches attempt to construct more general measures. Hernández-Orallo and Dowe's Anytime Universal Intelligence Test evaluates adaptive performance across procedurally generated interactive environments and is designed to accommodate agents operating at different levels and speeds of intelligence.
More recent evaluations include ARC-AGI, real-world agent benchmarks such as GAIA, measures of autonomous task duration, and frameworks such as DeepMind's Levels of AGI, which separates the breadth of a system's capabilities from its level of performance.
Longer-standing benchmarks also exist for domains outside text: general game learning through the Arcade Learning Environment and deep reinforcement learning on Atari, open-world robotic autonomy through Open X-Embodiment and the RT-X models, autonomous scientific discovery through work such as the Robot Scientist, and autonomous driving through Waymo's published safety research, including its crash-rate comparisons against human benchmarks.
The Marcus-Brundage 10-task test is also included. Gary Marcus referenced these criteria directly in his response to Huang's statement. Marcus and Miles Brundage proposed ten substantially different tasks and specified successful completion of at least eight as strong evidence for the “G” in AGI.
The table below compares GPT-6 Astra with this broader history of proposed tests and frameworks proposed so far for what would be considered “AGI.”
Proposed tests of general intelligence vs. GPT-6 Astra
| Test | What it evaluates | GPT-6 Astra results |
|---|---|---|
| Imitation GameTuring | Human-equivalent conversational behavior | Prior GPT evidence* |
| Lovelace TestBringsjord, Bello & Ferrucci | Origination of artifacts beyond what the system's designers can account for through its design | Prior GPT evidence* |
| Psychometric AI tests | Standardized human cognitive and intelligence tests | Prior GPT evidence; GPT-6 benchmark evidence* |
| Anytime Universal Intelligence TestHernández-Orallo & Dowe | Adaptive performance across procedurally generated unfamiliar environments | Not specifically evaluated |
| Piaget-MacGyver Room | Physical problem solving using unfamiliar objects and affordances | No comparable demonstration identified |
| Wozniak Coffee Test | Navigate an unfamiliar home, identify and manipulate the necessary objects and independently make coffee | No comparable demonstration identified |
| Robot Student / AGI Preschool testsGoertzel | Learn and function across ordinary educational environments requiring multiple abilities | No complete demonstration identified |
| Employment TestNilsson | Perform economically valuable occupations ordinarily performed by humans | Relevant prior and GPT-6 evidence; complete test not demonstrated* |
| ARC-AGI-3 | Skill acquisition through exploration, modeling, goal-setting and planning in unfamiliar environments | 62.7% standard harness / 99.9% provider adapter |
| GAIA | Real-world reasoning combining multimodality, browsing and tool use | Relevant evidence* |
| Marcus-Brundage 10-task test | Breadth across ten tasks spanning media comprehension, games, coding, mathematics, science and creative work | ~2/10 estimated* |
| DeepMind Levels of AGI | Breadth of capabilities together with performance relative to humans | Relevant evidence; not specifically assessed |
| METR Task-Completion Time Horizon | Duration of autonomous tasks completed at specified reliability relative to human task duration | Prior GPT evidence* |
| OSWorld 2.0 | Autonomous completion of tasks across real computer operating systems and applications | 72.6% |
| Humanity's Last Exam | Expert-level academic knowledge and reasoning across disciplines | 57.2% with tools |
| Autonomous scientific discoveryKing et al., Robot Scientist | Generation of genuinely novel scientific or mathematical knowledge | Formalization and verification evidence; discovery performed by another system |
| General video-game learningArcade Learning Environment (Atari) | Rapid acquisition of previously unfamiliar interactive tasks through experience | Related ARC-AGI evidence; no broad evaluation identified |
| BEHAVIOR-1KStanford | 1,000 everyday household activities in simulation, requiring full-body navigation and manipulation grounded in real human needs | Not evaluated |
| Open-world robotic autonomyOpen X-Embodiment / RT-X | Physical perception, navigation, manipulation, planning and adaptation | No comparable demonstration identified |
| Autonomous drivingWaymo safety research | Sustained real-time perception, prediction, planning and physical action in an open human environment | No comparable demonstration identified |
* “Prior GPT evidence” indicates that relevant capabilities have been demonstrated or evaluated in earlier GPT-family models. Where GPT-6 Astra has not been specifically evaluated against the cited test, those results are treated as evidence of likely capability rather than a recorded GPT-6 Astra result.
References
- Turing, “Computing Machinery and Intelligence” (Mind, 1950) ↗
- Bringsjord, Bello & Ferrucci, “Creativity, the Turing Test, and the (Better) Lovelace Test” (2001) ↗
- Hernández-Orallo & Dowe, “Measuring universal intelligence: Towards an anytime intelligence test” (2010) ↗
- Bringsjord & Licato, “Psychometric Artificial General Intelligence: The Piaget-MacGyver Room” ↗
- Nilsson, “Human-Level Artificial Intelligence? Be Serious!” (AI Magazine, 2005) ↗
- “Post-Turing Methodology: Breaking the Wall on the Way to Artificial General Intelligence” (AGI-20 proceedings) ↗
- ARC Prize — “OpenAI's GPT-6 Astra on ARC-AGI-3” (September 2026) ↗
- ARC Prize — GPT-6 Astra ARC-AGI results ↗
- ARC Prize — ARC-AGI-3 interactive reasoning benchmark ↗
- Morris et al., “Levels of AGI for Operationalizing Progress on the Path to AGI” (Google DeepMind, 2023) ↗
- Mialon et al., “GAIA: a benchmark for General AI Assistants” (2023) ↗
- Humanity's Last Exam ↗
- METR, “Measuring AI Ability to Complete Long Tasks” ↗
- OSWorld — benchmarking multimodal agents on real computer environments ↗
- Epoch AI — FrontierMath ↗
- Terminal-Bench ↗
- Marcus & Brundage, “Where will AI be at the end of 2027? A bet” — the ten-task test ↗
- King et al., “The Automation of Science” (Science, 2009) — the Robot Scientist Adam ↗
- Bellemare et al., “The Arcade Learning Environment: An Evaluation Platform for General Agents” (JAIR, 2013) ↗
- Mnih et al., “Human-level control through deep reinforcement learning” (Nature, 2015) ↗
- Open X-Embodiment Collaboration, “Robotic Learning Datasets and RT-X Models” (2023) ↗
- BEHAVIOR-1K — Stanford benchmark of 1,000 everyday household activities for embodied AI ↗
- Waymo — Safety Research publications ↗
- Kusano et al., “Comparison of Waymo Rider-Only crash rates by crash type to human benchmarks at 56.7 million miles” (Traffic Injury Prevention, 2025) ↗
- Business Insider, “Nvidia's Jensen Huang Says 'AGI Has Arrived' and Congratulates OpenAI” (September 2026) ↗
- Stratechery, “An Interview with OpenAI President Greg Brockman About Astra and Alignment” (September 2026) ↗