# OpenAI Says AGI. The People Whose Test It Beat Say No.

*AI · 5 September 2026 · RECORDx NEWS*

Watch: https://www.youtube.com/watch?v=KFzThGuwrdc
Page: https://news.recordx.co/story/openai-says-agi-the-people-whose-test-it-beat-say-no/

164 headlines said AGI. The people whose test it beat said no.
OpenAI released GPT-6 Astra on 3 September 2026 and its president, Greg Brockman, closed the launch briefing with "Welcome to the AGI era." The number that travelled with it was 99.9% on ARC-AGI-3, a benchmark built to be hard for machines and easy for people.
ARC Prize, who make that benchmark, then published their own post. Verbatim: "While we believe Astra represents meaningful progress towards generalization, we are not claiming that it is AGI." They also published TWO numbers, not one. The 99.9% (for $18,817) comes from a Provider Adapter harness that uses OpenAI's own built-in context management. On the provider-neutral standard harness, the same model scores 62.7% (for $26,098).
That is not a debunking. Francois Chollet, who designed the test, moved his own AGI forecast forward and called Astra "a step-function change in model capability for interactive reasoning problems." It is simply that the number in the headline and the number on a level playing field are not the same number.
And while the AGI argument ran, something else happened that actually changed what the model is allowed to do. Astra is the first OpenAI model ever to cross the "Critical" threshold for cybersecurity risk under the company's own Preparedness Framework. It scored 100% on ExploitBench against 78.5% for GPT-5.6 Sol, and found two previously unknown zero-day vulnerabilities during testing, which OpenAI says it is disclosing. The public version refuses to generate proof-of-concept exploits; OpenAI says it will loosen that for vetted defenders through a programme called Daybreak.

## The reporting

Claim: OpenAI released GPT-6 Astra on 2026-09-03. Its president Greg Brockman closed the launch briefing with "Welcome to the AGI era." The headline number everywhere is 99.9% on ARC-AGI-3. Angle — the correction is the story, and it comes from the benchmark's own authors. Two things nobody led with: ARC Prize, whose test it beat, refuses the claim. Verbatim, on their own blog: "While we believe Astra represents meaningful progress towards generalization, we are not claiming that it is AGI." And the 99.9% is one of two numbers. On the provider-neutral standard harness Astra scores 62.7%. The 99.9% is on a Provider Adapter harness that uses OpenAI's own built-in context management. The AGI argument is the one everybody is having. The one that actually changed what the model is allowed to do is the other one.

## Verdict criteria (fill at 24–48h, Studio ▸ Engagement)

Rank on unique viewers; the gate is Stayed to watch ≥ ~70%. What this reel tests that no previous one has: a tech subject on a channel whose pool has been world news and US politics. GROWTH #6 says subject variety is a reach lever and the ~500-unique ceiling belongs to the US-politics pool. This is the first real probe of a different pool.

## References

- https://youtube.com/shorts/KFzThGuwrdc**

Reporting drawn from: ARC Prize (arcprize.org/blog/astra); OpenAI "Path to Astra" and "Responding to the next frontier of critical cyber capabilities"; Fortune; Axios; Computerworld; CSO Online; The Hacker News, Axios and Fortune

## Transcript

OpenAI released GPT 6 Astra this week, and its president closed the launch briefing by saying, Welcome to the AGI era. The headline number was 99.9 % on ARC -AGI 3, a test built to be hard for machines and easy for people. The people who built that test have now published their own answer.
ARC Prize wrote, We are not claiming that it is AGI, and they published two numbers, not one. The 99.9 came from a harness that uses OpenAI's own context management. On the neutral one, the same model scores 62.7.
That is not a debunking. The man who designed the test moved his own forecast for AGI forward and called the result a step change. It is simply that the number in the headline and the number on the level playing field are not the same number.
And while everyone argued about AGI, something else happened that actually changed what the model is allowed to do. Astra is the first model OpenAI has ever rated critical for cyber risk under its own safety framework. It scored 100 % on a benchmark for writing exploits and found two previously unknown vulnerabilities while it was being tested.
So it ships refusing to write attack code, with looser rules promised later for vetted defenders. That is the part that came with a restriction attached. The AGI question is the argument everybody is having. The cyber rating is the one that changed the rules.

## Topics

openai, ai safety

## Related

- [He Quit Anthropic To Warn About AI. Its Safety Lead Agreed.](https://news.recordx.co/story/he-quit-anthropic-to-warn-about-ai-its-safety-lead-agreed/)
- [He Refused The Pentagon. Seven Others Signed.](https://news.recordx.co/story/he-refused-the-pentagon-seven-others-signed/)
- [OpenAI Froze Its Biggest Training Run. DeepSeek Shipped.](https://news.recordx.co/story/openai-froze-its-biggest-training-run-deepseek-shipped/)
- [OpenAI, Anthropic And Google Want To Write Their Own AI Rules.](https://news.recordx.co/story/openai-anthropic-and-google-want-to-write-their-own-ai-rules/)
