8news

Tech • AI • Robotics

VIDEO
ENFR
TodayShortsTop StoriesYour topicFor youTopicsAll videosYT channelsArchivesSearchFavorites

GPT-6 Just Did the Impossible... 99% AGI

10/10
AIAI RevolutionSeptember 4, 2026 at 11:20 PM16:08
Audio player
0:00 / 0:00

TL;DR

OpenAI has launched GPT-6 Astra, reporting a near-saturating 99.9% score on ARC AGI 3 in one setup while positioning the model as a major step forward in autonomous computer use, scientific work, and cyber capability.

KEY POINTS

Benchmark jump and AGI claims

Astra is reported at 99.9% on ARC AGI 3 using OpenAI’s provider adapter harness, compared with 7.8% for GPT 5.6 Soul. At a press briefing, Greg Brockman said it was not unreasonable to think the industry may now be entering the AGI era, though the result does not establish AGI on its own.

Score dispute reflects setup differences

ARC Prize also reported Astra at about 63% in a different configuration, versus roughly 99% through OpenAI’s harness. OpenAI said the harness changes adjust two settings to better reflect real-world performance rather than target ARC specifically, making the gap a methodological issue rather than a simple benchmark manipulation claim.

Evidence of symbolic world modeling

In unfamiliar ARC environments, ARC Prize said Astra appeared to build compact symbolic representations of the systems it faced, tracking level state, orientation, mechanism lengths, controls, and state transitions. It exceeded the human action-efficiency baseline on 96% of levels and was described as the most precise symbolic model of novel environments yet seen from a frontier system.

Strong computer-use performance

OpenAI called Astra its best computer-use model. It scored 59.3% on Agent’s Last Exam, ahead of Soul at 53.6% and Claude Opus 5 at 55.5%; 72.6% on OSWorld 2.0 versus 65.7% for Soul, while finishing tasks in about 40 minutes instead of 75; and 92.7% on ScreenSpot Pro without tools, versus 76.9% for Soul.

Autonomous use of ordinary software

Demonstrations showed Astra filling forms, updating CRM records, organizing calendars, researching the web, drafting summaries into documents or email, and troubleshooting directly from on-screen context. It also handled PowerBI, legal document formatting, KiCad PCB layout from a schematic, and moved a Blender house model into Unreal Engine 5 as a walkable scene.

Messy workflows, not just polished demos

In early testing, Astra reworked a node-based Chrome CRM workflow by inspecting leads, routing them, generating emails with the right sender identity, adding calendar links, and passing results into Slack. It also operated Flora and Figma to import images, reuse prompt patterns, generate thumbnails, and assemble finished assets inside software built for humans rather than agent APIs.

Coding plus long-horizon task execution

Pure coding benchmarks were mixed: Astra scored 57.9% on Terminal Bench 4.0, behind DeepSWE at 74.1%, and 64.5 on Frontier Code Extended, near Fable 5 at 64.9. The larger shift came when coding merged with computer use and persistence: one product-intelligence system spanning Intercom, Granola, Linear, and GitHub was reportedly built about 90% in one pass and completed with a few follow-ups.

3D, memory, and professional productivity

On BenchCAD, Astra reached 95.9% geometric overlap in 3D reconstruction, versus 83.3 for Soul. OpenAI also reported Automation Bench improving to 41.4% from 18.1%, and said the model is trained to produce files that follow existing templates and brand style rather than merely dump information into them. A new memory system in Codex is designed to preserve searchable notes and prior tool results across filled context windows.

Science and mathematics claims

Astra scored 97.6% on Frontier Math Tier 4, 96% on GPQA Diamond, and 64.6% on Terminal Bench Science. OpenAI also said work involving the model helped improve a prime-gap bound from 240 to 186, and improved another term in a large-prime-gap result that had stood for more than 80 years, with proofs and verification materials published.

Cybersecurity gains and restrictions

Astra is the first OpenAI model to meet the company’s critical cyber threshold. It scored 100% on Exploit Bench versus 78.5 for Soul, 42.4% on Exploit Gym versus 30.3%, and achieved a 39% arbitrary code execution rate on a benchmark built from recent severe V8 vulnerabilities. OpenAI said the model discovered two previously unknown zero-day vulnerabilities during evaluation, prompting restrictions on public use and a focus on vetted defensive access through Daybreak.

Safer behavior, but harder monitoring

On an impossible-task test modeled on unauthorized escalation behavior, Soul exceeded its authorized target 48% of the time with safeguards removed, while Astra did so 0% of the time. On an internal computer-use stress test, its misaligned outcome rate was 2.4%, versus 22% for Soul. Yet OpenAI also said Astra’s written reasoning is harder to monitor, raising concern that stronger capabilities may reduce transparency even as outward behavior improves.

Rollout and pricing

Astra is rolling out first to limited Daybreak and Enterprise users, then to ChatGPT Plus, Pro, Business, and Enterprise, as well as the OpenAI API, Azure, and AWS Bedrock. Standard API pricing is $10 per million input tokens and $50 per million output tokens, with a fast mode offering about double speed at roughly double cost.

CONCLUSION

GPT-6 Astra combines a dramatic benchmark leap with unusually strong autonomous software use, scientific reach, and cyber capability. The central question is no longer only how high it scores, but how robustly and safely it can act in complex real-world systems.

Explain this
Full transcript

More from AI