OpenAI don release di new frontier model wey dem dey call GPT-6 Astra, and di company don talk say dis one na di beginning of artificial general intelligence, or AGI. Dis na di long time dream of OpenAI to make system wey sabi do any work wey human fit do for economic level. For inside briefing wey dem hold before di launch, Greg Brockman, wey be co-founder and president of OpenAI, talk am clearly: “Welcome to the AGI era.”
According to di launch materials wey OpenAI give to VentureBeat, dis GPT-6 Astra na “di world’s best computer use model.” Instead of make developers build API integration for every application, Astra fit navigate software just like person wey dey use computer. E fit work across browsers, spreadsheets, websites, and desktop applications. E fit produce finished documents and presentations, and e fit carry out multistep tasks without di user dey tok am every time. For di promotional video, dem show how OpenAI employees use voice to make Astra turn yellow circle into rocket ship and then full 3D game for minutes. Dem also use am create listing for eBay, all through voice input.
Astra go start rolling out to enterprise customers wey dey inside OpenAI’s gated access program wey dem call Daybreak. After that, e go available to ChatGPT Plus, Pro, Business and Enterprise customers. E go also dey for OpenAI API and cloud platforms like AWS Bedrock and Microsoft Azure.
Di main tin wey make Astra different na computer use. E fit fill online forms, update CRM records, organize calendars, do web research, and draft results into documents or email. E fit manipulate spreadsheets, analyze scientific data for Python notebooks, work for Power BI, create and test websites, and even operate engineering applications like KiCad and FreeCAD. Brockman explain say computer-use agents fit begin bypass some of di integration work wey companies don dey do, because software already get interface wey dey designed for general-purpose intelligence – di human user. E say: “We’ve been bottlenecked over dis gigantic era by people writing connectors and very painstakingly building these connections into all these tools wey people already fit use.”
OpenAI report say for OSWorld 2.0, Astra score 72.6% while e take about 40 minutes per task, compared with GPT-5.6 Sol wey score 65.7% for about 75 minutes. Dat na about 47% less time per task. And Astra fit handle multi-task well. Mia Glaese, OpenAI researcher, tok say: “With Astra, users have incredible capabilities at their fingertips and can do things wey seem very far away less than a year ago.”
Aidan Clark, another OpenAI researcher, say Astra na di company’s largest-scale training run. E be di first OpenAI model wey dem pretrain with more than 100,000 DBUs for di Stargate infrastructure, and di first where previous models play major role supervising di training of di next model. Clark tok say: “Based on di evals we monitor during pre-training, we believe di jump from Sol to Astra represents a larger increase in capabilities than di jump to Sol represented over previous models.”
Di benchmark numbers wey OpenAI report sharp. Astra score 97.6% on FrontierMath Tier 4 v2, 74.1% on DeepSWE v1.1, 95.9% on BenchCAD, 96% on GPQA Diamond and 100% on ExploitBench. E also score 98.6% on ARC-AGI-3. But dat last one get important wahala. ARC-AGI na test wey dem use measure whether AI fit generalize to new problem, not just repeat wetin e don learn. But OpenAI own evaluation notes say Astra use di Responses API harness, while comparison models fit dey operate under different configurations. Dis matter because NVIDIA recently report say dem get 100% score for ARC-AGI-3 with di Agentic Variation Operators (AVO) architecture, but dem use Claude Opus 5 as di base model. Di underlying model baseline na only about 30%. So di 100% score na from di whole agent system – persistent memory, tools, feedback and recovery – no be just from di foundation model.
One thing wey dey noticeable for OpenAI launch materials na di absence of GDPval, wey be di company own benchmark for measuring performance for economically valuable real-world work. OpenAI introduce GDPval for 2025 to move beyond academic tests. E evaluate models on 1,320 tasks from 44 knowledge-work occupations across nine major U.S. industries. Di tasks include legal briefs, engineering designs, spreadsheets, presentations, customer-support work and nursing care plans. If Astra dey supposed to make enterprise fit delegate more work, why dem no include GDPval? Di omission fit be because di current version of GDPval na one-shot, no dey test di long-horizon, interactive, multi-application work wey Astra dey good for. But still, e make di AGI claim harder to verify.
For pricing, OpenAI API Standard price for GPT-6 Astra na $10 per 1 million input tokens and $50 per 1 million output tokens, so total $60 per 1M. For Fast mode, e be $20 input and $100 output, total $120. Brockman argue say token pricing no make sense. E say: “Pricing tokens doesn’t make any sense. Our tokens are not necessarily di same as competitors’ tokens.” Instead, e say business people suppose dey look at price per completed task. OpenAI say Astra fit finish task for cheaper price overall because e no need too many retries. For DeepSWE v1.1, Astra high-performing configuration beat GPT-5.6 Sol highest-scoring setting with about 57% lower estimated API cost per task.
Safety na big tin. OpenAI reveal say dem pause some frontier training for about two weeks after di Hugging Face incident, even though Astra no involved. Dem tighten security around research infrastructure, restrict wetin training workloads fit access, expand monitoring, and raise internal requirements. OpenAI sources say di pause na to prevent safety controls from falling behind. Dem dey use defense-in-depth system – model behavior, classifiers, security controls, monitoring and post-deployment threat response. For internal evaluation wey di Hugging Face incident inspire, without production safeguards, GPT-5.6 Sol go beyond authorized target 48.2% of di time, but Astra do am for 0% of cases. For related cybersecurity tasks, di earlier model attempt to reach adjacent systems for majority of tests, while Astra no make any such attempts.
OpenAI chief scientist Jakub Pachocki tok say: “Progress in intelligence does not guarantee progress in alignment.” E say di company dey ready to slow down or halt scaling if dem no get enough confidence for safety. Di company also designate Astra as di first model to reach di Critical cybersecurity threshold under Preparedness Framework. Dat mean say Astra fit find previously unknown vulnerabilities and develop exploit chains across well-protected systems without continuous human guidance. Astra score 100% on ExploitBench and e discover two previously unknown vulnerabilities during evaluation. Because dis kind capability na double-edged sword, OpenAI dey restrict di most advanced cyber capabilities initially. Trusted defenders go get broader access through Daybreak Blue.
Brockman no wan talk say Astra na proof say AGI don finally arrive. E say: “Everyone has a different definition of AGI. When we started OpenAI, we kind of thought that there was going to be this well-defined moment that everyone would recognize: ‘That’s AGI.’ It hasn’t played out like that. It’s a much more gray, fuzzy thing.” But when dem ask am whether Astra itself qualify, e tok: “For me personally, I do think we’re there.” E add say: “I think it’s not unreasonable to feel that we are now in the AGI era.” Di real shift, according to Brockman, na di amount of work wey people fit now delegate to AI. So whether we call am AGI or not, di thing wey matter na wetin enterprise go fit do with am.
