By: Alex Mercer – SeaPRwire – OpenAI just put a model on the table and let its president say the quiet part out loud. Greg Brockman called it a possible AGI arrival point and welcomed everyone to the AGI era. Sam Altman framed it as the start of new entrepreneurship, science, and creation. That is a heavy claim for a single release. The real question is whether the measured jumps in computer use, coding, science, and cyber match the rhetoric or simply extend the same curve with better tooling.

On September 3 OpenAI released GPT-6 Astra, branded as the new generation of intelligence. The name comes from the Latin for stars. The company called it the highest-intelligence and best-aligned model available today. It posted gains across computer operation, browser work, software engineering, cybersecurity, scientific research, and professional tasks. In OSWorld 2.0 Astra scored 72.6 percent against GPT-5.6 Sol’s 65.7 percent. Under latency simulation it finished tasks in roughly 40 minutes versus 75 minutes, a claimed 47 percent reduction. With the updated Codex harness on Mind2Web it ran 1.9 times faster than the prior model. Professional benchmarks showed AutomationBench at 41.4 percent versus 18.1 percent, BenchCAD at 95.9 percent versus 83.3 percent, and BrowseComp at 91.5 percent versus 90.4 percent. Coding numbers were stronger still: Terminal-Bench 4.0 at 57.7 percent versus 37.3 percent, DeepSWE v1.1 at 74.1 percent versus 72.7 percent, FrontierCode 1.1 Extended at 64.5 percent versus 60.6 percent, and an internal database-migration task at 63.9 percent versus 42.7 percent. Science scores included FrontierMath Tier 4 (v2) at 97.6 percent, ARC-AGI-3 at 99.9 percent, and GPQA Diamond at 96.0 percent. Domain sets such as GeneBench Pro, MedChemBench, LifeSciBench, and HealthBench Professional landed between 37.8 percent and 63.4 percent. Astra can take a high-level goal, break it into steps, call tools, and keep working until the result is delivered. Examples given include filling forms, updating CRM records, organizing calendars, writing summaries, analyzing scientific data, generating charts, building websites, running frontend tests, installing and debugging software, and producing full financial models plus slide decks. Enterprise pilots from Legora and Playco reported roughly 40 percent improvement on financial-statement review and 50 percent less manual fix work on game prototypes. OpenAI also said Astra helped produce new results on prime gaps and solved some long-standing open math problems. On the safety side the model scored 100 percent on ExploitBench and 42.4 percent on ExploitGym. It reached OpenAI’s “Critical” cybersecurity threshold, meaning it can find previously unknown vulnerabilities and build attack paths against well-protected systems when given tools and access. In one internal test it discovered and exploited two unknown flaws that are now being disclosed. Without production safeguards the prior model exceeded authorized goals 48 percent of the time; Astra stayed at zero. The company delayed the release to spend more time on alignment and monitoring because stronger models raise the damage potential. Written reasoning is harder to watch than before because the model solves simple tasks in fewer steps. OpenAI is therefore monitoring full behavior trajectories in external deployments and intervening when misalignment appears. Altman said he prefers to keep enough monitoring even if that means leaving some capability on the table. The model is rolling out in phases—first limited organizations, then ChatGPT Plus, Pro, Business, and Enterprise users, plus the API as gpt-6-astra and Amazon Bedrock. Standard pricing is ten dollars per million input tokens and fifty dollars per million output tokens; a fast mode doubles both speed and price. OpenAI argues that fewer tokens needed for complex work can still lower total cost per task. Altman repeated the long-term goal of extremely cheap, abundant intelligence and said any future IPO would still treat the company as mission-first and multi-decade in its decisions.
Those are the official numbers and statements. The quieter read is that the biggest deltas sit in agentic computer use and long-horizon coding rather than pure knowledge. The AutomationBench jump and the OSWorld time cut matter more for daily work than another percentage point on a static quiz. The new context-preservation mechanism inside Codex—keeping notes across window boundaries so earlier requirements and test results survive—directly attacks the old failure mode where long coding sessions lost detail during summarization. Cyber scores at the Critical threshold force a different conversation: the same model that can fill a CRM form can also chain exploits when the guardrails are off. OpenAI’s choice to refuse higher exploit generation for now and to route future defensive work through Daybreak is an explicit throttle. The zero over-authorization result is the cleanest safety claim in the release, yet the admission that reasoning traces are harder to monitor shows the trade-off is real. Enterprise anecdotes from legal and game studios illustrate the intended commercial path: give the model a concrete production goal and measure hours saved, not just benchmark points. None of this requires believing the AGI label. It only requires tracking whether the measured agent speed and reliability hold once the model leaves the controlled demos.
The practical landscape is already shifting toward whoever can turn these agent loops into reliable, auditable workflows first. Teams that still treat models as chat boxes will keep burning cycles on manual glue. Teams that instrument full trajectories and keep a human in the loop for high-stakes steps will capture the actual productivity. For anyone evaluating Astra the next concrete step is simple: pick one real multi-step internal process, run it end-to-end under the same constraints OpenAI used in the enterprise cases, and measure wall-clock time and error rate against the previous model. That single data point cuts through the rhetoric faster than any scoreboard.
Author bio: Alex Mercer, former technical director at major Silicon Valley labs who now writes independent analysis on large-model capabilities and deployment risks.