4D·

Is Astra really as good as everyone says?

OpenAI announced on September 3, released GPT-6 Astra , which shatters all previous AI benchmarks and can generate a playable first-person shooter map in 28 minutes.


Key metrics from OpenAI’s own release:

  • FrontierMath Tier 4: 97.6%
  • ARC-AGI-3: 99.9%
  • OSWorld 2.0 (computer usage): 72.6% with about 47% less time per task than its predecessor, GPT-5.6 Sol, at 65.7%
  • ExploitBench: 100% compared to 78.5% for Sol (previous model)
  • Artificial Analysis Intelligence Index: 61.2 compared to 60.9 for Sol and 65.7 for Claude Fable 5.1


However, there’s a catch: Astra isn’t designed to be a better chatbot, but rather as computer operator:

It fills out online forms, updates customer records in a CRM system, organizes calendars, conducts research in the browser, creates documents, spreadsheets, and presentations based on templates, and tests finished websites for errors.


Here’s what users have already built with it in the first week:

Astra no longer explains programs like Blender or Unreal Engine—it operates them itself :

  • Game map in 28 minutes: Riley Brown had Astra build a first-person shooter map via Codex with full computer access. The log shows 28 minutes and 16 seconds , 20 edited files and 80 automated tests passed —including weapon and audio timing as well as explosion effects.
  • A playable world requiring no prior knowledge: Matt Shumer, neither a 3D artist nor a game developer, had a survival world generated in Unreal Engine and set up as a playable character.


The point here isn’t the graphics, but the duration: Tasks spanning hours and days that the model pursues on its own have been the biggest challenge so far.


Astra is also the first model that OpenAI has classified as “Critical” for cybersecurity:

It can find and exploit unknown security vulnerabilities without a human guiding every step. The released version therefore refuses to perform offensive tasks such as building exploits; this capability is only unlocked through the Daybreak defender program. Sam Altman spoke of a “new capability level” for the model.


The leap forward, then, lies in tasks that run for hours and require the management of multiple programs. In terms of pure cognitive ability, the gap is small—in the independent Artificial Analysis Index, Astra is practically on par with its predecessor and trails Claude Fable 5.1.

Those who just chat will hardly notice a difference. Those who delegate work certainly will.


What does this mean for the AI sector?

A model that runs for 28 minutes straight consumes many times more resources than a short chat session and drastically increases token usage. Agent-based work is thus primarily a source of demand for more computing power—and that ends up with $NVDA (+0,1%)the hyperscalers ($AMZN (+1,99%)$GOOGL (+2,04%)$MSFT (+0,7%)) and the neoclouds ($CRWV (-0,95%)$NBIS (-1,76%)$IREN (+0,64%)), which lease out this capacity.


Sources:

https://openai.com/index/gpt-6-astra/

https://openai.com/index/safety-overview-gpt-6-astra/

https://www.cnbc.com/2026/09/03/open-ai-astra-gpt-6-cyber.html

https://x.com/rileybrown/status/2095679352927056230

https://x.com/mattshumer_/status/2095596175705399482

12
4 Comentários

imagem de perfil
But Claude has already discovered several cybersecurity vulnerabilities on his own, well before Astra did. That's more of an insight into the competition than actually overtaking them, isn't it?
1
imagem de perfil
@Get_Rich_or_Die_Tryin Yes, that's right. Another major improvement in this model is its ability to use programs independently, which Astra does better than, for example, Claude Fable 5.
1
imagem de perfil
Pazzzi, did u test Astra?
imagem de perfil
Participar na conversa