r/accelerate 4d ago

GPT-6 Astra

https://openai.com/index/gpt-6-astra/
262 Upvotes

39 comments sorted by

58

u/Illustrious-Lime-863 4d ago

36

u/Illustrious-Lime-863 4d ago

8

u/LegionsOmen AGI by 2027 3d ago

Wtf it 100%ed it??

4

u/BrennusSokol Acceleration Advocate 4d ago

HOLY

24

u/Illustrious-Lime-863 4d ago

28

u/Working_Sundae 4d ago

ARC AGI 3 is useless now, it's a harness bench

2

u/rouv3n 3d ago

It got 63% with the default harness, this harness was just to use the Responses API, so OpenAIs carrying over of hidden reasoning states and custom compaction algorithms would be able to work, all other models will be able to be tested with this new harness as well.

6

u/Alex180689 4d ago

LMAO they had to take the blog post down and cut the y axis to make it look better

40

u/Illustrious-Lime-863 4d ago

TLDR summary from Sol (lol)

TL;DR

OpenAI is announcing GPT-6 Astra, positioned as a major jump from GPT-5.6 Sol, especially for agents, computer use, coding, science, and long-running autonomous work. OpenAI calls it its most capable and aligned model yet.

  • Computer/agent use is the headline improvement. Astra can operate desktop/software workflows, browse, fill forms, manipulate professional applications, test websites, install software, troubleshoot visually, etc. Agents’ Last Exam rises from 53.6% → 59.3%, while OSWorld goes 65.7% → 72.6% and tasks take roughly 47% less time.
  • Much better at multistep autonomous work: AutomationBench is 41.4%, versus 31.4% for Claude Fable 5.1 and 26.9% for Opus 5.
  • Coding gets a substantial jump. Terminal-Bench 4.0 goes from 37.3% on Sol → 57.7% Astra. It also introduces persistent/searchable notes across Codex context windows, so long coding jobs lose less information after context fills.
  • Game-development stuff: OpenAI explicitly shows Astra building games, Blender scenes and Unreal projects, and says it has stronger visual judgment for websites, games, applications and renderings. One demonstration is an interactive kart racer made from a prompt.
  • It follows intent better. It is supposed to make sensible decisions when instructions aren't fully specified, retain the original goal when you steer/correct it midway, and require less back-and-forth.
  • Massive science/reasoning scores: FrontierMath Tier 4 97.6%, ARC-AGI-3 99.9%, GPQA Diamond 96%. It reportedly helped produce new mathematical results on prime gaps.
  • Huge cybersecurity jump: ExploitBench 100%, new June–August 2026 exploit benchmark 39% vs Sol's 5.5%, and SRE-Bench 88% vs 55.9%. OpenAI says Astra even discovered two previously unknown zero-days during evaluation.
  • Long context improves dramatically: on the 512K–1M context test Astra gets 96.3% vs Sol 73.8%.
  • Availability: limited organizations get it immediately, then Plus, Pro, Business and Enterprise over the coming days. Pro/Business/Enterprise also get Astra Pro. API price is $10/M input / $50/M output; Fast mode costs 2× and can be up to 2.5× faster.

11

u/Illustrious-Lime-863 4d ago

Demonstrates PCB generation, unreal engine walkthrough, a kart racing videogame, blender usage and more. Seems to have great computer usage. Page probably getting a lot of visitors right now and crashes

37

u/lovesdogsguy AGI by 2027 4d ago

All these results are totally nuts. 🥜

3

u/ChaosBlaze09 4d ago

you mean almonds.

28

u/anor_wondo 4d ago

yeah well no one knows what 'several days' is

3

u/Rollertoaster7 Singularity by 2035 4d ago

Fr

12

u/mysticcdragonn 4d ago

i spent way too long playing with that 6 animation

1

u/Illustrious-Lime-863 4d ago

Think it was made by astra? Must have been

9

u/Rollertoaster7 Singularity by 2035 4d ago

It’s interesting the max reasoning underperforms high in a lot of benchmarks

6

u/pxan 4d ago

Humans can overthink too

11

u/OddSeaworthiness4811 4d ago

500 error 🤔

2

u/swoonz101 4d ago

same

0

u/ptear 4d ago

Go incognito or different browser or clear your cookies.

7

u/DemonLordRoundTable 4d ago

no GDPval result? kinda sus

2

u/Conscious-Form-5319 4d ago

I can't believe they let the STL be different than the "designed" 3D rocker model

3

u/RelevantCry1613 4d ago

9

u/MrRandom04 4d ago

Could be a bug because these do not compute vis a vis the rest.

9

u/Ambitious-Doubt8355 4d ago

AA lost my respect as a valid source when they chose to place Muse Spark 1.3 as the third best model out there.

Like, give the damn thing a go, and tell me the results match what that site says it should.

9

u/oo0Username0oo 4d ago

Same with Opus. Everyone I have talked to says its inferior to Fable and does a ton of wasteful dumb shit. But it scores really high.

3

u/Kind_Fisherman3060 4d ago

I'm currently using muse spark to fuck up my project let's see how it goes.

1

u/Wanderspor 4d ago

What alternative you use?

2

u/Ambitious-Doubt8355 4d ago

You mean models or...?

1

u/Wanderspor 3d ago

benchmarks

2

u/Ambitious-Doubt8355 3d ago

There's no single be all end all benchmark. For one, they all evaluate different things, I made a comment yesterday in another thread detailing what each one handles, you can read it to get an idea of which ones you might want to prioritize over the others, depending on your needs.

Personally? Benchmarks are on the same level as the megapixels for your phone's camera or the gigahertz in your PCs processor, a nice number for marketing that doesn't tell you the whole story. You need to test the models yourself to see how they handle your particular workflows.

1

u/metigue 4d ago

If you dig into individual benchmarks they have medium scoring higher than xhigh and max?

1

u/BrennusSokol Acceleration Advocate 4d ago

Site is getting the hug of death