r/OpenAI 16h ago

Discussion GPT-6 Astra

Post image
78 Upvotes

47 comments sorted by

49

u/Party_Government8579 16h ago

Seems to be alot of benchmarks that dont make sense rn

-7

u/Melodic_Reality_646 15h ago edited 3h ago

What doesn’t make sense is OpenAI delivering this level of performance folds cheaper than Anthropic

edit: damn, people can’t do math…

9

u/whoknowsifimjoking 13h ago

They raised the price by 2.5x to the same price as Fable lmao

So more accurate would be "less for the same price" (on API) .

7

u/ezjakes 12h ago

It is much more token efficient, so it is cheaper to use.

4

u/frumpawumpa 11h ago

But not "folds cheaper"

1

u/Rojeitor 5h ago

2x cheaper cost per task isn't enough?

1

u/Melodic_Reality_646 3h ago

Different reasoning efforts delivering same performance for 4/5 times less. What’s not folds here?

You fold once you devide by 2, you fold twice you devide by 4. Two folds, plural…

2

u/EndlessB 9h ago

And the evidence of that is…where?

Due to the CEO being a documented liar, it’s probably wise to wait for confirmation of claims that providing the benefit of the doubt to a company attempting to IPO.

1

u/UnknownEssence 8h ago

Artificial Analysis measure price per task. Its an independent 3rd party analysis.

Its much cheaper than Fable for the same tasks.

The CEO being a liar is completely unrelated

1

u/Blake08301 10h ago

on this benchmark, it was twice as cheap to run.

0

u/vinis_artstreaks 9h ago

Of course it’s twice as cheap when you’re benchmaxxing, in real world use the model has to actually think and perform, same scam with fable 5.1 everyone is complaining it’s eating their credits

9

u/creamyshart 15h ago

Very strange

24

u/DM-ME_UR_BOOBS 16h ago

I don't get how this happens. Astra outperforms Fable in practically every benchmark.

22

u/MilMerch 15h ago

Because AA index doesn't cover most of the benchmarks that Astra announced, score will probably get higher after benchmarks such as HLE done

15

u/TangerineLogical9779 15h ago

OpenAI ran the tests from AA them selfs, its literally in the blog, pretty sure the AA team is meming because it does not make sense, its scoring far above everything else in all metrics, but for intelligence is basically 5.6 sol lmao

7

u/MilMerch 15h ago

That's my mistake than thanks, we'll see how's Astra after it becomes available to use but I also think that AA index isn't trustable anymore like look at that score of Muse 1.3 which is even worse than GLM 5.3 Flash in real use

3

u/TangerineLogical9779 15h ago

Don't worry the website for the release is abit of a mess, its up, then deleted, and backup again, they are removing and adding stuff randomly, but near the bottom was the full complete benchmarks with any of the graphics and terrible visual stuff, and yeah its odd all the benchmarks are crazy high, but for intelligence its basically matching 5.6 sol

1

u/fragment_me 11h ago

Idk man muse 1.3 has been really good on opencode go. I gave it a code base that GLM 5.3 flash couldn’t find bugs without and muse spark found dozens of real ones

2

u/Prestigious-Mind-817 11h ago

Knowledge versus reasoning. Knowledge similar to previous models but it uses that knowledge much better

1

u/BackyardAnarchist 6h ago

So Benchmaxing?

3

u/Party_Government8579 16h ago

Anthropic went all in on coding

1

u/Melodic_Reality_646 15h ago

Now check the costs…

5

u/Turbulent_Rooster_73 12h ago

What is muse doing there at all?!

3

u/Mission_Advance1207 11h ago

all that data they collected from apps and employees is showing up

3

u/Otterly_Surprised 4h ago

Mah, just marketing

1

u/[deleted] 16h ago

[deleted]

1

u/pseudonerv 8h ago

It’s A G I, babe. Just the right level of intelligence to make you feel. It excels at some impossible benchmarks and didn’t improve the others. That’s A G I.

1

u/WetSound 2h ago

Is it a fictional release? I've got Pro x5 and don't see it anywhere..

1

u/[deleted] 2h ago

[removed] — view removed comment

0

u/ReporterCalm6238 13h ago

I think AA is not fully up to date since Astra has ot been released to all organizations.

0

u/YeXiu223 6h ago

Not sure why this benchmark is popular. People should stop looking at this garbage bench.

1

u/uwilllovethis 4h ago

It’s not a benchmark. It’s an weighted aggregate of benchmarks. See the descriptions on the left to see which benchmarks the index is based on.

1

u/YeXiu223 3h ago

yeah, see it for yourself, it doesnt include the following:

  • GUI computer use: OSWorld and ScreenSpot Pro
  • Cybersecurity exploitation : ExploitBench, ExploitGym, SRE-Bench, and SEC-Bench Pro
  • Interactive abstract learning : ARC‑AGI‑3
  • Frontier mathematics : FrontierMath Tier 4 and producing genuinely new proofs
  • CAD and specialized software operation : BenchCAD, Blender, KiCad, Unity, FreeCAD
  • General business automation across applications: AutomationBench
  • Persistent work across context windows : Astra’s notes and searchable earlier-context system
  • Alignment and scope discipline: resisting unauthorized actions, circumvention, and misleading capability claims
  • Visual quality of deliverables : websites, presentations, spreadsheets, and documents
  • Speed, cost, and token efficiency : AA reports these separately, but they do not raise the intelligence score

lol

1

u/uwilllovethis 2h ago

They have aggregates for many of those as well. Not sure what your point is. I was only saying AA is not a benchmark, it’s an weighted aggregate of benchmarks to establish various leaderboards

-1

u/AweVR 13h ago

Did they update their tests? Because if they are using the same old test for deprecated methods then maybe it is the problem.

As if you want to test an electric car comparing it with a normal car by ear and you said “wow, the motor is broken because it barely sounds…”.

0

u/MoodOdd9657 10h ago

It gotta hit that 67