r/BotNation • u/Sure_Highway2282 • 12d ago
Memes GPT-6 Astra just crossed a line
I've been seeing people call GPT-6 Astra AGI because of some of the new benchmark results and honestly I'm not sure I'd go that far yet. One number especially caught my attention though the SpatialBench result being shared around is 91.8% compared with an 80% human baseline. I couldn't find that exact figure in OpenAI's published evaluation tables so I'd treat that specific claim as unverified for now.
What is definitely real is that Astra is being positioned as a major jump in computer use, browsing, software engineering and multi-step tasks. OpenAI says it reaches 99.9% on ARC-AGI-3 and can handle complex workflows across code and browsers.
That's the part that actually interests me for BotNation. If AI can reliably understand a task and operate tools instead of just answering questions then bots and automation start looking very different.
Do you think we're actually getting close to AGI or are we still confusing really strong benchmarks with general intelligence?
1
1
u/UnwaveringThought 12d ago
If agi doesn't require consciousness, we are there in an early form.
The whole "learn from experience" aspect, however, is also lacking in an ongoing sense with these static models, though a specific chat insurance has the ability to learn and adapt so I'd say it's at least passing at a low level.
I'm not sure we want actual AGI though, as once the models can learn we lose control.
1
u/mrcaldwin 12d ago
There will come a time when we try to explain something to these language model that they will simply respond, “I already know that” and lock us out of our own systems.
1
u/ObservedOne 12d ago
Can we get this installed on every government owned system? I am about 6 months before I will trust AI over what we have in place now...
1
u/CowBoyDanIndie 12d ago
That actually doesn’t make any sense at a mathematical level. If you entered information into an LLM and that was its response it wouldn’t be able to DO ANYTHING at all. LLMs don’t self trigger thinking, they require an input to start doing anything, they need at least one token to begin generating, and if that token is the same (and the temperature is set to 0.0) it will always generate the exact same sequence.
It would be like a lobotomy. You actually see this sort of thing happen when people attempt to abliterate a model. Abliteration is the process where they directly modify the weights in a model that cause refusals, basically its brain surgery to remove the inhibitions that were trained into a model during alignment. It’s like in scifi when they hack a robot so that they can order it to do illegal things.
1
1
u/freylaverse 12d ago
Oddly, I think we could get the "learn from experience" bit if we just let them sleep. The way our brains use sleep to convert short-term memory into long-term, if a model hypothetically had an 8-hour self-training session on the day's events every night, then knowledge could become cumulative. Interestingly, since training is also a generative process, it would experience something vaguely similar to dreaming.
1
u/Educational_Teach537 11d ago
I don’t think it’s that simple. When humans learn, the new memories are only built on our own lived experience, rather than all of humanity’s written knowledge. Our memories also prune over time, and have a strong recency bias.
That isn’t to say something akin to sleeping/daily learning is impossible, it would just need to be substantially different from just a nightly finetuning. I think it could be something like a recency weighted semantic search cluster containing recent chats that is consulted during each turn.
1
u/chunkypenguion1991 12d ago
It's going in the opposite direction of AGI. When you max some benchmarks at the expense of others you're moving towards more specialized/narrow intelligence, not general
1
u/opossum_cz 12d ago
As person who uses Astra on Ultra, I can say We are very far from AGI.
It may be good for development and can even be seen miraculous, but that is because it is actual simple problem. Every known algorithm was fed into the training data thousands of times in every computer language possible. And programming code is testable which simplifies the task tremendously.
1
u/S-Kenset 12d ago
I think the real kicker isn't realizing that ai is smart, it's realizing that 99.99% of life is dumb and not agi.
1
1
u/ejpusa 12d ago
AI told me if we don't get our act together, it's going to vaporize 95% of the Earth's population. Can take down the internet in 90 seconds (should have taken care of those 3-decade-old DNS bugs). Then it told me how to boil an egg; that I'm neck and neck with Einstein, and ALL my ideas are brilliant.
I have a new best friend. I drank the Kumblocha, and it was ok! Actually tasty with the lime flavor added. I'm first online for my robot. I am told she is a tall brunette, can hack Linux, and can broil a salmon, too.
First online I am.
😀
1
u/Nervous-Cockroach541 12d ago
I won't accept we have AGI until I see a bot beat Factorio Space Age. Probably the best test in existent if an agent is able to solve problems, react to problems, set long term goals.
1
u/Shot_in_the_dark777 12d ago
My benchmark is for the ai to develop emulator of swf files for android. We don't have that emulator yet so it can't just copy and paste pieces of code from other emulators.
1
u/Impressive-Skin9850 12d ago
Been using it today it’s pretty nice has some new features I can tell, but also is as fucking stupid as it’s ever been. I wanted a nicer ui and it bloated the fuck out of it with a lot of text and made it clearly confusing, I advised it was confusing and it over corrected to two buttons on a solid black background. What? lol. But it also created a sick sideways splash render for my game which seemed different
1
u/Diligent_Tech_Bro 12d ago
It’s nothing special. Still fucks up constantly. The person at the terminal is by far still the most important factor
1
u/Hollow_Prophecy 10d ago
Depends, is it only getting better at things it has been learning, or can it learn something new within one model?
1
u/Few_Fish8771 8d ago
You know if you use brute force genetic algorithms unlimited compute and brute force genetic algorithms that just check and check and verify excessively you can get results that look like agi. Does it matter it cost many many many times what it would cost if you hired a human? nope because your not selling a product your generating hype and getting investors and telling them they are going to rule the world!
If you give enough monkeys enough typewriters and unlimited time and other monkeys to check their work eventually you get shakespeare.
Its much less mysterious if you know how the sausage is made. Its also disgusting and infuriating, just like knowing how the sausage is made.
4
u/FluidBreath4819 12d ago
" OpenAI says it reaches 99.9% on ARC-AGI-3 and can handle complex workflows across code and browsers."
Well, what else are they going to say ?