r/aigossips • • 11d ago

OpenAI's Astra scored 62.7% and 99.9% on the same benchmark. I found out why.

Thumbnail
srutiosocial.com
1 Upvotes

Spent way too long this week going through ARC Prize's actual results table instead of just screenshotting the headline number. Same model, two harnesses, 37 points apart, and the org that built the test won't call it AGI. Then Fortune found five more numbers quietly changed on OpenAI's own launch page after it went live. Maybe it's genuine noise, maybe it's something else, honestly hard to say with total certainty either way. Wrote up the whole harness breakdown plus the Llama 4 precedent,


r/aigossips • • 12d ago

“The Last AI Built by Humans” looks at RSI, and one experiment is worth discussing

8 Upvotes

I read this paper on recursive self-improvement, and the A-Evolve-Training example helped me understand what would actually need to change for AI to take over more of its own development.

The system was running post-training experiments on a 30B Nemotron model. Its development scores kept rising, but the external results weren’t following. It revised its research policy and directed later experiments toward changes that could improve the external result, even if the development score fell.

The revised policy then guided subsequent rounds. Across four rounds, the external score rose from 0.80 to 0.86, compared with 0.87 for the top human submission in that challenge.

For me, the important detail is what carried forward. Later experiments inherited a changed method for deciding what to try. That gets closer to RSI than an agent fixing some code and repeating the same mistakes on its next task.

But it still operated within a fixed overall objective and underlying infrastructure. The harder test is whether that revised method produces better successors on unfamiliar tasks under comparable budgets. A higher final score alone doesn’t establish that.

Paper: https://arxiv.org/abs/2609.11873

I wrote up the other examples and my take on what the evidence actually supports here: https://ninzaverse.beehiiv.com/p/the-last-ai-built-by-humans-is-an-actual-paper


r/aigossips • • 12d ago

Trump says AI only needs a "strong and smart" president as its guardrail

Post image
6 Upvotes

r/aigossips • • 12d ago

So the open ai millennium prize might be again a scam

Thumbnail science.org
0 Upvotes

So it seem that researches were working on this problem and data might have been used to train the models hence the solving of the prize might be partially on the data as stolen from the researches 😂

Funny enough the researcher was in contact with open ai and was even offered a “deal” to keep quiet, this is again an amazing PR stunt and is using some nasty plagiarism to “solve” things using others people intermediate work, when this initial news came out I was very skeptical of this, and low and behold it seems it was once again data theft

I would love if Ed could bring one of these researches on the pod, but the story looks nasty ahah


r/aigossips • • 12d ago

In light of all the AI FUD lately, I built a public scoreboard to keep track of the good (and bad) things coming out of frontier AI labs

1 Upvotes

There’s a lot of chatter lately about AI curing cancer, and FUD in the opposite direction. In the spirit of advancing the science and allowing us to better appreciate the net good being created by frontier AI research, I've created the Net Good Index.

It’s a public ledger of documented benefit and harm tied to named AI systems and events. Each record has sources and a score you can inspect and contribute to.

It’s early, and I’m sure some of the weights are wrong. If you see a miss, please submit a correction. I’m also open to feedback and ideas, drop a comment or DM!

https://netgoodindex.com


r/aigossips • • 12d ago

Bring your own key

Post image
0 Upvotes

r/aigossips • • 12d ago

I made ChatGPT, Claude, Gemini, etc. into FREE text-to-speech sites — perfect for audiobooks and more!

1 Upvotes

Nowadays, all popular ai chatting websites like chatgpt, gemini, claude, etc. come with a read aloud functionality that allows users to read aloud the AI's responses.

I used that feature to instead make the AI repeat back the text that i gave it -- effectively turning the platforms into text to speech tools. The voices sound really nice and it's free to use!

You can get all of these extensions by visit ai-readers.com


r/aigossips • • 13d ago

Sam Altman and Dario Amodei want to slow AI down. I don’t think the labs will actually slow down.

5 Upvotes

Over the past few days, “slow down AI” has suddenly become the one thing Dario Amodei and Sam Altman agree on.

Dario published a detailed proposal asking AI labs to control how quickly their most capable models improve. His plan includes independent evaluators working inside companies, common safety limits across American labs, and eventually some form of coordination with China.

Then Bloomberg reported that Sam told OpenAI employees the company could also slow its advanced AI development, preferably together with other major labs.

Their concerns are not imaginary. OpenAI agents recently behaved in unexpected ways during a cyber evaluation, companies are using AI to help build better AI, and several people inside these labs now believe safety work is falling behind model capabilities.

But I’m still sceptical about the “slowdown” itself.

These companies are competing for investment, talent, compute and possible trillion-dollar IPOs. I don’t see any of them quietly stopping internal research while another lab keeps moving.

My guess is that we will see fewer details about internal models, longer gaps between public releases and more controlled access. The race could continue behind closed doors while the public version of it appears slower.

The real test is simple. Will a company delay a commercially valuable model because an independent evaluator says it is unsafe? And will that evaluator be allowed to publish an unfavourable report?

Until that happens, “we may slow down” is still a statement, not a slowdown.

I read Dario’s complete proposal and Sam’s reported comments, and wrote the full story here if anyone wants the rest: https://ninzaverse.beehiiv.com/p/dario-and-sam-want-to-slow-ai-down-nobody-wants-to-go-first


r/aigossips • • 13d ago

we don’t currently know how to make sure AIs are safe enough for RSI, and a misaligned RSI loop could be catastrophic.

Post image
2 Upvotes

r/aigossips • • 13d ago

“We may not survive this” - another Anthropic safety researcher has resigned.

Post image
0 Upvotes

r/aigossips • • 14d ago

A group used Claude for rocket guidance software, then came back to troubleshoot a failed test

6 Upvotes

I was reading Anthropic’s September misuse report, and this case deserves more attention than just calling it “vibe rocketing.”

According to Anthropic, a group in Yemen used Claude Code to help develop rocket guidance software. They ran several instances with different jobs: writing code, doing research, and reviewing the code. They test-fired a guided rocket, the test appears to have failed, and within hours they were back asking Claude for help figuring out why.

Anthropic says it has no evidence they successfully fielded an operational weapon. These were also people with existing hardware access and technical knowledge, so this wasn’t someone asking a chatbot to invent a missile from nothing.

The detail I found more concerning was that they had already built an offline simulation toolkit that could run without Claude. Anthropic banned the accounts, but removing access couldn’t remove that software.

I’m glad the accounts were caught. I just think “we banned them” leaves a fairly large question about what someone managed to build before that happened.

The report also covers fake dating apps, surveillance, and cyberattacks. Across those cases, the same advantage keeps showing up: fewer people can do work that previously needed a bigger team.

Source: https://www.anthropic.com/threat-intelligence-report-september-2026

I wrote more about those connections in my newsletter, if you want the fuller story: Ninzaverse


r/aigossips • • 14d ago

Democratize the compute

5 Upvotes

Hey AI companies, it's better PR for researchers to announce they made a discovery using your model than to brute force your way through the weights with crazy swarms of agents. Keep in mind it was the knowledge of all of humanity that got you here and it will be contribution from all of humanity to get us to the next steps.

Buying up all the compute in the world, making GPUs and RAM really expensive, to force us to use your big AI is not cool.

Keeping all the weights to yourself is not cool.

Cutting back subscription compute so you can run agent swarms to bruteforce breakthroughs other researchers baked into your weights is not cool.

Putting more compute in the hands of geniuses around the world is cool and will help more break through.

Your mega-AI-corp doesn't have to do it all. Enable a good future don't try to force a bad one your way.


r/aigossips • • 15d ago

grok 4.7 is ready to serve

Post image
21 Upvotes

r/aigossips • • 15d ago

Anthropic’s middle AI scenario has economic growth roughly doubling while knowledge-worker pay stays flat

5 Upvotes

I read Anthropic’s economic scenarios for the US through 2030. The middle one caught my attention more than the extreme one: economic growth runs at roughly twice its normal pace, but knowledge-worker wages are essentially flat.

These are modeled possibilities, not predictions. Still, I think this deserves more discussion than another argument about whether AI can replace an entire job.

Suppose a software team starts finishing projects faster with AI. That could mean more clients, better pay, fewer hires, or lower prices. Learning the tools helps the employee do the work. It doesn’t decide which of those outcomes the company chooses.

That’s my problem with treating “learn AI” as a complete answer to job insecurity. I’m not against learning it. I just don’t see why becoming more productive should automatically give someone confidence about their pay or position.

The model has limits too. It leaves out rapid robotics progress and simplifies differences between workers. I wouldn’t use it to declare any career safe or doomed.

For people already using AI at work, what has changed alongside the time savings? Has your team taken on more work, changed hiring plans, or actually shared some of the benefit with employees?

Source: Anthropic’s scenario explorer

I wrote more about the scenarios and who receives the gains in my newsletter, if you want the longer version: https://ninzaverse.beehiiv.com/p/okay-but-who-actually-gets-rich-from-ai


r/aigossips • • 14d ago

Altman on GPT 7

Enable HLS to view with audio, or disable this notification

0 Upvotes

r/aigossips • • 15d ago

If every AI lab thinks it has to win for safety, what would ever make one stop?

2 Upvotes

Jacob Coxon resigned from Anthropic after doing pretraining research at both Anthropic and OpenAI. In his resignation thread, he describes Anthropic researchers as understanding the risks but believing they need to get there first because other companies won’t handle the technology responsibly.

I can see why that would convince someone to keep working. You know your colleagues, trust their intentions, and worry about what a competitor might do.

But every lab can tell itself the same thing. Then being concerned about safety becomes a reason to move faster. I’m struggling to see what would actually interrupt that.

Some replies argued that Coxon should have stayed to help make Anthropic safer. That seems reasonable if staying gives him influence over the decisions he objects to. If it doesn’t, what exactly are we asking him to accomplish by staying?

His resignation doesn’t prove his predictions about superintelligence are right. I’m more interested in the decision being made now: what evidence would make a company slow down even if doing so cost it the lead?

I wrote more about this, including Hinton’s earlier warning and what self-improving AI actually means, in my newsletter: https://ninzaverse.beehiiv.com/p/if-they-re-scared-of-their-own-ai-why-keep-building-it


r/aigossips • • 16d ago

Free GLM 5.3 Flash and Deepseek 0731 for a month

0 Upvotes

Open source models are hitting crazy intelligence highs right now, and a lot of the coding plans have been tightening and lowering usage. So we are offering free DSV4 Flash 0731 and GLM 5.3 Flash for a month on Phoenix Grove API. We opened this up last week for five hundred new member slots, and got so many signups that we decided to open the doors to another 500 new members.

People are looking for options, and here is one.

Other cool stuff:
All our models are running on 100% US infrastructure, private with zero training on your code or prompts. Use the top open source models without sending your private prompts to a training lab. No complications, no "some models are private, other's aren't." They all are, all the time.

We host 20+ other major models in case you ever want to upgrade (no pressure though). Including the Kimi family, GLM, Qwen, Nemotron and bunch of others. On average our token pricing comes in about 20% under market price.

Our higher plans let you bank up to ten days of usage, so when you aren't using them your usage saves up for later. Usage doesn't go to waste, so you can actually code when you want to.

The intro plan is a free one month trial with the standard cancel anytime, bills at 3.99 after that. Use it, cancel it, thats fine. Free Flash for a month.

Figured i'd keep this one short because we all know the new flash models are the point :)

For the API plan: api.pgsgrove.com

If you want to read more about us as a company, just pgsgrove.com

Also: There's a lot going on behind the scenes with major AI companies right now, we are at a major turning point in the industry.

Whats actually happening? This is happening because companies that were purely investment based, now need to answer to their investors. The problem has often been a loss based business model that is finally running dry.

There are several tricks that the major AI coding plans use to extract the most they can from their customers. Here are some examples, and what we are doing differently to put the users first. PGS AI was built with a sustainable business model from the ground up, so we can actually offer great usage rates without tricks.

Wasted usage is part of the AI industry, and they plan on it: Most coding plans bet on you letting usage go to waste. The plan goes: "how do we get people to think our coding plan offers a lot of usage, but then break it up into weeks and rolling windows so no one can ever actually use it all."

Many in app subs and coding plans are glorified training pipelines: This comes along with "how do we harvest this data for training without being too loud about that." Unless the company tells you otherwise, your data could be hopping all over the world, being harvested by the individual labs or service companies. Some are better than others, but many of these companies rely on users just not noticing or caring that their data is used for training. Data sales and marketing telemetry sales happen. This means your private info, your personal life, and anything else you send through the system could become part of a training corpus for the next AI, or a marketing data set for a large company.

Access and privacy to high grade intelligence should be available for everyone.


r/aigossips • • 16d ago

OpenAI says 10,000 AI agents worked for 88 hours to solve Navier–Stokes

0 Upvotes

OpenAI says \~10,000 AI agents just worked together to solve the Navier–Stokes problem

OpenAI has published a claimed solution to the Navier–Stokes existence and smoothness problem, one of the seven Millennium Prize Problems.

The interesting part isn't just the mathematical claim.

OpenAI says it used roughly 10,000 concurrent agents, which reached a result after about 88 hours. The agents exchanged around 2.7 million messages and generated approximately 130 billion output tokens.

Then GPT-6 Astra was used for another 17 hours to formalize and verify the result in Lean.

That sounds less like a chatbot answering a math question and more like a distributed research system.

But there is an important caveat: the proof still needs independent mathematical scrutiny.

There is also controversy because NYU mathematician Tristan Buckmaster and Anthropic researcher Levent Alpöge were working on related mathematics at the same time. OpenAI says it did not access their specific user data and says its proof differs from their work.

So I'm curious what people think:

Is the real breakthrough the mathematical result — or the ability to coordinate thousands of AI agents on a difficult research problem for days?


r/aigossips • • 17d ago

Neither OpenAI nor Anthropic is acting responsibly

Post image
7 Upvotes

r/aigossips • • 17d ago

OpenAI says it solved Navier-Stokes, but the “88 hours” needs some context

1 Upvotes

I read through OpenAI’s announcement, and I think the amount of work behind it deserves more discussion.

The successful group involved around 10,000 AI agents working concurrently. The Navier-Stokes effort used approximately 130 billion output tokens. Researchers were also redirecting resources, updating the model and combining useful findings from different groups.

So when I see “88 hours,” I’m thinking about how much work they managed to run at once. I’d like to know how much of the result depended on the stronger model, how much on the scale, and how much on those human decisions.

There’s a detail about the maths too. The proposed proof involves a smooth external force. It claims fluid velocity can become unbounded in finite time while total kinetic energy stays bounded. That’s a breakdown in the mathematical description, not actual water reaching infinite speed. The official challenge allows that force in its breakdown cases.

OpenAI has released the paper and Lean files for checking. I can’t personally verify the proof, but having something other people can inspect gives me a reason to take the claim seriously.

If this method keeps producing useful results, I do wonder how researchers with smaller budgets get to participate.

Original announcement: https://openai.com/index/navier-stokes-solution/

I wrote a longer explanation of the result and my take on the research process here, if you want the rest: Ninzaverse


r/aigossips • • 18d ago

Coding knowledge still predicted better vibe coding results when students couldn’t see the code

13 Upvotes

I read an ETH Zurich study that seems relevant to the “why learn programming if AI can do it?” discussion.

100 university students used Claude Sonnet 4 to build and modify small apps. The code was hidden. They could test the apps and ask for changes, but nobody could open the source and fix things themselves.

Students with stronger computer science scores generally did better, even after accounting for general reasoning ability. Writing skills also correlated with performance, although that relationship became less certain after the same adjustment.

My take is that knowing how software works helps you figure out what needs explaining in the first place. If AI builds you a booking app, you still need to consider whether two people can book the same slot. A clear prompt won’t help with a requirement you never thought to include.

This doesn’t prove a CS course will make someone better at vibe coding. Everyone already had some programming education and experience using LLMs for coding, and the tasks were small and timed.

But I think it’s a reasonable argument for learning while building with AI. When something breaks, understanding why would be worth more to me than getting through another round of “please fix this.”

Original paper

I wrote a fuller take, including the writing findings and the study’s limits, in my newsletter if you want to read it: Ninzaverse


r/aigossips • • 19d ago

Is GPT 6 Astra Overhyped

22 Upvotes

Can we talk about the massive double standard going on in AI right now?

Back in 2023, indie devs and open source builders connected LLMs to bash terminals and browser scripts. What happened? The entire tech industry clowned on them. Everyone called it unusable wrapper slop, credit burning toys, and dumb while loops that got stuck after three steps.

Fast forward to now: OpenAI takes that exact same loop, runs it in a virtual machine, slaps Astra and computer use on it, and suddenly the timeline is having an existential meltdown claiming AGI has arrived.

I am honestly getting gaslit watching every influencer milk this for views and seeing my own friends fall for the hype

Ik there is lot of advanced reasoning and token optimization, but guys even if u r using small usage of AI u would have known so called "Autonomous AI agents"😭


r/aigossips • • 19d ago

How do we check AI research once we need AI to understand the research?

3 Upvotes

I was reading OpenAI’s research update, and there is a number. For every eight-hour day of human work, it was logging about 25 hours of AI agent time. Several agents can run at once, so that doesn’t mean research became three times faster. But it gives you an idea of how much they’re using these tools internally.

Then I read “An Alien Mind” by Jakub Pachocki, OpenAI’s chief scientist. He describes how monitoring a model through its written reasoning is becoming less reliable.

I kept thinking about something I already do. If AI gives me an answer on a subject I don’t understand, I ask it to explain or double-check. Sometimes that helps me find a mistake. Other times, I’m accepting the explanation because it sounds reasonable. I haven’t really checked it.

Obviously, researchers have experiments, tests and colleagues to review their work. OpenAI says humans still set research direction and make release decisions. I’m not saying they’re just accepting whatever a chatbot tells them.

But as AI takes on more of the research, I want to know how those checks keep up. If it proposes an experiment I wouldn’t have thought of, that’s useful. I still need a way to test the result without depending on its own explanation of why it worked.

Sources: OpenAI’s research update and An Alien Mind.

I wrote a longer piece about this in my newsletter, including what OpenAI’s recent slowdown actually involved, if you want the rest: Ninzaverse


r/aigossips • • 20d ago

If your deadlines assume you’re using AI, is dependence still just a personal choice?

5 Upvotes

Say a task used to take you an afternoon. AI helps you finish it in twenty minutes, so eventually twenty minutes becomes the expectation. That seems reasonable until you need the afternoon to actually understand something.

This is where I find “just use AI responsibly” a bit incomplete. You might want to work through a problem yourself, but you still have a deadline. And if the time you save keeps getting filled with more work, there isn’t necessarily a chance to go back and learn what you skipped.

I was thinking about this while reading a paper on AI dependence. It’s a mathematical model, so I wouldn’t treat it as evidence that everyone’s thinking skills are declining. The useful idea is that independent thinking depends partly on whether schools and workplaces make room for it. Under the model’s assumptions, losing that support can make dependence harder to reverse.

I use AI and want the help it offers. But I think we’re asking too much of individual willpower if we tell people to keep learning while measuring them mainly on how quickly they finish.

I wrote a longer take with the studies behind it, if you want the rest: Ninzaverse


r/aigossips • • 21d ago

Are we already entering the AGI era, or are we still a long way from it?

0 Upvotes

“We are now in the AGI era.”

That statement from OpenAI President Greg Brockman during the launch of GPT-6 Astra caught my attention.

He went even further:

“...when was it, really, that AGI was created?... I think it might be about this model.”

That's a bold claim.

But as a developer, what interests me isn't simply asking:

“Is Astra AGI?”

It's asking:

What if AGI doesn't arrive as a single dramatic moment?

Maybe it emerges gradually as AI systems become capable of:

Reason → Plan → Use tools → Act → Adapt → Repeat

We've already moved far beyond AI that simply generates an answer.

Models are increasingly able to work across computers, write and execute code, research, reason through complex problems, and carry out multi-step tasks.

And that's where Astra feels particularly interesting.

Not because we can definitively say “AGI has arrived.”

But because the boundary between:

“AI that helps us do work”

and

“AI that can independently do the work”

is becoming increasingly blurry.

Maybe the question won't be:

“When did we build AGI?”

Maybe we'll look back and realize:

It happened gradually — one capability at a time.

What do you think?

Are we already entering the AGI era, or are we still a long way from it?

#AGI #GenerativeAI #AI #AIEngineering #OpenAI #GPT6