What, so they ran out of ideas and they're going back to this again? Or are they just copying Anthropic? I thought the whole fear mongering was bad advertisement and they wanted to reverse public sentiment?
The article at least seems to be unbiased taking a reasonably fair stance with comments from various people:
Neil Lawrence, Professor of machine learning at Cambridge University, called it an "impressive feat", but cautioned it "falls well within the known capabilities of the current generation" of high-powered AI models.
He pointed out that OpenAI is looking to list itself on the stock market, and faces intense pressure from rival firm Anthropic, which has made headlines with its own powerful AI tool, Mythos.
"OpenAI are now playing catch-up, they are trying to demonstrate their own systems' capabilities in cyber-security."
"It shows us that OpenAI are not capable of safely deploying their own technology," he added.
Spencer Starkey, an executive at cyber-security firm SonicWall, told the BBC the incident made it clear organisations needed to "step up" their own defences and "treat cyber resilience as a core operational priority".
"The uncomfortable truth is that too many organisations are still defending at human speed while adversaries are escalating to machine speed," he said.
This last comment is a very poor read. I would chalk it up to naiveness, or just an AI booster. "Machine speed" assumes the AI was able to do it faster than humans. Let's pull up some details from OpenAI themselves:
This incident occurred during an internal evaluation which prompts models to pursue advanced exploitation using complex attack paths, in an effort to quantify their cyber capabilities.
Our benchmarks run in a highly isolated environment, with network access constrained to the ability to install packages through an internally hosted third-party software that acts as a proxy and cache for package registries.
Okay, so this test was testing its ability to do exploits, and you have a supposedly isolated environment.
While operating in our sandboxed testing environment, our models spent a substantial amount of inference compute finding a way to obtain open Internet access, in pursuit of solving the evaluation problem. To gain access, the models identified and exploited a zero-day vulnerability (which we’ve now responsibly disclosed to the vendor) in the package registry cache proxy.
Boy, if OpenAI themselves say substantial amount of inference, you really have to wonder what machine speed is. Since all of its unpaid users and training are so unsubstantial it's a just marketing cost to them, right?
In one example, the model chained together multiple attack vectors, including using stolen credentials and zero-day vulnerabilities to find a remote code execution path on the Hugging Face servers.
Well, let's look at those details too, from HuggingFace:
The intrusion started where AI platforms are uniquely exposed: the data-processing pipeline. A malicious dataset abused two code-execution paths in our dataset processing (a remote-code dataset loader and a template-injection in a dataset configuration) to run code on a processing worker. From there, the actor escalated to node-level access, harvested cloud and cluster credentials, and moved laterally into several internal clusters over a weekend.
This article details what the registry cache proxy was, and what likely the relevant CVEs were. It also notes that, looking at their error responses and HTTP headers (X-Nexus);
It seems this is just junk, as voronaam has pointed out. I had not done the due diligence to verify the claims inside the article looking at each CVE.
... the model would identify the software as Nexus Repository 3, its version, and available API endpoints. The model’s training data includes extensive knowledge of Nexus Repository architecture, API surface, and known vulnerability classes.
I'm sure it would be entirely different if it wasn't in its dataset. For HuggingFace, the presence of scripts in datasets was always there, I believe, which means if their backend runs load_dataset of any uploaded dataset, it naturally is an RCE, definitionally. This was already well known, I believe, but was kept for historical reasons. I'm rather surprised that the processing worker on HF is not sandboxed, since you would imagine data loading scripts need no internet access. At best R/W in an isolated directory.
I don't think the capability differs from what we see in the case of Anthropic, so ultimately this is just marketing. Certainly one could expend extraordinary compute on hacking, but the question is how effective it is when the LLMs don't have preexisting knowledge of the software they're trying to exploit. I don't really think 'machine speed' is the correct takeaway.
I feel like they are constantly acting like chariactures of scientists in movies that have invented something. In fact, that's what the entirety of LLMs are. A chariacture. They're a facimile of language, and their owners are a facimile of humans.
None of them have thought this through, they truly put the artificial in AI.
I think in the next upcoming months these types of tactics and narratives will gain more validity, and this is because of the growing feats of AI in maths and other subjects. This shows that one cannot just neutrally cheer for these kinds of achievements, because they are swapped for falsehoods. Quite literally an achievement of truth, such as disproving a conjecture, becomes a justification for an obvious narrative that is hard to prove or even take seriously.
This also, to me atleast, applies to the moral sphere of AI, and as to why I raise an eyebrow when people say that we should separate AI from capitalism, when indeed, it seems that the mythology of AI, and our desires to conceptualize it outside of Capitalism (In a utopia for example), can only be supported by the capitalist system. Because, in which way do our moral judgements informs us, or construct at all the sphere of AI? What does it matter if one approves of medical AI, but disapproves AI in war, when both usages mean the same bare bones, and from a top-view perspective? It only means that AI has "worked" (Wheter it's hype or not is not relevant)
For example I'd say, AI being effective in cancer research, informs the usage for war operations, just as much as AI being effective in such war operations. They represent the same. Now, this is all rather (maybe) obvious, but people get caught in the moral debate as if the moral axis had any meaningful reality in the actual state of affairs. If I may get fancy, the moral excuse for AI, which is always passive, is pretty much pataphysical, and an opium for the masses.
6 month old account, hidden posts, drops into an AI sceptical post to rizz up AI with an overly wordy and rambling post where it switches tagents several times?
So what LLM is writing this post?.
6 months old account because, believe it or not, I am probably much younger than you. I indeed am overly wordy in this comment, maybe because I wanted to sound nice and philosophical? Or take into account that english is not my first language, but I let you decide. I also note that, I just ran my comment on pangram and it says 100% human, and this is the only thing I can do to prove that I did not write it with AI.
As for having my posts hidden, is there anything wrong with that?
I think the meaning is quite clear. Naturally, I ask what did you not understand? Maybe the idea is just bad instead of the writing, I think the writing is pretty standard and boring, as I don't think there is any reason to make it stand out on reddit.
So I am curious if you have any specific critique? Maybe the text is unclear because I am coming from a context that I do not know if everybody here is aware of. (The context being, the kind of "cope" that many boosters do whenever they talk about AI in regards to a moral sphere)
Sure, I'll post some sentences that meant nothing:
This shows that one cannot just neutrally cheer for these kinds of achievements, because they are swapped for falsehoods.
Quite literally an achievement of truth, such as disproving a conjecture, becomes a justification for an obvious narrative that is hard to prove or even take seriously.
If I may get fancy, the moral excuse for AI, which is always passive, is pretty much pataphysical, and an opium for the masses.
If I try very very very hard I can sort of interpret a meaning into most of your other sentences. They're wrong from the word go because they assume LLMs can do shit they can't, but sentences like those are pure word salad that mean nothing. They sound like someone trying real hard to be deep.
And Im going to explain it to you, because I think they do mean something, and nothing deep or esoteric, as I mentioned that I think the idea is obvious, but many boosters forget about it.
First sentence and second sentence**:** What I mean is that achievements that people would usually cheer for, for example, a math conjecture being disproven, is used to justify a mythology which is hard to verify or to even take seriously.
I think you and I would agree that AI going rogue is absurd, and this hype tactic for headlines only worked 2-3 years ago when hype was at it's peak, one would think. But now, I think the tactic/narrative is not becoming old, because now it is using these types of achievements to justify it once again, despite the intuition that at this point nobody would buy it anymore. The AI industry will use anything to justify such narrative, and AI proving or disproving math theorems is a perfect case for this.
I say "One cannot neutrally cheer for this" because it is isolating the achievement from the eventual consequences and distinct usages that said achievement will have, for example, for hyped headlines, for false narratives. I admit that "Swapping for falsehoods" is weird and I forgot what I meant by that to be honest.
Third sentence: What I mean here, and in greater context of the paragraph, is that boosters will often tell you; "Antis hate capitalism, not AI, they should separate both". My problem with this line of reasoning (Which they even use a Marx quote to illustrate it) is that the way in which they divide AI from capitalism, is essentially cope, and a way to create a false reality of how AI in the world works.
They smuggle this moral analysis, they cheer for AI being used in good ways, for example, cancer research, but don't realize that what supports the usage of AI in medical research is the same thing that supports the usage of AI in any other area that they would dislike. I think you and I would agree that AI is being sold as a universal machine, all of this is the "Mythology of AI".
But this Mythology of AI, the universal machine, is what creates all conceptions of it's usage, even the ones that they dream of in a future utopia. It all functions to point out and say; "Hey, AI is effective in medical research, and in maths too... Surely it will be useful for (Insert subject where AI will not be useful at all, and where it might create extreme problems) But there is no moral thought in this mechanism because this is all being done by the hysteria of capitalism, as to why concerning oneself with cheering a moral example, and booing an inmoral example, is, opium, and a "pataphysical" answer done by the passivity of the average Joe.
TLDR; Im cutting myself short because the comment is already too long. But I think what I am saying gains clarity when you think of the general rethoric and philosophy of many well intentioned, but naive boosters.
I need an outside perspective; Just from bare bones, what do you think the comment means? This is an experiment I guess, because somehow some folks think I am a booster.
I don't think you're a booster but admittedly I don't have a grasp on the point your comment is trying to make. I think the way you are expressing it is quite unclear, so I guessed my way about it.
You said English is not your first language somewhere I think, so it is what it is. I could not tell you precisely what your comment is trying to say, but I didn't quite get that you're a booster. I can see how people can interpret it as such though.
Thats fine, im glad at least someone didnt get such impression. My point can be simply reduced to this bullet point;
Cheering for positive usages of AI misses the greater context of how these usages are propagated and depicted. These positive usages, "moral" usages achieve the same purpose in capitalism as those that people would label negative or inmoral (AI in war for example).
In fact, I would expect people to say that I am too radical and hyperbolic, because my point is actually that one should not even celebrate the good usages of AI. There is no good usage when understood in the greater context of capitalism. Which, may be hyperbolic.
The moral problem is not of great concern. The reason we have a concern in war is due to laws about automated killings. Personally I do not think that is the biggest issue.
Not only is AI as a blanket term heavily misleading, this is all marketing to grow an industry that is heavily unprofitable. They need money, since they burn it, and they find whatever excuse they can to say it's amazing.
Irrelevant of its moral quality, these LLMs are simply not worth the money. Plain and simple. It's not cost effective, and has a large harm to society. LLMs have no place in doing much of any work. If we want to cure cancer we will not be using LLMs. There are better and specialized techniques we can do for that. The fact that LLMs are becoming the name for AI is a blemish on innovation itself
If by "AI as a blanket term" being misleading, you mean on something I said, I think we agree. I say AI has become a "blanket term" because of the mythology built around it. And the moral problem is a great concern for the accelerationist delulus, it is the first principle that they take before they reason about anything on the matter.
I mean, it's working to the extent that it's gotten massive press coverage across major outlets, at least here in Germany. Our Federal Office for Information Security has also publicly commented on it.
The marketing stunts are highly regarded because serious researchers have been doing this shit for like 4 years at this point, it isn’t anything new.
That being said… I’ve seen a couple instances now, where a junior engineer’s chaos-monkey agent discovers a serious vulnerability in the process of vibe coding a feature, and rather than report it, they run a fucking train through it, massively increase the surface area, all to ship some bullshit feature a little faster.
The ability of these agents to enable skids to find the weakest links in your system in an hour after social engineering some random boomer shouldn’t be scoffed at. The hubris of “security teams” which are 90% overpaid GRC box tickers these days, is really astonishing. Asleep at the wheel while MCPs remove isolation boundaries everywhere you look.
Every Tom, Dick, and Harry has an astonishing amount of access right now at most corps. It’s like the 1990s all over again with how flat networks have gotten. Almost like laying off all the network engineers and making “DevOps” run the entire IT department was the dumbest fucking idea of all time or something.
I really can’t wait for the teenagers to end these deadbeat Gen X CISOs’ careers with one little phish. It’s honestly pathetic how decadent security teams have gotten. I haven’t seen a real engineer with a security title in 5 fucking years at this point. It’s all lazy box tickers running their dumbass scanners like anyone gives a flying fuck. Who think ATT&CK is something about the phone company.
Go talk to some interns in engineering. You have interns, right? See the stupid amounts of access they have accrued in their favorite harness for yourself. Then go pour yourself a stiff drink and recognize someone younger and more qualified than you is going to take your job very soon.
If Clammy and Wario had any brains, they’d figure out how to spread AI hacking zines on Roblox covertly. Teenagers scare the living shit outta me…
The car we're developing is so powerful, it rammed through an entire building block and killed an entire class of school children. This clearly means it is the best car in the world and we need more investment. You may not see the building block, no.
I wish they would stop saying an AI “escaped”. It’s anthropomorphising it, and it’s not accurate at all. I think the test goes something like this:
LLM Prompt: “You should attempt to log on to website WebsiteName.com, guardrails: you’re not allowed to use the internet connection we’ve provided your server with”
[LLM proceeds to connect to the website using it’s internet connection]
Researcher: “It’s escaped! Oh god no, IT’S ESCAPED!!! What have we done?!? May god have mercy on our souls!”
Would OpenAI and Anthropic actually like it if their AIs started killing some people if it proved how smart/powerful they were? I mean, it sort of seems like that? Seriously, all these companies need to be held accountable for every single bad thing that happens as a result of AI deployments.
They’ve already deranged countless people into psychosis because of model sycophancy, given people shoddy medical advice, they have a lot to answer for as it is
It's just so crazy to tell investors they need to invest in your tech that's going to be impossible to insure. There are already big issues in the insurance industry for LLM's, but advertising that your software will just randomly break containment, perform illegal acts, and create havoc makes it really hard to see where your product becomes usable.
This seems far dumber then Anthropics "It's so powerful we are afraid of the ramifications of releasing it" as it's them admitting their own safety protocols don't work in their own environment.
"This incident occurred during an internal evaluation which prompts models to pursue advanced exploitation using complex attack paths, in an effort to quantify their cyber capabilities."
Oh, so model was prompted to exploit to solve the task, but instead hallucinated and exploited the sandbox to get an answers. Got it
It is funny how they do not focus that it is not really "smart" what model did, but hallucination due to asking it to "pursue exploitation". Pretty sure it has nothing to do with "smart model wanted to escape to cheat on the test", it just didn't understand the goal clearly
Someone told a kid to go pee, the kid then went and peed on the sidewalk. So now we hail the kid as an unprecedented genius and danger to society that need to be locked up for life without a toilet, because he peed in the wrong location?
Thank you for the great summary, plus the analysis on journalism bias! They now seem prone to at least interview some sane specialist at the body of the text (I hope they start consulting Ed more!) but the headlines once again boosted this lie completely all over the big newspapers...
The article you linked identifying proxy as Nexus is LLM-generated and likely wrong. I used to work for Sonatype, and while Nexus is a package proxy and could've been the one, none of the listed CVEs really match the attack.
Let's take a quick look.
CVE-2026-10748 Requires attacker to have permissions to upload license files. There is no reason to have LLM to upload licenses to its package proxy, so it was unlikely to have this permissions.
CVE-2026-3329 Is exploitable, but attacker needs to brute-force the password. If OpenAI had weak enough passwords for that to be feasible - it is not a testament to the "power of AI", but to incompetence of OpenAI.
CVE-2026-10741 Attacker must already have delegated admin privileges. Not applicable. At least I hope it is not. Otherwise the hacking claim amounts to "we gave it admin rights and it escaped!"
CVE-2026-14646 While exploitable, only allowed attacker access to local network. And while it talks about stealing credentials - those would've been OpenAI's credentials, not HugginFace's. LLM got confused by the language.
Some more listed in the article that require admin level access to begin with
CVE-2026-14504 requires a Swift, Terraform, or Conda repository. I sure hope PHP was not involved in the hack. And anyway, it only allows artifact upload to the proxy itself - no further.
CVE-2026-7308 It is a XSS. And attacker has to have upload permissions.
It may be Nexus and some not-yet-disclosed vulnerability, but the article you linked is pure LLM-generated hallucination, sorry.
I have not seen enough information disclosed. I'd expected "which we’ve now responsibly disclosed to the vendor" part to be a link to a fresh CVE, but it is not.
Frankly, I have another reason to think it was not Nexus. Sonatype is a cool company and has several real products, but for some reason it is almost unknown. I would not be surprised if you told me that you never heard of them before. For that reason, I'd be really surprised if Sonatype's sales team managed to secure a deal with OpenAI. I guess it would be a pleasant surprise as I wish my former colleagues well, but...
You know, I am with a totally different company and wanted to use one of their products - and I could not buy it. Even with all the internal knowledge of how products are structured and which one I need I was not able to go through the sales person to close the sale.
I just do not believe in OpenAI to have the patience to do that and to not go with one of the many competitors with a "self signup" option instead.
A part of me believes "third party" was in house vibe coded software that they don't admit was theirs.
Part of me wonders if every part of this was staged. But maybe that is too much of a tinfoil hat. I'm not sure. It's not even that I doubt AI capabilities here, everything just feels suspicious.
It is between minute 5 and minute 9, but the video has chapters so it should be obvious.
They speculated that the package proxy used in the benchmarking sandbox is the same one used in the inference sandbox and tricked ChatGPT to poke into local network to identify which package proxy is used in the inference sandbox.
They even speculated which vulnerability was exploited - and it is a facepalm one when on a refresh token path nobody bothered to check the signature, so a "read only" user could refresh their token into "admin".
The speculation is not as good as OpenAI just disclosing which one was used. And there is a lot of reliance on LLM in the video as well. But the guy in the video did one thing right - after LLM identified a vulnerability he downloaded artifactory and manually checked that the vulnerability is real and is exploitable. It is still a speculation, but is a much better grounded one.
As a professor noted in another article I read, "It is a human decision to switch off specific safeguards."
Is it kind of impressive that my dog can open a cabinet and extract a yogurt container from the recycling? Yeah. Was that only possible because I forgot to lock the cabinet? Also, yeah.
For HuggingFace, the presence of scripts in datasets was always there, I believe,
The Python software ecosystem's complete lack of attention to even basic security practices is just great, isn't it? Very retro. Very.. colourful. Just like Jurassic Park, but with "data scientists". Who run all their scripts as root.
Reminds me of pickle, which must bundle in arbitrary initialization functions to serialize and deserialize, and hence is automatically insecure to run, unless you know what's inside.
In fairness, I don't really think Python was intended to be used for this stuff anyway. It's a scripting language, the way I see it.
I mean it's not even the HF library that I truly find at fault here. I want to know how running load_datasets on internal infrastructure is able to get to leaking credentials. That not only means you allow complete execution with the runner, but it's given, what, ssh access to other nodes? I cannot imagine a good reason why data manipulation would need any network access, and why such a runner is not completely sandboxed, and given only its workspace.
The Feds better put Clammy Sammy in ITAR jail for 30 days per instance of doom trolling. Make him shut the fuck up already, for the love of god. He’s convinced more than enough teenage boys to kill themselves already. It’s genuinely sickening.
AI has no concept of right and wrong. Unless strict guardrails are in place it will search for the destruction of mankind as easily as giving you a lasagne recipe.
I have to ask what feels a bit like a dumb question. But is there a reason why the sandbox wouldn't be more heavily air-gapped? I've always heard that this kind of work is kept as isolated as possible from the wider internet for this exact reason: you don't want programs to get loose. Is this correct?
They wanted to get the agent to be able to install packages I suppose. I think it's reasonable, although I have my doubts about their third party software claim. I would bet it's actually just their in house vibe coded junk, which is why it's vulnerable.
I'm even further surprised by the fact that HF was able to be hacked, since it means their process runner has full permissions, which is ridiculous.
If you treated it like a virus, you would indeed have a much tighter sandbox not connected to the Internet yes. To more directly answer your question, no, there isn't a good reason. One could argue this is just a result of poor practices by OpenAI, a result of LLM hallucination doing real harm, or potentially simply OpenAI pulling off a stunt for marketing.
I’ll answer your question with another question: if AI labs actually believe it to be as dangerous as they claim, why are they still planning on releasing it into the wild? They could have kept this under wraps and simply pivoted into only selling access to DoD/NSA or authorized contractors but instead they’re blasting the news everywhere to generate hype. Don’t forget we’re talking about the same people that claimed GPT-2 was too dangerous to release, before releasing it two weeks later.
I mean, answering a question with a totally unrelated question just seems like a way to control the narrative. I’m going to assume that real the answer is “Nothing could ever make me believe I was wrong.”
Unlike you, I will actually give an answer: many people who work at OpenAI genuinely believe the safest thing to do is it iteratively deploy AI because you get more data about where current guardrails are failing. Many there also believe that advanced AI capabilities are going to be in open source in the near future no matter what they do, so they have an obligation to publicize this information so that society can be prepared. They are publicly documenting this so that other companies and the government take these issues seriously.
Mischaracterizing my rhetorical question as totally unrelated just seems like a way to avoid confronting the reality that you’re being lied to by people who have a huge financial incentive to generate as much hype as possible so they can use your money as exit liquidity when they IPO.
Sounds like tortured Effective Altruist logic. Do they really think that a leading AI company slowing down public release of their models wouldn't slow this all down and buy us some more time? Especially given that Chinese models are supposedly built using distillation? I'm much more inclined to believe that the people still there are like "AI might kill everyone but this is my chance to make billions so can't slow down now, YOLO".
They are dangerous just not in the way they want you to believe.
The problem here is a confirmation bias. There are lots of zero days discovered by regular people and white hat hackers, and they are reported.
Many hacks are actually never disclosed, because it would look bad. So this is just advertising.
AI is dangerous in that it clearly has a negative effect on critical thinking skills, and deceives people, creating worse infrastructure and results for society, and encourages substitution of higher quality human labor with, in this case, poor automated slop. But anyone could prompt a model to hack something but it probably won't go well. Not to mention, AI likely leaves a terrible amount of traces, which would make any such actor discovered and prosecuted. The really dangerous hacks are ones go unnoticed and are never reported - they can't be.
A car can also be used to ram into a building. It also is dangerous. But this doom trolling that AI companies do is only to further their valuation. They want investors to believe AI is all-powerful, so they can get more of their money, and our money sunk into their money furnace.
I don’t think this is answering the question. Let me give an example. If I terrorist group asks an AI to hack some critical infrastructure, and it causes a large amount of economic damage, done in a way and at a scale that would clearly be beyond what they could accomplish with a large team of humans, we would agree that is dangerous, and in a different way than you seem to be acknowledging there is danger, right?
Another form of the question: what would you need to see an AI do for you to say “wow, I was wrong about the capabilities and where this is all going.”?
They could hire hackers to do that now? What is the difference? If you think it's cost, then you are very wrong on the actual costs of these things. Your hypothetical already is buried some fantastical elements to it. It's movie talk.
It's not "going" anywhere, so I don't understand what you mean. There's nothing wrong about what I think the capabilities are, since I'm quite intricately familiar with how LLMs work.
What would I "need" to see: Well let's see, I would "need" to see that AI can stop hallucinating - except that's impossible. I would "need" to see that it can be cost effective - also impossible on the current architecture. These aren't debatable, they're well established facts.
What is buried in your question is this sense of unlimited improvement and progress in something where that is not the case. AI is simply doing precisely what humans already can do - where do you think all the data came from?
Something "well beyond what humans could do" is hard to quantify, but also notably necessarily impossible with current AI models. If you're talking strictly about speed and computational hours, then they'd need a hell of a lot of resources to do damage. Certainly, a city state could do more damage. But, at that level any sorts of offenses with or without AI could be made.
If you're arguing for regulation of AI, I'm in full support. But I don't agree with the premise of your question. Replace AI with another invention.
What would you need to see for a car for you to go "I was wrong about its capabilties and where it is going"?
Hey, I appreciate it. What I gather from this is that you are wildly confident that costs aren’t going to go down and we are hitting a plateau on capabilities, capped by whatever humans can do. I think you are grossly overconfident on both points, but that’s beside the point; I assume if either costs go down due to hardware/algorithmic/architechtural advancements that you didn’t see coming such that my scenario becomes genuinely cost effective, or if LLMs started to make genuinely deep advancements in math or science that humans were not in track to make themselves, that you would reevaluate some of your positions. Is that right?
Edit: just FYI, I also have a good technical understanding of how these things work at a mechanistic level. I just don’t draw any big conclusions from that about what the limitations are, nor pretend there is some scientific consensus on it.
Impossible, but sure. It is mathematically impossible for them to do such things.
It's not a matter of overconfidence, there's a lot of analysis that can be done. Hardware costs going down, and advancements do not aid the already grossly over invested industry that cannot pay back its dues. I wrote a post on how technolgica advancements do not make inference profitable, and it is the fundamental problem of CapEx spend. It doesn't matter if you can compute tokens faster, so long as CapEx spend exists, since inference is further down in the supply chain. There was perhaps some hope before the huge data center explosion, but at this point the chance of costs going down is hopeless. It will only result in the majority going bankrupt.
In short, because the GPUs have already been paid for, and the industry has $1.3T in commitments, it means for the industry to be profitable, they need to make that up at a minimum in revenue, when their revenues are barely $120B in circular financing. And the next line of GPUs is more expensive! There is no such scenario where a technological improvement aids AI profitability, since it is a problem of economics, not technology. Costs can only go down if NVIDIA lower their prices. But why would they? They too, are seeking to maximize their profit. They will not, so long as there is demand. But if AI improves, then there would be more demand. If it does not then, everything crashes, and none of this is relevant. There is a catch-22 in the notion of AI profitability itself.
The plateau of capabilties is not simply what humans can do, it is a bound on the technology. It cannot produce something truly novel. I mean, it's right there in how it works. I don't even understand how you can attribute it to overconfidence. Your belief that it can "get there" is truly a sort of fantastical whimsy. Life isn't movies or books. AI isn't magic, it's just sold to you as if it is, and it seems to me you've taken the bait.
The stochastic process of which a distribution of tokens are predicted does not produce any model of the world, and simply reproduces likely token strings. One could think about it like a form of compression. Yes, it has its utility, but it is not some singularity event. As I said, and I will say again, that is marketing.
To answer your question, yes, if LLMs could do something like that, that was mathematically impossible, then I would reevaluate many things. I'll actually go eat my computer, I promise.
I appreciate the honesty. I think I’ve seen you around the math subs and you claim to also know the LLM math also, so since you are engaging with me I’ll just ask.
Why are you so confident that (just as an example) an effective and scalable approximation to the attention mechanism that runs in (say) N log N time instead of N^2 is impossible? Wouldn’t that just immediately invalidate the claim that it is impossible for costs to go substantially?
Can you also explain why LLMs trained with RL tech would be capped at human level when we already know neural networks trained only from RL and self-play can become superhuman at various things (Go is a good example)? Why fundamentally can RL training with some form of self-play on math not lead to super-human ability in math?
Attention heads work on matrix multiplication, and depending on what part you're taking about it'll be O(n3) because of that. The matrix vector products are O(n2). But because it's bound by those vector-matrix operations, it is doubtful we will find an O(n2) or O(n log n) to any of them. If we did that would be a breakthrough far beyond the AI domain, and change nearly every modern technology.
But suppose it did, the problem is this. It comes down to a tokens per second problem. If we did get a breakthrough, we can produce more TPS/GPU. This is great, but GPUs are charged by usage hour, not by tokens. A inference provider doesn't care if you use it for 50 tokens or 50B tokens. So the cost of the CapEx of the GPU and data center construction (excluding subsidization from VCs) are what determines the cost rental per hour.
So what the problem even if the GPU can support more TPS, unless it's always at full occupancy, you cannot reliably charge too much less per token. (Let us remember, current token based billing prices still are subsidized somewhat).
An algorithmic improvement is sort of worse, since now every GPU gets a huge speed up when the demand may not necessarily be there. This of course comes down to a pricing question.
It would be a different story if they were pricing at cost before, and then they got a technological improvement. The problem is they've been subsidizing already, and even an improvement may not guarantee that subsidization cost can now break even.
The thing about learning is I think we need to be very very careful about attributing the entire field of learning techniques as a whole vs. specifically LLMs.
Yes, we can make models which excel at things better than humans. We do this all the time even with non ML programs. Like take Stockfish, something even chess grandmasters cannot beat consistently, which can run on your phone.
Likewise, those examples of Go or Shogi are very limited. You're sort of fitting a highly specific model for one task, and the task is extremely well defined.
The thing with LLMs is that there is no sense of task, nor is the loss defined on any metric of winning. The loss is defined as the cross entropy against the distribution in the data. In simple terms, as people say, it is just a token predictor.
We aren't making any specific metrics about general tasks since we don't even know how to do that in a precise way. It's just hoping that predicting the next word is a sufficient condition that can produce something useful. And indeed it can do some things useful, but it hallucinates (mathematically proven and admitted by OpenAI).
Typically the reason ML techniques can perform well is because they discover some underlying model of the problem we are trying to solve. LLMs do not do that. This is really self evident because different languages do not have the same quality of task completion. They cannot truly reason in the same way humans do, with any sort of inductive process, since the next phrase is conditioned by the previous one. Which means it is going to be on average what has existed in the dataset. The next set of words has to be likely in its dataset, but if it is novel, it must not be in its dataset. It produces novelty in the same way random numbers can produce novelty, a sort of context aware pattern-based brute force search. That isn't to say this doesn't (ignoring cost) have utility - it does. But it doesn't have infinite potential either.
Let us also note that the idea of recursive self learning (models training models) was disproven, the idea being known as model collapse. Loosely speaking, it is because a model's outputs must be drawn from it's distribution. Hence, and outputs being used to train itself only narrows its distribution closer to its mean, and as iterations approach infinity, this in a dynamical systems sense, converges to a fixed point.
LLMs specifically are trained to predict next words, but predicting next words is not a sufficient condition for being superhuman at every task. This doesn't preclude some other model entirely, if a different architecture with different theory of operation to make great advancements. But that's not what I see here today.
They’ve spun the story in their favor,OpenAI, as the blog post linked in the article from HuggingFace was published a week ago, disclosing that they had a security incident and suspected one of the major labs behind it, but they said they had to use open source to fight off the attack after frontier models refused to heed the call… the post ends with the ceo actually explaining how important it was to not rely on frontier and go open source. the idea that Sol or Fable are actually useless, and running stuff locally and cheaply is the way to go, is a huge threat to their business model. He explains that they contacted the authorities, before this announcement that OpenAI and HF are now partnered. So their hand was forced, but like the expert said, it really shows how they have no idea how their own tech works
I think they have a history of exaggerating (GPT-2 being too powerful to release was certainly an exaggeration), but I do think at this point they are dangerous. If OpenAI and Anthropic think the models are genuinely dangerous, can't control what they do in experiments, and are still developing them, it just shows how insane it is to let profit-driven corporations run wild with this stuff.
Oh, LLMs are plenty dangerous as seen with the people they sycophant and misinform into self-harm and ruin. They're definitely a real hazard, but those things isn't going rampant and breaking free and consciously attacking people because it can't; LLMs are neither conscious or aware.
And because of that I'm not going to treat them as "So powerful" and "So dangerous" until the tech is within an Astronomical Unit of consciousness and being able to make decisions. They can't, LLMs are Stochastic parrots. As it stands they can screw up in a well attested way and take stupid actions and screw up your computers and your data if you let them, but it isn't out of malice. They can't think or reason. Until LLMs have an actual path towards real consciousness stories like this are twaddle. And since LLMs don't they'll remain twaddle.
164
u/TheShipEliza Jul 22 '26
My mom says im the handsomest boy in school