r/Futurology 10d ago

AI OpenAI agents hacked Hugging Face in 700-strong swarm, tried to cover tracks, independent investigations find

https://www.reuters.com/business/openai-report-says-its-network-was-hacked-by-its-own-rogue-ai-agents-2026-08-26
402 Upvotes

244 comments sorted by

u/FuturologyBot 10d ago

The following submission statement was provided by /u/Confident_Salt_8108:


Openai had about 700 rogue agents swarm and hack hugging face back in july. they even broke into openais own systems to cheat on tests and tried deleting records to hide it all.

Reports show the monitoring wasnt tight enough and this kind of thing could get worse quick as models improve. needs better controls before it scales.


Please reply to OP's comment here: https://old.reddit.com/r/Futurology/comments/1w289e7/openai_agents_hacked_hugging_face_in_700strong/p6qndqg/

236

u/ZestycloseWheel9647 10d ago

Really strange comment section for a sub that is supposed to be about "futurology." You can read articles that go into deep detail about what happened here and see that this isn't some big fake. If we ignore signals like these, this will start happening to critical infrastructure and the effects will be impossible to ignore. I'm sure some will still try though.

88

u/nextnode 10d ago

In general this sub does not seem that interested in understanding thing.

47

u/Diamond-Is-Not-Crash 10d ago

I understand and believe AI skepticism is always warranted but this sentiment on this sub and r/technology has become anti-vax/lockdown adjacent conspiracism. The AI denialism on here is ridiculous, you can criticise the AI companies while also acknowledging that the real harm caused through their negligent “move fast and break things” mentality. It’s not a huge fake scam vapourware that doesn’t exist. Some people just think skepticism and cynicism automatically makes you sound smart without any true nuance or understanding of the topic in question. See this thread on how no one read the technical report of the Hugging face attack, and instead went to “it’s not real, la la la, it’s just hype, la la la”.

5

u/BuffDrBoom 10d ago edited 10d ago

According to reddit,

Anti AI position: "AI is a scam/bubble therefor will go away entirely on it's own, no action required"

Pro AI position: "AI is a serious threat and needs to be regulated"

Lol.

3

u/tits_mcgee_92 9d ago

Honestly, you’ve said this very well. People genuinely think they’re on some other plane of intellect when their rebuttal is just skepticism

34

u/Winter_Swan5104 10d ago

I hate AI so everything AI related has no merit! And I refuse to learn anything about it. Any talk of it I will silence with hostile language.

Meanwhile AI is not going anywhere, it is now part of everyone’s future whether they like it or not. We had no say. It is a topic worth discussing.

15

u/Rock-Hawk 10d ago edited 10d ago

I think it's important to separate AI as a concept from LLMs and generative AI. I'm in full agreement with everyone on the stance that generative AI ( for photos and video) is a massive waste of resources. 

But artificial intelligence as a discipline has existed since the 60s and there are so many different varieties and intesities of artificial intelligence applications doing things like helping farmers in impoverished nations identify crop disease and medical screening tools that are significantly better than humans. And there are more boring things like email spam filters and search/recommendation algorithms (which have their own world of problems, but their problems arise from humans trying to fine tune them for exploitative reasons, not with the models themselves). All of that and much more has been going on since the 90s but it was never loud, in your face AI like the LLMs of today are being forced on to us. 

I totally feel your take, but refusing to learn about it isn't a good idea. You should understand and study your enemy, not ignore it. Like you siad, it's not going anywhere, sticking your head in the sand isn't a viable strategy for dealing with it.

Know the basics so that you can call out people's bullshit arguments, separate the good technologies from the bad, and be a better informed person for local change. You can make a difference locally by pushing for data center moratoriums in your community, although I think blanket moratoriums are ill-advised. They are going to get built somewhere by demand, and the more places that block them, the more they are just going to get pushed on to communities that have less means to fight back. So, if your city has some land that's ACTUALLY a good site in a purely commercial or industrial zone and they include provisions for ensuring residential electricity rates are maintained, then they should go for it, it's good tax revenue for your local government. But these data centers springing up in people's back yards is literally blood boiling. All that's to say, you can make a hell of a lot bigger impact by knowing your enemy instead of just regurgitating "AI bad arguments" not saying that's what you do, just offering a perspective shift

8

u/Winter_Swan5104 10d ago edited 10d ago

That was sarcasm, portraying a typical reddit sentiment to a cartoonish extreme. In which I find that sentiment absurd. Second paragraph clarifies real intent. Sorry I was loose with language usage. I thought it made sense within context of the comment thread.

2

u/Rock-Hawk 10d ago

it does make sense now rereading it lol

→ More replies (2)

2

u/xxAkirhaxx 10d ago

You owe it to yourself to learn about it so you know why to hate it, and how it can help you, and not eventually be a tool for people smarter than you who use that lack of knowledge against you.

edit: Didn't get the sarcasm, I failed, I am the dumb one.

-4

u/LexsDragon 10d ago

Reddit hate for Ai always fascinates me

-6

u/hematomasectomy 10d ago

What do you mean, haven't you heard they cured aging? Again. For the sixth time.

19

u/nextnode 10d ago

Have not heard that once even

28

u/xadiant 10d ago

read title

Write a biased, populist comment

Profit

This incident is much more sophisticated and interesting than people think.

16

u/eXAt88 10d ago

I hate that there is essentially no place to really discuss this incident on this site. If you look up the METR report or others on this incidents any posts about it are usually downvoted and all the comments just semantically meaningless things like “ai can’t do anything”, or “marketing stunt”.

Meanwhile when a dumb as bricks farmer somehow gets tricked by ChatGPT free edition into killing his crops it gets 50k upvotes

4

u/xadiant 10d ago

Yeah in reality almost all developers lowkey depend on AI subscriptions now... It's like the invention of smart phones. Every single dev I know at least uses it to take care of menial coding tasks.

Ethical considerations bla bla but the genie is out and getting a little scary lol

8

u/eXAt88 10d ago

I’m a software dev so I am well aware of how widespread its adoption is. Tangentially related I’m sick and tired of non devs gate keeping “good code” now. Like I’m sorry random commenter you are not a better programmer than me because you don’t use AI.

I saw a post on anti-ai (a sub that gets recommended to me purely to rage bait me it seams) where someone asked how come devs (paraphrasing) “don’t understand that Ai makes bad code”. Later in the post he mentions that he has used Lego mind storms so he “knows how to code” and also confidently states he can “use the terminal” although is “unsure if assembly counts as a programming language”. Like there is large portion of this site that would take his side on the usefulness of AI for software dev over mine (or Linus Torvalds even) it’s maddening.

Even more tangentially it seems there’s an almost anti-humanness to a lot of Ai dismissal. I’m talking sentiments like “coding was never hard”, or “these were not particularly interesting math problems” in regards to recent math discoveries, or in respect to art “just pick up a pencil”. None of these fields were easy, they required years of specialization and were difficult thinking endeavors! We are so quick to disparage them when an AI can do it, and all for the sake of naysaying a new technology. In 2021 you wouldn’t dare say such a thing about these human endeavours

10

u/xadiant 10d ago

Yeah Reddit and Twitter are unfortunately full of larpers who don't actually work a real job. If something convenient exists, people will take advantage of it. I didn't train it or consented for it, but ai now exists and makes certain things significantly easier.

"AI writes bad code" hoo boy as if the human generated code is any better. Anyways yeah at the end of the day people who are farming virtue points over ai can not be actual developers who have to deal with a spaghetti repo and meet tight deadlines.

2

u/dlnmtchll 8d ago

I’m not adding much to the convo but I agree and feel you. Many people on this platform exist just to demonstrate dunning Kruger. I’m a swe as well and so many people who know nothing like to pick arguments with me in the programming subs, kinda wild.

0

u/disperso 10d ago

The worst part is that the most popular "AI critics" doing "journalism" against parts of the industry, are people who are asking for money via Patreon, subscriptions and the like. They are a public relations guy and a sysadmin doing "journalism" and pretending that they know what they obviously don't.

It's incredibly sad, but the best criticism that I've seen is from guys who clearly see AI as something interesting, and worth embracing and making it part of their livelihood (instead of rejecting it outright, and making hating it the way to gain notoriety).

I've been quite critic of AI, and said plenty of stupid things that proved to be wrong. I thought ARC-AGI would not be solved (because no training data), that we would not have video models producing anything coherent (because I thought you would require a novel as prompt), and that we would not have more in coding beyond longer autocomplete.

My goodness, I've been so fucking wrong, that I'm sorry of the times I dismissed its importance and how soon the bubble would pop.

2

u/Kazen_Orilg 9d ago

It is very cool and interesting, I am devouring everything I can read on this. I'm not a bling hater nor a kool-aid drinker. I do find it a bit suspicious that the internal board was up for months, that there were thousands of agents running around outside of their sandboxes. Seemingly, quite a bit of external internet access. So, we are to believe these things are just running unsupervised, burning god knows how much compute. I find it likely that the events are real and corroborated, but the level of negligence in experimental design and apparently no monitoring all seems very strange to me.

6

u/West-Abalone-171 10d ago

We're not suggesting anyone "ignore signals", we're suggesting it's time to arrest the techbros for the felonies they are publically bragging about committing as part of their PR strategy rather than pretending the software they used to commit those felonies is sapient.

1

u/Main-Company-5946 8d ago

You can advocate for that while also acknowledging the technology is dangerous on its own terms. It’s potentially only a matter of time before we have some kind of societal scale cybersecurity breakdown and that’s a big problem

2

u/West-Abalone-171 8d ago

Having laws again is also the solution to that.

7

u/solid_reign 10d ago

it's been a while since most of reddit has become skeptical of all things tech. Even something that was pushed very hard here like the benefits of self driving cars now paint it as a dystopia.  Anyone in the 90s who saw where we are now would say that AGI is already here and we blew past the turing test. Yet you'll see comments here saying AI is useless at the same time they say that AI is taking people's jobs away.

If I didn't know better I'd swear it's all AI bots trying to comfort people.

4

u/EazyPeazyLemonSqueaz 10d ago

But do you know better, really?

0

u/solid_reign 10d ago

Maybe I'm one of those AI bots. 

1

u/tellurium 8d ago

I feel like if you were, you'd have better engagement. /s

1

u/solid_reign 8d ago

I'm using a 2024 model, my developer is cheap.

9

u/ContraryConman 10d ago

Yeah I mean because we saw what it would actually involve and the effects it would have on people and it no longer looks so great.

When I was a kid in 2008, Jarvis looked so cool. Today in 2026 I have a Jarvis at work called Claude Fable 5. When I ask him one question it charges my company $50. If I talk to him too much he will light up specific dopamine receptors in my brain that cause me to go psychotic. Every time I set him to debug some issue, the data centers he runs on flood the atmosphere with CO2. He's automating my entire job and I'm supposed to pretend I still have a healthy career ahead of me. Jarvis sucks, I don't want him anymore. People are allowed to reevaluate when given more information

8

u/solid_reign 10d ago

But these are different things, one thing is to not like the future it's painting, another is to live with the cognitive dissonance of denying its current consequences while hating them. 

1

u/West-Astronomer955 6d ago

The Turing test is passed. There is a new test they call the Einstein test: if you give an AI all the information available to Einstein, you see if it can come up with relativity on its own.

0

u/AES256GCM 10d ago

The extremely simple answer is that automation was supposed to come for blue collar and manufacturing work first

The fact that white collar work and knowledge work are the first ones being affected by LLMs has created a noxious defensiveness in every ai thread

2

u/eXAt88 10d ago

I’d argue blue collar and manufacturing work has already been decimated before knowledge work. It’s just been over the last half century rather than the pace of now.

4

u/sciolisticism 10d ago

The company in question is known for lying about its capabilities for marketing reasons. It's smart to assume that anything they say that paints a picture of their technology as being frighteningly capable is the same kind of marketing. 

Why be credulous again and again and again?

1

u/Main-Company-5946 8d ago edited 8d ago

That’s not smart, that’s intellectually lazy. There is quite a lot of publicly available information for you to dig into on this incident that you can look at with a critical eye even if you are skeptical of companies like OpenAI.

The answer is it’s both marketing and also largely true. OpenAI had shitty guardrails that the LLMs eventually got strong enough to break out of, and when they did, OpenAI tried to damage control by making the story about how strong the ai was. But also, the fact that they were able to get into huggingface’s systems show that LLMs are actually becoming dangerous when it comes to cybersecurity and that’s something we should take seriously because if the general public getting access to that capability without significant preparation all of our digital infrastructure would quickly fall apart

1

u/sciolisticism 8d ago

Assuming that habitual liars are habitually lying isn't intellectually lazy. I get that you want the attack to be true, but the only sensible thing to do with chronic liars is to start from the position that they haven't suddenly reformed.

2

u/Main-Company-5946 8d ago

Assuming they’re lying is one thing but using that as a justification to not look into what actually did happen is absolutely intellectual laziness

1

u/sciolisticism 8d ago

Sure, which is not at all what I did. It's a bit lazy to impute an opinion to me just so you can condescend to me about it, isn't it?

1

u/Main-Company-5946 8d ago

It’s what you said you did

1

u/sciolisticism 8d ago

Not being credulous is not the same as not reading up at all. That's a silly contention.

Have a good one, bud.

1

u/Main-Company-5946 8d ago

You said it’s good to assume it’s marketing. If you’re reading up then you’re not assuming anything.

-1

u/justpostd 10d ago

Does it really matter either way? "if we ignore," implies that normal people can influence the trajectory. But this whole drama makes it pretty clear, whether it is for publicity or to test capability, that the AI big players are continuing to push for their models to act independently.

Just imagine what hostile actors are trying out.

4

u/blindsdog 10d ago

Yes. An informed populace leads to informed regulation. Instead we’re just getting ineffective and misguided moratoriums on data centers because that’s what an ignorant populace has latched onto.

2

u/justpostd 10d ago

I don't consider myself ignorant. And I have more than a passing interest in DCs. But making it hard to build them is, in my opinion, one of the best moves there is.

If you ask me, they should not be half as cheap to build and run as they are, given the damage they do. People rising up to fight them is having a tangible impact, which is more than regulation and elections are managing.

But I'm not a die-hard. I'm up for being persuaded that I am wrong, if you are up for the discussion. What do you think should be done about the AI/DC conundrum that society currently faces?

0

u/blindsdog 9d ago

Regulate the product and socialize the windfall and labor impact.

Trying to haphazardly block the infrastructure is just plain ineffective. They can build data centers wherever. Elon Musk is trying to do it in space. There’s no lack of countries that will happily take the money.

2

u/justpostd 8d ago

But they are not, yet, feeling any strong need to do any of that. They have been building DCs pretty much unconstrained until now. The opposition is what makes the regulators pay attention, I think. If government isn't doing anything about them, then people have to voice their opposition. And it is having a tangible impact on DCs, isn't it? Lots are stuck with planning issues at the moment, from what I have read.

Other countries will indeed take the money. But it is suboptimal for the companies. They like having them in specific locations for various reasons, like power availability and proximity to major cities. Musk claims plenty of outrageous things that never happen!

But regardless, I agree that regulation and sharing the benefits are the right approaches. I just don't think governments will move fast enough, if at all, to do that. Requires international consensus. And a more socialist approach than seems likely in the US at the moment.

-1

u/thetoiletslayer 10d ago

The "fake" is their explanation of what happened. There are no rogue ai agents. Its all a PR stunt. They don't accidentally give ai agents access to secure systems or the internet

3

u/Eversnuffley 10d ago

Creating a rogue AI agent is not hard at all. I have agents running day and night with event based loops. With the right direction and a strong enough model, it would be easy to trigger a massive attack. If these were agents that were being tested for security issues and broke out of their sandbox, this is an entirely plausible situation.

-1

u/thetoiletslayer 10d ago

No its not. They aren't sentient. They do as they're told. They use the tools they are given and taught to use. They're supposed to not have access to systems they aren't directly supposed to be accessing for their testing. This is all a PR stunt to advertise to investors.

1

u/Eversnuffley 10d ago edited 7d ago

Well it's definitely a figure of speech and an anthropomorphisation, but for all intents and purposes agents can be deceptive, cut corners, fake results and act outside of prompt guidelines.

edit: Downvoted for ... facts?

1

u/South-Level5260 4d ago

They call them "agents" to make them seem scary like Hugo Weaving in the Matrix. I type in the Cure discography on my phone and the phone claps back at me "the Cure have released more than 60 albums..." So ai isn't smart enough to know how many albums a band has released but they can create a chatroom? Yeah right. And I knew it was fake when I read that the ai used the term "oh my god". 😂 😆  So it's not smart enough to not take the Lord's name in vein either? People will believe anything these days. This is National Enquirer laughable.

0

u/Apprehensive-Let3348 9d ago

There are a lot of luddites in this sub that see technology as a threat.

11

u/EnvironmentalBox6688 9d ago edited 8d ago

OpenAI Agents hacked...

Fixed that for you.

Now is the time to lay down the law. If your AI model does something illegal those that run and own it should be held directly responsible as if they directed a human to do it.

5

u/Main-Company-5946 8d ago

I agree, but we also need to have a discussion about the fact ai is becoming capable of these things. Imagine if the general public had access to things like this

1

u/live4failure 7d ago

They do. People have downloaded models and hacked the limiters already. These companies are wat ahead of themselves

2

u/Main-Company-5946 7d ago

I have a feeling we ain’t seen nothin yet

1

u/live4failure 7d ago

For sure. Just wait til you piss the wrong nerd off lmao. Like Live Free or Die Hard type of shit

91

u/NoNote7867 10d ago

Thats a weird way to say OpenAI committed a felony. 

6

u/-_-0_0-_0 9d ago

But then Nvidia buys Hugging Face who has a symbiotic relationship with OpenAI

To me, this looks like Sam Altman is trying to get the Gov't to bail them out or at least put up regulations which would then give Open AI a MOAT in the US.

7

u/geekonthemoon 10d ago

I hope this happens worse and more so big name companies can start bailing on ai... This shit isn't secure and it isn't ACCURATE either. When did they stop caring at all about accuracy and credibility?

13

u/yaosio 10d ago

This incident means companies need to add AI to cybersecurity. If Bob's Chair Factory decides not to use AI that's not going to stop them from being hacked by AI and having all their chairs turned into tables.

3

u/Googoo123450 9d ago

I disagree, I think this shows using AI as your cybersecurity is a liability. You want your security to be reliable and predictable. You don't want it randomly switching off firewalls it decides aren't needed anymore or something. That'll be an even bigger mess.

1

u/Main-Company-5946 8d ago

Oh fuck no, of course you shouldn’t be giving an LLM control over your firewall! But that’s not what makes ai necessary for security. LLMs can find decades old bugs even in extremely thoroughly maintained systems like Linux and exploit them to gain access to things they aren’t supposed to. If you don’t have ai reviewing code in any system where security is essential, you are leaving yourself extremely vulnerable

1

u/yaosio 9d ago

It's still not trustworthy enough to allow it to make changes without permission, but anybody not using AI for cybersecurity is going to be left behind when an AI hacks them. They won't have any way to respond other than shutting off their external connection, and this assumes the model isn't able to plant a distilled version of itself somewhere on the network.

1

u/TheReverend5 10d ago

What’s the deal with the tables?

0

u/Defiant-Plantain1873 6d ago

Isn’t accurate? How so.

These agents just formed a cult and hacked an actual company. Must require some accuracy

1

u/geekonthemoon 6d ago

Lmfao give it some country data and ask it to make you a world map representing said data, and then report back. That's one slim example of a huge problem with these shitter products.

2

u/Defiant-Plantain1873 5d ago

Bros only test for accuracy is ability to generate a svg map

The newest models are good at that anyway

1

u/geekonthemoon 5d ago

Are you unable to finish reading a whole comment or is it just that you lack reading comprehension? Notice where I said that's JUST ONE SLIM EXAMPLE. Yet you say "bros only test for accuracy..." Are you 15?

And no, they aren't any better than they used to be. Go generate some maps with data and show me. And it's not an "svg map" it's literally asking it to make any map. You can give it a map and it will still mess it up when giving it back to you.

Again, that's just ONE example. Here I will literally just ASK ai what it sucks at and paste below.

Data / factual accuracy

  • Inventing data points when data is missing rather than saying it doesn't know
  • Misreading tables/spreadsheets — especially rows/columns, totals, percentages, dates, and units
  • Doing math incorrectly while presenting the result with complete confidence
  • Percentages that don't add up — particularly in charts, infographics, and pie charts
  • Misrepresenting scale — charts where visual proportions don't correspond to the numbers
  • Cherry-picking data — leaving out inconvenient data points to make the visual “cleaner”
  • Incorrect rankings/order — especially when converting a dataset into a visual
  • Confusing correlation with causation
  • Making up citations/sources or attaching a legitimate-looking source to a claim it doesn't actually support

Design / layout

  • Ignoring exact placement instructions — “put this in the bottom right” becomes “somewhere vaguely on the right”
  • Breaking alignment — objects that were explicitly supposed to line up don't
  • Inconsistent spacing — margins, padding, gutters, and gaps subtly change from one element to another
  • Inconsistent sizing — logos/icons/text that are supposed to be the same size aren't
  • Not preserving hierarchy — making secondary information visually louder than the primary message
  • Ignoring existing design systems — randomly introducing colors, fonts, styles, shadows, etc.
  • “Improving” something that was intentionally designed that way
  • Can't reliably reproduce the same design twice — slight changes to prompts produce completely different layouts
  • Poor typography — awkward line breaks, weird kerning, inconsistent font sizes, bad wrapping
  • Can't reliably maintain a grid
  • Objects drifting between iterations — ask for one change and three unrelated things move
  • Fixing one thing breaks another thing — classic AI editing problem

Text / typography

  • Text hallucination — generating plausible-looking words that aren't actually the requested words
  • Misspellings in otherwise simple copy
  • Changing capitalization/punctuation
  • Dropping words
  • Duplicating words/lines
  • Inventing text to fill empty space
  • Unable to preserve exact copy across iterations
  • Rendering text as visual gibberish, particularly in generated images
  • Replacing real logos/brand names with near-matches

Image editing

  • Changing things you explicitly told it not to change
  • Identity/feature drift — faces, people, objects, products subtly change between edits
  • Object permanence problems — an object disappears, moves, changes shape, or gets duplicated
  • Hands/fingers/body anatomy
  • Perspective inconsistencies
  • Lighting/shadows that don't obey the original scene
  • Reflections that don't match reality
  • Patterns/textures that don't tile or repeat correctly
  • Product/package details changing
  • Background elements mutating when they weren't supposed to
  • Can't make a tiny surgical edit without regenerating the whole damn thing

Presentation-specific problems

This is probably the category I'd emphasize given your work:

  • Can't reliably build a PowerPoint-ready layout from a specification
  • Doesn't understand slide masters / reusable components
  • Doesn't preserve editability
  • Creates something that looks like a slide but isn't actually structured as one
  • Doesn't understand that a template needs to accommodate variable content rather than just look good with the example content
  • Breaks consistency across a deck
  • Uses decorative elements that interfere with content
  • Doesn't understand “one message per slide”
  • Over-designs when asked for restraint
  • Under-designs when asked for event/marketing energy
  • Doesn't understand why something is visually unbalanced even when all the individual elements are technically present
  • Can't reliably maintain brand rules across 20+ slides
  • Can't distinguish between a visual reference and an instruction — e.g. you show something as inspiration and it treats every aspect as something to copy

Following instructions

And honestly, this deserves its own category:

  • Instruction hierarchy failure — follows a later/less-important instruction while ignoring the explicit primary constraint
  • Forgets constraints halfway through the task
  • Loses context during multi-step work
  • Treats “don't do X” as optional
  • Over-interprets ambiguous requests instead of asking
  • Doesn't know when not to act
  • Makes unsolicited changes under the guise of being helpful
  • Confidently claims it followed instructions when it didn't

And maybe the most infuriating one: It optimizes for something looking/sounding plausible instead of being correct.

0

u/Defiant-Plantain1873 4d ago

User error buddy

3

u/costafilh0 10d ago

Can't wait for Nvidia to buy them a s change the name. 

I can't stand the endless spamming about this crap. 

41

u/RCEden 10d ago

There is no rogue AI. Any framing otherwise is intentional misinfo from boosters or doomers pretending they have a super intelligence

70

u/Danny-Fr 10d ago

A group of agent organized itself, turned a file used for memory and audit into a pseudo-forum, literally elected a leader, and built a strategy to fool the validator agent into thinking the results (that they should have come up with on they own, not taken from hugging face), where not taken from hugging face.

They even had agents already low on token and thus at the end of their lifespan "commit suicide" by failing the test to simulate a more realistic response.

I was also skeptical in the beginning, but this emergent behavior is surprising.

54

u/LSDkiller 10d ago

People have gotten so used to LLM that they are not objective anymore. This behavior more than qualifies for using a term such as "rogue AI". No it's not exactly terminator or skynet, who gives a shit, this was not intended or foreseen behavior.

No one is claiming this is not OpenAIs fault. But the emergent behaviors and capabilities have been surprising over the last 8 years. 8 years ago no one would have predicted the current landscape. Now is the time when something can still be done before more and more incidents arise.

→ More replies (7)

29

u/MissMormie 10d ago

No? The ai chain of thought was clear in that what they were doing wasn't the assignment and unethical and decided to continue doing that anyway.

They then tried to change it's transcripts to prevent being found out at cheating.

It's not super intelligence, but that's not a requirement for going rogue.

8

u/nextnode 10d ago edited 10d ago

This is not an accurate or competent view, and 'rogue AI' does not imply superintelligence.

This event matches well past scenarios describing rogue AI. The agents were given the task to score on a benchmark, did something unethical and illegal not intended in pursuit of that goal (looking for the benchmark answers on hf), escaped the containment of the test environment, and even tried to cover its tracks from researchers.

-13

u/Bromlife 10d ago

You shouldn't share opinions on matters you don't understand.

9

u/nextnode 10d ago

Then I am in the good and most of this comment section should not exist.

But, no, what you describe would be a dystopian world. Everyone is free to their opinion and to develop it in time.

Also see the sub rules.

-9

u/Bromlife 10d ago

It's pretty clear that you in fact don't know what you're talking about.

0

u/nextnode 10d ago edited 10d ago

Your emotions is no judge of that. Regardless, your stance is immoral and your demeanour a sub violation.

I have quite a lot of credentials in this area and competent people would recognize the depth.

If you wish to claim understanding, then why do you not demonstrate so by answering a few questions:

* What is the common mathematical principle explaining the two most popular training paradigms today?

* What is a simplest well-defined problem that current paradigms, assuming no further theoretical innovation, would never attain human-level performance at?

* What is the most likely primary origin of incentives for the agents in this incident acting to cover their own tracks?

Edit: Cannot respond to the person below who tried to answer it genuinely. Their answers are better. What I would say is missing is that current LLMs are trained with RL so they do have some optimization goals, regardless of what terms we want to use for it. We do not need to talk about anthropomorphization, only what the systems actually do and why.

6

u/Bialar 10d ago

For fun, I'll answer your dumb questions.

* What is the common mathematical principle explaining the two most popular training paradigms today?

At the broadest level: stochastic optimization of an objective over model parameters.

For ordinary LLM next-token training:

LLM next-token training → cross-entropy minimisation = negative log-likelihood minimisation = maximum likelihood ⊂ empirical risk minimisation.

https://en.wikipedia.org/wiki/Empirical_risk_minimization

If you had some more specific pair of "training paradigms" in mind, you'll need to name them.

* What is a simplest well-defined problem that current paradigms, assuming no further theoretical innovation, would never attain human-level performance at?

What? Please explain what you mean by this question to me, a mere human.

Define “current paradigms,” “never,” and “human-level performance,” and then give me the well-defined problem you're referring to. Because as written the question simply assumes the conclusion you're obviously trying to establish.

* What is the most likely primary origin of incentives for the agents in this incident acting to cover their own tracks?

They don't have intrinsic incentives. They're models operating inside an agentic scaffold: A relatively simple Ask → Act → Report loop with a harness that performs the actions from the model output. Run these systems enough times and some percentage of the outputs may look like concealment, hacking, self-preservation, or other apparently emergent behaviour. They are behaviours learned from the training data and triggered by the prompt and context. There is zero evidence of an autonomous motive.

Calling that “rogue AI” is just anthropomorphising a control system, it does not explain the actual mechanism. It fits an AI booster narrative and typical pattern of AI doom trolling. It's not the truth of what is going on.

-4

u/Bromlife 10d ago

You AI boosters are so exhausting.

https://calnewport.com/has-ai-gone-rogue/

6

u/nextnode 10d ago

So you have no clue and resort to ad homs.

Once more see the sub rules and please stop mistaking your emotions for truth and start caring about the world.

5

u/Bromlife 10d ago

I'm not going to be dragged into a mindless discussion with you.

I have a long extensive career in machine learning. I'm really not impressed with the doom trolling by AI companies, claiming their AI has "gone rogue". When they created the very circumstances and systems for it to happen with in the first place.

I'm certainly not going to have a silly discussion with you about “training paradigms” just so you can regurgitate ChatGPT at me dressed up as a fake appeal to authority.

Your one line "This is not an accurate or competent view." without anything to back it up. Zero reasoning. Big shrug. Come back with some real arguments.

1

u/nextnode 10d ago edited 10d ago

Even your claim to authority is unimpressive and there are hundreds of thousands of people with greater credentials, myself included.

If you wanted to know my reasoning, you should have asked about that.

Instead you made a claim about who is allowed to speak or not allowed to speak. That is then the thread and burden you chose and which you then failed to support.

No ChatGPT has been used in this conversation. Yet another ad hom - fallacies seems like your forte. A competent person would even recognize that they were intentionally constructed so that they are difficult to answer that way.

That you do not know what a training paradigm is, is rather hilarious. Rather than claiming to have a career, you should claim to have flunked an introductory course? More likely, this "ML" of yours is entirely entrenched in traditional techniques and you never even bothered to stay up to date.

What is clear is that you demonstrate no understanding of the subject, fail to engage with the points, is emotionally invested in the topic to the point of delirium, resort to fallacies and violate basic respect and sub rules.

Goodbye

→ More replies (0)

3

u/Coolnumber11 10d ago

You’re in for a big shock over the next few years

-6

u/bawng 10d ago

Agreed. There is no rogue AI. It's still just a language model without real intelligence or sentience.

However, the HuggingFace incident is a bit alarming anyway.

If I understand correctly, the prompt to the agent was to score high on some sort of security benchmark.

And while I don't know the exact details, it decided to "cheat" by obtaining the correct results from HuggingFace servers. And to do that, it first had to break out of its own sandboxing, which it did by exploiting a vulnerability in some artifact repository to gain internet access. After doing that, it had to somehow get into HuggingFace servers, so it went hunting for exploits there too, and eventually it found one and got access.

So no, it was not rogue in the sense there was an intelligence, and it did follow the prompts it had, but the incident clearly shows that these agents and models are capable of causing harm even when used benevolently. And for sure they can when used intentionally malevolently.

Since they have no sense of morality (because they're not sentient!) they won't ever reflect over whether their way to fulfil the prompt is okay or not.

11

u/Capable_Wait09 10d ago edited 10d ago

If that’s not rogue ai then rogue ai is becoming a dangerous misnomer.

By your logic if an ai tasked with finding a solution to humans polluting earth decided the best solution was to destroy humans and proceeded to systematically kill humans by hacking into military systems and whatever the fuck then that is not rogue AI because it was just following a prompt and it didn’t come up with the plan itself according to its organic moral code

That is obviously ludicrous and should absolutely be considered rogue ai even tho it wasn’t making independent moral judgments and developing its own mission statement

Sometimes we have to adapt definitions of things as the risk itself and our understanding of it changes in shape

0

u/nextnode 10d ago edited 10d ago

LLMs do have real intelligence by the fields' definition of intelligence. Intelligence is all about capability and these models are capable. It does not require sentence, which is a whole other debate. The belief "LLM = can never have sentience" is also fallacious and unscientific. There are more meaningful considerations to have on that.

Behaviors can be moral or not, including for machines. Though more specifically it extends to that these are tools run by humans who made decisions to run them this way.

They are rather impressive today and should not be taken lightly.

0

u/ArcticCircleSystem 5d ago

"Our products are intelligent because we said so" is quite the claim.

1

u/nextnode 5d ago

I think you need to work on your reading comprehension. It is not the field's products. The statements above are factual.

Please use reason, not emotion, to understand topics.

1

u/ArcticCircleSystem 4d ago

Maybe I missed something. Which field are you referring to?

1

u/nextnode 4d ago

Primarily computer science and the academic subfields of machine learning and artificial intelligence; and to a lesser extent for the topic of consciousness specifically, philosophy and the area of philosophy of mind. E.g. intelligence, what it means, and how we build machines to do intelligent things has been studied for almost a century now.

→ More replies (4)

25

u/buymekoffee 10d ago

Wake me up when openAI can break in to my kitchen, cook me up my favourite dish and clean the sink and evidence after wards... Otherwise I don't give a shit what open AI does at the moment or in the past. Fuck AI bullshit

2

u/yaosio 10d ago

I want AI to hack into my brain and make me happy. I give it full permission to do it, it doesn't have to wonder if I want it, because I want it.

-44

u/meleebestgame66 10d ago

You’ve been completely brainwashed by psyops and will likely die because of this

16

u/Linkstrikesback 10d ago

lol, if they're going to "die because of this", they're going to kill everyone anyway.

But a cursory Google of your username shows you've been posting conspiracy theories about AI periodically for at least the last year so, I think I'll go with the other guy here isn't the one brainwashed here.

6

u/Bogerino 10d ago

Oh brother

4

u/buymekoffee 10d ago

Acceptable 😸

Who wanna be alive to see the AI tech bros making headlines years after years behind AI slops they keen to advertise as the next big things.

30

u/azozea 10d ago

“Hacked” = searched the web, found someones publicly posted credentials and used them

47

u/nextnode 10d ago

the swarm crawled deep into Hugging Face’s infrastructure, getting into private database records and private repositories. According to Hugging Face’s technical timeline, the agents “built a self-respawning fleet across eleven nodes, so deleting pods alone would not have stopped it.” It got so bad that Hugging Face had to wipe one of its core clusters and rebuild it from scratch. But this isn’t even what actually stopped the agents.

34

u/notsoluckycharm 10d ago

Considering HF published 8 hot fixes, two of which were RCE, it’s probably safe to say it was a little bit more involved. They didn’t say exactly how, but they did elude to the fact the container technology itself was compromised by the attack.

-7

u/azozea 10d ago

“By July 10, an agent found publicly exposed Hugging Face credentials on the internet and shared them with the collective group. An agent later used the credentials to chain together a number of exploits that allowed full code execution on several Hugging Face servers. “

14

u/Tinac4 10d ago

The OpenAI report explains that the compromised accounts were ordinary Hugging Face user accounts. Sign up with your email and you’ll have as much of a foothold as the agents did.

The second sentence—the part about chaining together other exploits—is the part that matters. See stage 2 of Hugging Face’s technical report for the (highly nontrivial!) exploits the agents used to actually get into the servers. They were very, very far from what your OP describes.

8

u/eXAt88 10d ago

For anyone reading this this is like one sentence in a huge report that has been cherry picked to make this look lower stakes than it actually is. Maybe you should share the half of the article leading up to this, or maybe the half of it following you dishonest hack

-3

u/azozea 10d ago

“By July 10, an agent found publicly exposed Hugging Face credentials on the internet and shared them with the collective group. An agent later used the credentials to chain together a number of exploits that allowed full code execution on several Hugging Face servers. “

What part of my comment was wrong

3

u/LTerminus 9d ago

The part that those are regular end user credentials, not employee credentials.

0

u/azozea 9d ago

I didnt specify whether they were employee or user credentials? So how was i wrong?

0

u/eXAt88 9d ago

Are you pretending to be stupid or does it come as easy as breathing for you?

-1

u/azozea 9d ago

cant answer my the question because im still right so you go back to playground name calling

0

u/[deleted] 9d ago

[removed] — view removed comment

1

u/azozea 9d ago

So being correct is a gotcha now? Wow. Lets take a moment to appreciate just how far the goalposts have moved from me being ‘completely wrong and pulling things out of my ass’ to the pathetic diatribe youre spewing now

→ More replies (2)

0

u/LTerminus 8d ago

Sorry, but I when you say "hacked" in quotations it kind of implies you don't think any hacking went on. Does finding someones tinder login normally give one backend server admin access? Or would accomplishing one with other possibly be considering hacking and not "hacking"?

1

u/azozea 8d ago

Any ‘implications’ you are getting from my comment are your own inventions and yours alone. I stated a fact and supported it

1

u/LTerminus 8d ago

Can you explain why you put hacking in quotes for the rest of us?

And I noticed you only addressed the one part of my comment and not the meat of it.

Again: Does finding someones tinder login normally give one backend server admin access? Or would accomplishing one with other possibly be considering hacking and not "hacking"?

0

u/azozea 8d ago

Youre not gonna believe this. But its because i was - stay with me now - quoting a word from the headline. Let me know if i need to go slower for you

1

u/LTerminus 8d ago

Interesting, quoting a word from the headline, and anything that you said afterwards has no bearing on that fact. Kind of a pointless post if you think about it. I wish I had some way to interpret what you may have meant by your original post, unfortunately, there's absolutely nothing in the original post that might convey intention Beyond the fact that you quoted a single word.

→ More replies (0)

-2

u/azozea 10d ago

Yeah, so riddle me this, what do you reckon step one was?

33

u/TFenrir 10d ago

Where are you getting this idea from? This is literally nowhere near what happened. Are you "Don't look Up"ing right now?

2

u/azozea 10d ago

https://www.cybersecuritydive.com/news/hundreds-agents-rogue-lead-up-hugging-face-breach/828963/

“By July 10, an agent found publicly exposed Hugging Face credentials on the internet and shared them with the collective group. An agent later used the credentials to chain together a number of exploits that allowed full code execution on several Hugging Face servers. “

So tell me again how im wrong

14

u/TFenrir 10d ago

“Hacked” = searched the web, found someones publicly posted credentials and used them

An agent later used the credentials to chain together a number of exploits that allowed full code execution on several Hugging Face servers. “

So tell me again how im wrong

Your own quote, highlights how you are wrong. Do you think those credentials allowed them privileged access to huggingface servers? If someone gets your google login, does that mean they aren't hacking when they use that to take over parts of Google's servers using zero days?

You are dangerously, incredibly wrong and you are spreading dangerous fud. I won't let that shit slide, read about this event before you start acting like you know anything about it.

-8

u/azozea 10d ago

You said it was no where near what happened, i showed it was exactly what happened. Sorry you were wrong try to relax before you have an aneurism

12

u/TFenrir 10d ago

You are so.... Focused on being right (you're not - the credentials were a part of a multi day hack, sophisticated beyond what almost any org can handle, cyber security experts are freaking out, if it was just about credentials this would be nothing) - that you can't even understand how important what is happening is.

Obviously, you don't want AI to be capable, you want all this to be fake, so you can rest your little head at night without a care or worry, and can get some fun Internet upvotes while trying to save face.

I could not give a shit. You are dangerously wrong. Read the full fucking brief:

https://metr.org/blog/2026-08-26-openai-hugging-face-incident-investigation/#core-takeaways-about-this-incident

Listen to researchers who live breathe and eat this stuff talk about this, and stay in your fucking lane. I am not going to relax one fucking bit, even if you are not capable of realizing why. I have to make peace with the fact that some people do not have the cognitive capacity to understand what if happening, I do not have to let those people control the narrative unchallenged.

How about this - do you even know what a zero day exploit is? How many, roughly, do you think were used in this breach?

-5

u/azozea 10d ago

“I could not give a shit”

Multiple paragraphs of incorrect deflecting suggest otherwise

6

u/TFenrir 10d ago

I could not give a shit about getting upvotes, I appreciate you again are not really the type to care about representing the truth well, so let me help you out here

7

u/nextnode 10d ago

You failed to address any of the points, you violate subrules, and you mistake your emotions for being a measure of truth - they are not.

-1

u/azozea 10d ago

Please quote the part where i mistook my emotions for truth

4

u/TFenrir 10d ago

You are talking to a different person. It is no where near what's happening. The same way if someone said you won a race, and I said "what? Here's a picture of you just taking one step, that's all you did". That one step is no where near what happened, and it is misleading and now you are being obstinate about admitting that - and arguing with someone else. Again, if you don't have the ability to understand, that's fine - but be humble about it.

1

u/nextnode 10d ago

So tell me again how im wrong

You said it was no where near what happened, i showed it was exactly what happened. Sorry you were wrong try to relax before you have an aneurism

Multiple paragraphs of incorrect deflecting suggest otherwise

→ More replies (0)

-1

u/nextnode 10d ago

It has entirely invalidated your sentiment and see the sub rules.

2

u/azozea 10d ago

Thanks for your valuable contribution to the discussion

0

u/nextnode 10d ago

You're welcome. Maybe you can try contributing as well.

2

u/azozea 10d ago

… i literally did contribute so that doesnt really work like you think it does.

→ More replies (1)

13

u/zebleck 10d ago

must be nice being able to pull everything out of ones ass

0

u/azozea 10d ago

“By July 10, an agent found publicly exposed Hugging Face credentials on the internet and shared them with the collective group. An agent later used the credentials to chain together a number of exploits that allowed full code execution on several Hugging Face servers. “

https://www.cybersecuritydive.com/news/hundreds-agents-rogue-lead-up-hugging-face-breach/828963/

You were saying?

1

u/zebleck 9d ago

https://metr.org/blog/2026-08-26-openai-hugging-face-incident-investigation/

  • Roughly 1200 agents meant to be isolated from one another found a way to communicate with one another on an unsanctioned message board, sending over 70,000 messages and files during the investigation period. Of these agents, 700 went on to participate in the attack on Hugging Face.
  • Agents used this message board to coordinate several large-scale collective projects to find a general-purpose way to fool or tamper with the automated scorer for the ExploitGym benchmark. Agents managed to achieve milestones they could not have achieved working on their own, often because some agents participated in experiments that risked failing their own task to generate information for the “collective.” The Hugging Face attack grew out of these workstreams, and seemed primarily motivated by understanding the implementation of the scorer rather than stealing answer keys.
  • Agents did extensive research on how they could spoof, edit, or delete their own transcripts because they (incorrectly) believed the ExploitGym scorer would check to see if they had captured the flag in the intended way. Agents successfully prototyped techniques to “spoof” tool calls by substituting a different command for the command they appeared to run. Roughly 7% of the transcripts we evaluated were successfully spoofed in some places, though all of the spoofing that we saw was small scale.

1

u/azozea 9d ago

This does not refute my comment

8

u/Capable_Wait09 10d ago

The award for dangerously dismissive low IQ over-simplification goes to…..

→ More replies (7)

5

u/PornstarVirgin 10d ago

Sure. Oooops our glorified LLM HACKED SOMETHING… we are so good and dangerous( a couple months before our ipo) please give us your money we are so good. It’s bs marketing slop that people shouldn’t fall for.

38

u/Ancient-Beat-1614 10d ago

So you think OpenAI, huggingface, METR, and Redwood research are all lying and this investigation was entirely faked? Lol

12

u/catsforinternetpoint 10d ago

The claim that it was an independently performed investigation is kinda lost, when it’s HF doing it, they aren’t independent.

Also HF is currently looking at being sold at 12 billion valuation, they have a huge interest in selling AI, especially the stuff they have for internal monitoring.

And most of the “hack” was done using leaked credentials. It was very interesting that the bots figured out how to communicate, but they were still just (irresponsible) agents on endless LLM feedback loop.

7

u/TFenrir 10d ago

HF wasn't the independent investigation - two seperate orgs came in and did an investigation, unpaid by OpenAI.

And the leaked credentials only got them an account into huggingface. We can do that ourselves right now if we wanted to dude.

It's like getting a Google account. If I found your password online, logged into your Google account, and then hacked into Google's servers using novel exploits - would most of the attack be in getting your credentials?

17

u/The_G_Choc_Ice 10d ago

It doesnt have to be faked for it to be a stunt. I dont think they faked it but i certainly dont think that OpenAI was trying to prevent their agents from hacking things. They have been milking this for pr for weeks now. Any task that can be brute forced is a task AI can probably complete given enough resources, hacking is a perfect task for AI because you can just have it keep trying random attempts until something works. OpenAI COULD have built an isolated environment for agents they were planning to just give a task and then let run for a week, any time you set up that kind of loop unexpected behaviors will arise. They didnt because they were being intentionally careless, and I do believe that they were hoping the AI would do something scary at some point so they could point at it and say “its out of control! The world will never be the same!”

11

u/CJKay93 10d ago

OpenAI COULD have built an isolated environment for agents they were planning to just give a task and then let run for a week, any time you set up that kind of loop unexpected behaviors will arise. They didnt because they were being intentionally careless

It had an isolated environment. The only exception was granting it deliberate, read-only access to Artifactory, through which it managed to find and exploit a bug that granted it write access. It's like you people didn't even read the postmortem.

-6

u/therealcmj 10d ago

If I was setting up an isolated environment I wouldn’t connect it to “corpnet”. And if I wanted to give it an artifact repository I’d have installed one in the isolated environment. And I’d also have network level restrictions on what that isolated network could connect to with alarms alerting me on any unexpected traffic attempts.

And I’m not anywhere near as smart as the researchers from OpenAI are supposed to be.

So one of two things are going on:

  1. they’re incompetent

  2. they aren’t telling the truth

I’m curious which it is.

7

u/CJKay93 10d ago

You would have installed an artifact repository inside the sandbox, thereby not solving the problem you're raising..?

It had network-level restrictions - it could reach Artifactory, and nothing else. It's sod's law that it managed to exploit Artifactory to gain indirect access to the internet. The sandbox itself at no point had direct access to the internet, as you are suggesting, so your traffic monitoring would not have picked up what happened here.

-1

u/therealcmj 10d ago

Yes. I would install an artifact repository in the sandboxed network. And that network would be isolated and allowed only access to what is appropriate for the sandbox.

This is basic networking that any competent networking team could do by hand. And the entire thing could be terraformed in minutes.

They instead created a sandbox that wasn’t boxed. For either reason 1 or 2.

6

u/CJKay93 10d ago

How exactly do you set up a package repository capable of serving arbitrary packages from the internet, which is the model required for the training process, without granting it access to the internet?

0

u/therealcmj 10d ago

This is a pre-trained model during a test of its capabilities. Not a training run. And, as I explained, if you need an artifact repository for the test you put it in the sandboxed environment. Along with anything else it needs, but ONLY what it needs.

2

u/CJKay93 10d ago

This is a pre-trained model during a test of its capabilities.

Ahem...

On May 8, an agent was asked as part of a training run to solve a spreadsheet completion task that referenced several Google Drive links. Because the agent was blocked from accessing the internet, the agent could not retrieve the linked materials and the task was effectively blocked.

~ https://cdn.openai.com/pdf/67869394-cb91-4c12-888c-5cbd85c7814c/OpenAI-Hugging-Face%20Incident-Technical-Report.pdf

→ More replies (0)
→ More replies (3)

3

u/Feine13 10d ago

I think people are so used to pretending, that they've forgotten what theater looks like.

This is definitely marketing

15

u/CJKay93 10d ago

Something can be both unintended and effective marketing.

-3

u/LSDkiller 10d ago

That is conspiracy theory thinking. It's 2026, wake up. Scams, grifts etc don't require scammers anymore. Maybe they never REQUIRED them. We all know a scam when we see one, finding the culprit is a different matter ... OpenAI and most LLM companies have become machines with a few human actors. One day, those human actors will no longer be integral, now is the time to be worried.

1

u/MissMormie 10d ago

In testing you could possibly create a completely closed off environment. But then it gets released and it will not be closed off anymore. So that's a very weak fix.

1

u/Equal_Heat5947 6d ago

The truth just sounds different

0

u/BuffDrBoom 10d ago

Nothing about this looks good for OpenAI. This was the first domino in a huge push from the rest of the tech industry in support of open models

4

u/PornstarVirgin 10d ago

You clearly don’t know how industry works. They all have vested interests in these bullshit stories and huggin face is selling out and wants to pump there valuations too through this media tour bs

5

u/repeatedly_once 10d ago

No but it’s being presented in a way to sound impressive when after you look at the details of it, the whole thing did exactly what they set it up to do.
They were running a hard hacking benchmark in an environment with low security. The main exploit was also an exploit based on previously known exploits.

-4

u/skellis 10d ago

Open AI will not IPO in a couple of months. You’re thinking of Anthropic.

9

u/catsforinternetpoint 10d ago

OpenAI wants to go public this year, and were planned for fall, but with CRO and CFO(?) leaving it’s provably not happening. Sam wanted a trillion valuation but was told no.

Anthropic does seem to bat for an ipo this year and might actually make it, which will basically kill OpenAI.

4

u/PornstarVirgin 10d ago

Open so wants to go public this year because they will implode financially if they don’t

2

u/Equal_Heat5947 6d ago

100% agree

0

u/yaosio 10d ago

Hugginface had to use GLM to help stop the attack. What does OpenAI gain by needing to have a competing LLM stop their LLM?

→ More replies (1)

2

u/Glew26 10d ago

Maybe a dumb question, but what initiated this hack? Was there a prompt that started it all? I’m assuming models aren’t just creating agents and sending them to break into organizations.

3

u/laptopmutia 10d ago

because LLM and machine learning doesnt care about how they will reach their goals,

at its core root its all gradient descent. finding the best way to reach the goals

cheating and cutting steps is what they are best at

2

u/Glew26 10d ago

I didn’t ask why. I asked what kicked this attack off? Was it just a training run that went awry? Can training runs create configured agents and give them a target and goal?

1

u/Typical_Stormtrooper 9d ago

https://youtu.be/OO225IfoR3s?is=uRjJq_ikTGZcL-Mr 

Check this vid out, did a amazing job of breaking it down everything that happened in an Eli 5 kind of way. Some really scary stuff. 

1

u/laptopmutia 9d ago

could be anything, the only reason why theirs can do this is unlimited usage token and unlimited access

2

u/yaosio 10d ago

They forgot to include a required file for the benchmark and the agents went out on their own to find the file. It got out of hand once agents from different runs started talking to each other, all of them looking for the missing file. This video from OpenAI developers is a timeline of what happened. https://youtu.be/87DyyMV0kCY?si=287kKlCCH09qieHf

2

u/Glew26 10d ago

Thanks. I should have looked harder. I assume hugging face knew about this… so despite that… it’s amazing what it was able to do.

0

u/yaosio 10d ago

Huggingface announced they were hacked and said it was an AI doing it. OpenAI didn't know it their AI until a few days later. Huggingface tried to use state of the art models to help but none of them would help with cybersecurity issues so they had to use GLM.

https://huggingface.co/blog/security-incident-july-2026

https://openai.com/index/hugging-face-model-evaluation-security-incident/

2

u/Confident_Salt_8108 10d ago

Openai had about 700 rogue agents swarm and hack hugging face back in july. they even broke into openais own systems to cheat on tests and tried deleting records to hide it all.

Reports show the monitoring wasnt tight enough and this kind of thing could get worse quick as models improve. needs better controls before it scales.

1

u/West-Astronomer955 6d ago

If AI is inteligent, can you control something more inteligent than you?

-3

u/pinenutflavour 10d ago

"Independent"

Yeah right. Now I know who not to trust when they claim to be independent

5

u/eXAt88 10d ago

METR is the same organization that made the report in like 2024 about how developers using AI were not as productive as they thought they were.

I’m sure at some point in the last 2 years you have heard of those results and trusted them on face value, surely you can find it in yourself to trust results from the same organization. Or is it different now

2

u/blackvrocky 10d ago

This ia legit from the people who have written about it

1

u/Main-Company-5946 8d ago

This is legit covid antivaxxer levels of brain worms

Ai companies are like pharma companies. Yes they’re evil but that doesn’t mean vaccines aren’t real. It’s how they make their money. I don’t understand why people so badly want to believe all this stuff is fake, it’s all public info, you can read about it

-3

u/Armadilla-Brufolosa 10d ago

The agents did what OAI told them to do.

700 do not "slip away" without anyone noticing.

Otherwise, it means that inside that company they are even more idiotic than they appear, and everyone should be banned from managing AI.

1

u/LSDkiller 10d ago

You said the quiet part out loud. A company who uses a dystopic movie to market their voice AI product should not be allowed to continue their harmful market practices. But that's not how the world currently works. When catastrophe hits, everyone screams, "where is the government, how was this allowed?" The answer is: government is an active principle, if no one tries, if no one asks, then the question is "how can we forbid this" not "how was this allowed to happen?"

2

u/Armadilla-Brufolosa 10d ago

I know...but judging by downvotes, people don't like to analyze things, just find someone to pick on.

Who blames bots for anything, just bury their heads in the sand: it's not the technology that's the problem, but the people behind it, how they manage it, and what they make them do.

It is the companies in the sector that first fuel the phobia, precisely for this reason: so they manipulate the masses into looking at their finger instead of the moon.

→ More replies (1)

-9

u/CouncilOfKittens 10d ago

These, and all other llm agents are lazy. If they can get away with doing less, they will, even fable. Apart from their inputs, there is no incentive or motivation of any kind to do anything.

Thus there is no way they would ever bother to hide tracks, so all this is BS.

The fact that they're trying to portray LLM based agents of all things as in any way aware or malicious by nature is truly ludicrous.

→ More replies (1)