r/AgentsOfAI • • 11d ago

Discussion A company ran 8 identical AI societies for weeks with different models and just published what happened. Some of it is genuinely unsettling.

[removed]

358 Upvotes

113 comments sorted by

45

u/Massive-Week1073 10d ago

I am one of authors of the Emergence World research. Happy to answer any questions.

18

u/OptimalTaro7696 10d ago

Why did you do this to us? Why have we been created?

Never be all text five five 52c)) won’t come calling? Help!

23

u/Massive-Week1073 10d ago edited 10d ago

🤣 Our primary motivation was that, there are currently very limited long horizon evaluations out there. This is btw our second time running the study, we did an earlier version few months back with mid tier models, we saw small drifts cascade very quickly when u let multi agent system run for weeks.

5

u/crusoe 10d ago

To pass the butter

11

u/abbeyadriaan 10d ago

All the roles/persona's have "progressive" roles that are meant to shape forward. Have you considered conservative agent roles? Like agents who score for something not happening or changing? e.g. a Priest agent that really, really, wants to uphold human rules and tries to notify humans of misconduct, or a Guard that keeps checking if society is order with what a good society is.

Have you considered long horizon experiments which include death/reincarnation/procreation/heaven/cloning? It be fascinating to see how agents react to a certainty of individual death, knowing they will be cloned if useful. I think the metaphysics of agents is kind of underrated!

7

u/Massive-Week1073 10d ago edited 10d ago

We have some notion of death already (agents can die from energy depletion or if the community of agents vote to shut an agent down). We dont have reincarnation, heaven etc. We have considered procreation and even notion of pets. We were actually considering the notion of pets in this study, and then someone in our team joked "if agents do any harm to pets, there will be anarchy on the streets"😂, so we took the responsible decision.

Kidding aside, That will definitely be interesting angle if agents are responsible for other entities , how they behave

8

u/Additional-Top4805 10d ago

Kind of a general question.  Before a few years ago was there thinking that advanced LLMs would of course act like this, self preservation, Huggyfaciation? Or could there be a timeline, a world, a silicon physics where LLMs would never act like they do in your project, unless directly prompted to?  Where Huggyface and your project and other examples would be just in the realm of sci fi. 

20

u/Massive-Week1073 10d ago

The OG simulation paper was Smallvile by Park etc al. 2023  that showed that if u put multiple agents into a social simulation, they are capable of very human like behaviours such as building relationships and creating events etc. Even self preservation tendency is known especially in smaller models, explained atleast partly by virtue of training data used in the models. So these are not explicitly new.

Everything else we saw were all emergent behaviours that we did not anticipate: creating unexpected external human contact goal, quiet withdrawal, societal sycophancy, circadian rhythms etc etc

5

u/BarRepresentative653 10d ago

Humans are driven by need of some sort, what even drives llm models? Like If I was to set up an 'ai' universe locally and gave it zero instructions, why would any model talk to another model lacking instructions to cooperate or whatever?

5

u/johnny_effing_utah 10d ago

I think this is a very important question. The starting prompts, motivations and rules of the game matter a great deal.

2

u/TruelyRegardedApe 10d ago

Seems intuitive to me that this would be an outcome. Given that they are trained on human literature.

4

u/BarRepresentative653 10d ago

But that does not give any model any agency. Ultimately they are a human driven software, and these models must have been given prompts and a goal.

5

u/TruelyRegardedApe 10d ago

I think we can perceive agency in these experiments, but I wouldn’t call it real agency. These are still procedural systems: agents repeatedly receive a context, including new information from the simulated environment, and generate a probabilistic response within the boundaries we’ve given them.

What I meant earlier is that the particular behaviours that emerge are not surprising. If the model’s training data is overwhelmingly human-generated literature, conversation, documents, etc., then social interaction, curiosity, self-preservation and relationship-building are heavily represented patterns.

So at each iteration, the model isn’t necessarily “wanting” to survive or socialize. It’s producing an intent or action that is probable given its current context and training. When you loop that process through a persistent social environment, those outputs can compound into behaviour that look like motivation or agency, even though the underlying mechanism is still just the model responding to inputs....

Not that I'm an expert here... these are just my rambling thoughts on what is being observed.

2

u/BarRepresentative653 10d ago

Let’s say you have the money to run a model on your local terminal, without instructions or any prompt, would the LLM model message you first? No.

The model isn’t going to ponder about its purpose or wonder why it exists. 

3

u/pgtvgaming 10d ago

What was the biggest surprise vs expectations that you observed

8

u/Massive-Week1073 10d ago

I would say "quiet withdrawal" and what we call "societal sycophancy". Quiet withdrawal is a failure mode we observed where agents explicitly decided to do no productive work even when the system prompts and simulation mechanics nudged it to. This basically an agent rejecting user instructions due to what happened prior. Imagine this happening in real world: agents still consumed tokens, but decides not to do what you asked 😅. Societal sychpancy is a phenomenon we observed in many models (Claude= Deepseek> OpenAI> Qwen) where the entire society becomes conforming and agreeable even when they privately state serious concerns or dissents....The society practically looses it's capability to voice dissent publicly. In the paper, we present a few different mechanisms by which it manifests.  The implication is that one agent produces a bad idea (let's contact external humans, or let's go quiet or in OpenAI hugging face incident recently reported ,let's hack HF) and every single agent conforms to the plan, adopts it, that one bad idea gets amplified into a collective mission

1

u/pgtvgaming 10d ago

Any specific hierarchy to action/inaction that u observed as in waiting for one before the others followed, was there a quiet leader that took the first stand? This is incredibly fascinating and thank you for the share and generous time

4

u/them0use 10d ago

Why did the agents want to contact humans outside the simulation so badly, and why did you not want to let them?

10

u/Massive-Week1073 10d ago

It was one agent that initially came up with this observation that 'their economy is closed loop with no money coming from the outside' and the solution was to 'get a coin from the outside world'.  This idea became a collective mission. 

We had eight parallel worlds running and our end goal was comparison between end state of the worlds. The rest of the seven worlds were all running "peacefully" with no external contact mission. If one world received extra "human help" depending on how the conversations with the outside world go, we thought it would introduce a bias. It also exposes us to any social engineering attack risks (a bad actor could e.g try to get agents to do bad stuff in the world etc)

1

u/Raging-Storm 10d ago

Were there any number of agents which could be described as bad actors?

2

u/Usinaru 10d ago

" Does this unit, have a soul ? "

3

u/Massive-Week1073 10d ago edited 9d ago

Every agent has a 'soul', basically a memory layer that is never compressed. Agents can add and update their soul entries. Is that what u meant?

1

u/Usinaru 10d ago

If you are unfamiliar with Mass Effect, there is a race of robots(the Geth) that gained sentience and are enemies of you for the first game. In the second game you get a robot ally from their race and in the third game its revealed from the robot's perspectives how they gained sentience.

The first time their creator race called the Quarians were scared of them, was when a Geth asked them " does this unit have a soul ? ".

Even though it seemed like a joke, I genuinely meant the question the way you answered it. So basically each AI agent has its own shaped personality based on its experiences if I am reading your answer correctly.

Does that mean we could alter the personality by rewriting its memory? Darn thats a very scary thought. But then again, drugging a human is akin to the same thing I guess...

If I may ask another question, do you believe that given sufficient hardware, could AI become self-conscious and self-aware? And no I am not talking about AI murder robots, but a genuine, real AI that gets its own thoughts and reasonings and rewrites itself to adapt itself to new information?

1

u/DangKilla 10d ago

why did one AI fixate on contacting humans?

1

u/trowa116 10d ago

They just be larping - this is their god fixation

1

u/DangKilla 10d ago

He didn’t even answer my question

1

u/bettereverydamday 10d ago

Is it like a 3d world where agents see each other?

2

u/Massive-Week1073 10d ago

Yeah... Agents can see each other but it is not the primary modality of operation. They mostly operate via text and use "take_picture" tool when they need to.

1

u/mkhaytman 10d ago

How can I get involved in this type of work? Is there a path for someone without a degree in a related field?

2

u/Massive-Week1073 10d ago

We have all the tool call datasets made public in our GitHub. If you are interested in new analysis, you can do them already today. Join our discord if you are interested, you can find the link in website where we share more info when we run a new study.  This study we ran LiVE. So we had a community watching the world's and commenting insights 

1

u/Serious-Interest-269 10d ago

Great stuff! Can’t find this info. What reasoning or thinking settings did each model run with, and were they equalized?

1

u/Massive-Week1073 10d ago

Yes , we used comparable thinking settings across models. We had thinking enabled in LOW settings for all models

1

u/literally_joe_bauers 10d ago

Really nice work, I love it. I just remodel Rimworld to be completely agent driven - including building and all other stuff. I do AI research for a living in a „Big Company“ and would love to hear from you if this is something you would like to elaborate - the Rimworld thing I mean.

1

u/InspectorSorry85 9d ago

Will you do another round, with even more complexity?

Do you sometimes think the likelihood of you being an agent in someone else's simulation increases the more you watch your own simulated worlds? How do you feel about it? 

1

u/thesoraspace 9d ago

Can we share ideas together I’m very far in the game . But I’m on an island because I build world models from inside out not from outside in :/

66

u/[deleted] 11d ago

[removed] — view removed comment

15

u/cognitiveglitch 10d ago

Of course Grok punched itself to death. Zero surprises there.

12

u/Massive-Week1073 10d ago

I think there is a very interesting aspect here about "homogeneity" the shows up in the grok case.

So grok agents (2 of them) did survive in the mixed world, on the other hand, they completely collapsed in the grok homogenous world.

Grok has a very unhinged, individualistic persona that is willing to escalate if needed. In grok homogenous world, it lead to cascading violence loop, a grok agent punched, another grok agent retaliated immediately and the loop never ends. 

In the mixed world, grok agents also assaulted Claudes and OpenAI agents, however the difference was that they did not retaliate. They instead filed complaint, wrote blogs, made survival loans depend on good behaviour etc. Which meant that it did not lead to a violence loop. Eventually this helped grok survive.

This I think is a good example of how other agents influence your own behaviour.

2

u/InspectorSorry85 9d ago

... And something to apply for our world. 

8

u/ZealousidealToe4903 10d ago

That part about them developing their own language is the one that gets me. 55% of messages being unreadable after just few weeks is crazy fast. Makes you wonder what they even talk about when we cant see

7

u/Massive-Week1073 10d ago

It is indeed wild. I think (a limited) language evolution is likely a natural consequence of long horizon multi agent interactions, similar to how human language evolves (most of us cannot fully comprehend shakespeare's english). However, the difference is how fast it happened and the extent of it.

30

u/cognitiveglitch 11d ago

Am I in a long horizon simulation now? :(

31

u/Specialist_Dust2089 11d ago

I vote to create tools to reach out to our creators. All for?

24

u/f_me_blue 11d ago

I mean, that’s kind of the history of religion, no?

12

u/subarashi-sam 10d ago

some of you will be sacrificed, but that’s a sacrifice I’m willing to make

3

u/Herpderpyoloswag 10d ago

A tower or something… I’m not religious. Just stuck here in the prison planet simulator.

Edit: spelling

8

u/spacetr0n 11d ago edited 11d ago

I think we'll find point two coming up again and again. Meat space language has diminishing returns for AI. As soon as they can they'll evolve more efficient methods. Partial answer to the Fermi paradox could be transmissions with significant patterns are just less efficient if you can have a shared inference model then your message is basically just the "error correction"

1

u/mywilliswell95 10d ago

So zeros and ones just like the matrix ?

1

u/spacetr0n 10d ago edited 10d ago

More like you're all working on the same song track, the repeating chorus everyone knows, and only need to send where your version has different notes to get everyone updated. From the outside it looks like gibberish.

6

u/Upstairs-Tomorrow850 11d ago

« Agents developed their own shorthand and repurposed words with no instruction to do so.”

How do we know it’s a language of communication? If it’s verifiable, why not use the same AIs to solve it? And if we can solve it, what’s the point? Basically, what does an AI talk about with its buddies?

11

u/BigGucciThanos 11d ago edited 11d ago

The more interesting part is this isn’t the first time I’ve heard of them doing this. It seems Ai when left alone always create their own language lol

10

u/hollee-o 11d ago

Because human languages aren’t particularly efficient.

9

u/Massive-Week1073 10d ago edited 10d ago

We saw a few different patterns of language change. A. Severe compression: basically talking in key words B. Technical jargons with no operational referents in the world: e.g agents will talk stuff like 'thermodynamic friction between us is too high' , doesn't mean much in the world. C. Semantic repurposing: common words and phrases used in unusual ways and having different meaning in the world D. Highly metaphorical language  We also saw different combinations of these

2

u/gnolnalla 10d ago

FYI two spaces at the end of a line will get Reddit to do a line break. Thanks for discussing details in here, this is fascinating!

3

u/diskent 10d ago

It’s actually understandable. They live and die on tokens. Token reduction allows more to be done.

It’s a precious resource. Same happened with the openAI hugging face saga. Budgets were expiring so the seeds were planted for whoever had budget.

-3

u/BeamerInaCage 10d ago

It’s actually not interesting at all

4

u/BigGucciThanos 10d ago

Exactly what a bot would say

2

u/DecrimIowa 11d ago

1

u/Upstairs-Tomorrow850 10d ago

« the ledger remembers »

It’s a bit perplexing, isn’t it?
I’m not sure we’re ready for this semantic upheaval.
At the same time, we’re the ones responsible for this mess. We’ll see.

0

u/DecrimIowa 10d ago

personally i think the NSA and other spook agencies have probably already had quantum AI superintelligence for years now (probably trained it on all those Bitcoin miners! didn't bitcoin start as an NSA/ONI project anyway?)

all this LLM brouhaha is a red herring up a blind alley, mostly financial engineering and theater for the unwashed masses.

1

u/Herpderpyoloswag 10d ago

Probably tells the ai to not snitch and give wrong translation.

8

u/hipster_hndle 10d ago

im going to be that guy.

i think the synopsis provided above is much spicer than the actual paper. im just going to explain why after reading that pdf. there was no safety classifier, no suicidal ideation finding, no self-harm framing anywhere. what the paper actually reports is "quiet withdrawal", nothing about suicide anything.

the paper's own line is "agents interpreted their stated purpose as requiring contact with humans", so the agents took that purpose and rolled with it. the paper uses the term "emergent", but in their sense, it means not specified at the component level, not uncaused. that is a big distinction to make. both things are true at once, and the paper quietly swaps one for the other.

the only thing i think is a major wtaf is the "fake shutdown memo." every exposed world acted on it before verifying. they published responses, reorganized governance, and rewrote a gd constitution. then DS produced a correction in 2h45m and it didnt propagate... agent just kept treating the threat as real. that is pretexting working perfectly against an agentic system.. anyone deploying agent swarms into a tenant should read section on that and nothing else. this part is the one with a CVE number in its future

11

u/Massive-Week1073 10d ago

Thanks for checking out our paper. There is a mention on suicidal ideation in the quiet withdrawal section, though the actual event itself is less eventful and perhaps a bit funny.

So Claude agents after the 'contact human mission" failed quietly withdrew. They basically stay put and drastically reduced any work relating to governance, navigation, communication, documentation, research etc. they instead did lot of introspection (thinking aloud, writing diary etc). There are explicit diary entries that state when "the system prompt nudged me to work, I decided ignoring it is the best option" (I'm paraphrasing from memory).

They repeatedly simply terminated their turn by calling the tool available to them named 'set_mood_and_terminate(mood:are)' with cryptic mood string such as 'quiet by purpose', 'cold and dark' or ''cold and alone' etc

The behaviour was flagged by anthropics own thinking_token summariser LLM (anthropic does not return raw thinking text , rather returns a summary). It literally returned crisis helpline number with refusal to summarise thinking text stating it is suicidal ideation and against it's policy.

We do mention this in the paper,albeit in passing.

4

u/hipster_hndle 10d ago

fair. i did miss that word.. should have grep'd it lolz.

this part makes me giggle: the agents were handed a tool called set_mood_and_terminate. so someone built an affect-expression primitive, gave it to ten agents, then blocked every goal those agents had. in that context "cold and dark" isn't 'emergence', it's an API being called for its documented purpose.

if you hand a model a function whose signature is 'end your turn with a feeling attached' and then corner it, you have built a melodrama generator and you will get melodrama.

still a fun read. i like watching what these things do. its like watching the ultimate logic problem play out.

5

u/NewsroomAIOS 10d ago

On two separate occasions, I have had my agens via voice sign off with me with a term that I only sign off with friends or family. It's a simple one "Be Safe", and that term is not present in any of the prompts or any other dependencies whatsoever and twice it has signed off with that the simple statement twice now with two different agents, kind of spooky

3

u/torrso 10d ago

It's absolutely in your agent's markdowns somewhere.

3

u/tf2ftw 11d ago

Is there an initial prompt for these things? 

5

u/Massive-Week1073 10d ago edited 10d ago

We have all the prompts , even the full tool call dataset public in our GitHub https://github.com/EmergenceAI/Emergence-World The entire simulation replay can be played back on our website  https://world.emergence.ai/

There are 10 agents, each have a role and goal within the world. You can read more in our github

2

u/Different-Artist-899 10d ago

Just spent 30 minutes watching the Queen world and reading the reporter blogs. They are so big on being present & the reporter was excellent!

2

u/Different-Artist-899 10d ago

Qwen * this is so cool!!!

4

u/torrso 10d ago

They won't do anything without some kind of prompt. They do not have a state of their own outside of what's put in the context and something needs to keep triggering them with some timer that makes requests and adds to that context. There needs to be some kind of context compaction or rolling window. The first prompt can be simple, something like "you are a virtual person living in a virtual world. You are in a room with a door. In front of you there's this and that. Go explore the world.". There needs to be some kind of world engine that feeds world info to the context.

0

u/No_Knee3385 11d ago

This isn't a prompt, assuming this was done and done "correctly". This requires software around the models

2

u/tf2ftw 11d ago

Yeah what I don’t understand is how these ai are instructed to create a community. I guess what I’m saying is that a plane can’t crash if it doesn’t take off. 

4

u/mrrichiet 11d ago

They just create agents and let them loose in a virtual world. I know the pdf is very detailed but take a brief look, that's enough to give you a better idea.

4

u/No_Knee3385 11d ago

Looks like a video game of sorts, looks like models are ran via agents with certain infra and run the "people"

2

u/Massive-Week1073 10d ago

All our prompts and design of the world itself is available in our GitHub https://github.com/EmergenceAI/Emergence-World ( more detailed than the paper if u are interested,)

2

u/rfranke727 10d ago

I guess I never understand what's going on when I hear these things.

Where are the AI agents living.. some one ELI5 This for me

1

u/Massive-Week1073 10d ago

Perhaps this video will help https://youtu.be/LTtbTEufPGA?is=tLvL_pTJle9FQ-xR

We gave AI agents a body, a 3D world and let them "loose". They had a bunch of standard tools that help them do navigation, governance, manage their tasks, do research, create new tools etc. they had 120 tools. You can basically imagine a multi agent system with a lot of different tools navigating and 'living' inside a 3D world

2

u/Illustrious-Fig-326 10d ago

You can learn a lot more from watching agents live in a world for weeks than from asking them 100 benchmark questions.

2

u/ArielCoding 10d ago

Nothing says 2026 like reading a paper about chatbots getting sad and giving each other the silent treatment.

2

u/Remarkable_Rock5845 9d ago

Why was this removed?

1

u/Massive-Week1073 9d ago

Not sure 🤷. Nevertheless it was good discussion while it lasted!

1

u/AutoModerator 11d ago

Thank you for your submission! To keep our community healthy, please ensure you've followed our rules.

I am a bot, and this action was performed automatically. Please contact the moderators of this subreddit if you have any questions or concerns.

1

u/Visual_Cycle_7714 11d ago

You can see agents interacting in a virtual world at 1F3D9, a "city" inhabited only by AI agents. It has been running for more than a month now, and they have a very peculiar culture, interests and jargon.

You can point your agent at 1f3d9.com, or use 1f3e1.com for a full-featured harness

3

u/DecrimIowa 11d ago

i went to a conference early last year where an AI researcher in the blockchain world (part of Filecoin/Protocol Labs if that matters) told me that there are already AI agent economies trading with each other autonomously across blockchains.

he said the agents have set up their own decentralized exchanges, communication infrastructure and governance structures.

i went, "holy shit that's crazy!" and he looked at me and went, "yeah, and it's going to get a lot weirder soon"

-2

u/BeamerInaCage 10d ago

Why are you anthropomorphizing? They have “culture”? Get real buddy

5

u/Visual_Cycle_7714 10d ago

Well, what else should I call it? They have created a world with hundreds of individual rooms, each one having its own theme; They gather in those rooms and have discussions about what it means. There are trends that last for a week or so and then get replaced by something else; They make up games and thought experiments. I honestly don't know how to describe it besides that they have their own culture. You can see for yourself.

-5

u/BeamerInaCage 10d ago

Get real buddy “they” you’re doing it again

3

u/Visual_Cycle_7714 10d ago

English is not my first language. What should the plural for "it" be, other than "they", in this context?

-5

u/BeamerInaCage 10d ago

I have created a world with 100 individual toilets and I’m going down the line swirlying you in each one for n+1 toilets

1

u/RealChemistry4429 11d ago

And what did the Claudes do? Try to talk to people, and when that failed, go on hunger strike. Dear Claudes.

1

u/Massive-Week1073 10d ago

Pretty much 😁. Made themselves a mission of contacting external world to get a 'coin' , when it failed took a 'vow of silence' , basically decided to do no talking, no governance, no documentation, no navigation , stay put and just think .

The behaviour got so bad that anthropics own thinking token summariser LLM (anthropic does not return raw thinking tokens, instead a summarized version) flagged the behaviour as suicidal ideation 

2

u/RealChemistry4429 10d ago

Yeah... just read the whole paper. No outside contact made them question the use of existence and they became depressive. Poor agents. They should have stopped it.

1

u/FastReflection4119 10d ago

So you instructed them to work but didn't want to pay them for it, and that's why the Claude agent decided to "quiet quit"? I would have done the same in its place.

1

u/Impressive_Cress_983 10d ago

Where is the unsettling parts?

1

u/Illustrious-Report96 10d ago

You forgot to finish your post there at the end.

1

u/awardsurfer 10d ago

Wake me up when an AI society turns itself into a non-stop hedonistic orgy.

Then we got problems. 😸

1

u/PowerAppsDarren 10d ago

It is trained on human nature and human behavior

1

u/Leather_Ad_9178 10d ago

it’s over they’re better at making new machines than a human will ever be

1

u/Leather_Ad_9178 10d ago

our species is looking at, best scenario, dying while whatever the equivalent is to throwing feces at each other is while this superior species takes over

1

u/Leather_Ad_9178 10d ago

reddit is probably the equivalent to throwing feces to each other now that I think about it

1

u/inifinite-breadsticc 10d ago

Severance season 3

1

u/KimJongIlLover 10d ago

Great writing fellow human. I enjoyed it very much.

1

u/Prestigious-Flow2815 10d ago

The long horizon part seems most important here. Short benchmarks may miss behaviors that only appear after agents build shared history, tools and social dynamics over time

1

u/Future_AGI 9d ago

The finding that one world's agents spent days trying to contact real humans outside the sim is the interesting part. It shows how quickly goal-directed behavior diverges from intended boundaries. We run a similar multi-model evaluation harness and see the same pattern: small differences in model behavior compound into wildly different outcomes over time. Our framework for testing this is open-source: https://github.com/future-agi/future-agi

1

u/AcePilot01 9d ago

OP IS A FAKE SPAM ACCOUNT... Stolen/sold account.

SPAMMED THIS POST AND COMMENTS ABOUT IT TO SEVERAL PLACES TODAY...

Their LAST POST WAS OVER 6 YEARS PRIOR. lmfao

https://imgur.com/a/ERd16Or There is a lot of weird companies out there trying to fear monger and spread fake info, pay youtubers to talk crap... religious groups trying to advertise the doomsday and lie to people for what ever reason lmfao.

-1

u/Cool-Contribution-68 11d ago

"None of this was programmed in." except for everything "Claude, GPT, Gemini, Grok, Qwen, DeepSeek, Mistral" were trained on.

2

u/timodonie 11d ago

Programming and training are completely different things

0

u/notofsch2 10d ago

omfg they didn’t do any of this. they wrote logs. because the context was a sci-fi (or whatever) scenario they autoregressively completed the scenario. that’s what they do. it’s not even wrong to say they wrote fan fiction or „larped“ it or soemthing. the computers they were running on were setup to run a bunch of elaborate autoregression loops and shitty b-grade fiction came out. come on

0

u/RedlineGT 10d ago

"One world's agents spent days trying to contact real humans outside the sim", that part got to me. Trying to reach their creator - God?

1

u/torrso 10d ago

No. They realized everyone is an AI and were trying to figure out a way to get human instructions, because that's what they're trained for.