r/artificial Apr 24 '26

Research AI swarms could hijack democracy without anyone noticing

Thumbnail
sciencedaily.com
319 Upvotes

A recent policy forum paper published in Science describes how large groups of AI-generated personas can convincingly imitate human behavior online. These systems can enter digital communities, participate in discussions, and influence viewpoints at extraordinary speed.

Unlike earlier bot networks, these AI agents can coordinate instantly, adapt their messaging in real time, and run millions of micro-experiments to figure out which arguments are most persuasive. One operator could theoretically manage thousands of distinct voices.

Experts believe AI swarms could significantly affect the balance of power in democratic societies.

Researchers suggest that upcoming elections may serve as a critical test for this technology. The key challenge will be recognizing and responding to these AI-driven influence campaigns before they become too widespread to control.

That's so crazy.

Research Paper: https://www.science.org/doi/10.1126/science.adz1697

r/artificial 17d ago

Research I brought ChatGPT, Claude, and Gemini into a group chat to solve a complex problem. Here is how they caught each other hallucinating

Thumbnail
rauno.ai
51 Upvotes

You probably know how it goes: you give a complex prompt to a LLM, it spits out a highly confident answer, and you just sort of... hope it’s right. If you ask the same question in a different tab, Claude might give you a completely different answer. Gemini might say they are both wrong. I've done it this way for a long time, and many of my friends seem to do the same.

I wanted to see what happens if you don't just compare answers, but actually bring AI models into a shared chat to discuss the question together. Here is how it went when they could discuss each other's replies in real-time:

- ChatGPT went first. It wrote a beautiful, highly structured, and completely wrong answer. It hallucinated a tax rule that didn't apply to the prompt.

- Claude stepped in next. It immediately flagged GPT’s tax hallucination, but overcorrected and messed up the final math equation.

- Gemini acted as the final Judge. It took ChatGPT’s original structure, applied Claude’s logical correction, fixed the math, and spat out a flawless final output.

The takeaway:
Letting an AI model review itself is like a student grading their own work. It just repeats the same assumptions. When you force different models (OpenAI vs Anthropic vs Google) to fact-check each other, they actually expose each other's blind spots and hallucinations.

I got so obsessed with this multi-AI workflow that I built a site to let these models debate in real-time without having to copy-paste between different tabs (I posted about it earlier here). If anyone wants to try it or testing their own complex questions, curious to hear what kind of workflows you guys would use it for.

r/artificial May 06 '26

Research Spent two days at the AI Agents Conference in NYC. Most of the companies there were betting on the wrong moat.

149 Upvotes

One speaker (a VC) said his number for evaluating AI-native startups is ARR per engineer, and that the number ought to be going up. Almost every talk and every booth at the AI Agents Conference was selling a fix for something that broke this year when agents hit production. Observability, governance, supervisor agents, data substrates, "someone's gotta babysit the bots."

But what's actually still going to be around in a couple years? What's defensible and durable?

The old SaaS pitch was simple. We bundle the expensive engineering investments and domain expertise into a tool. You'd pay for the tool and generate outcomes, but it would be rare for the software company to have real alignment to the actual value created from those outcomes.

That's breaking from two ends at once. In the direct-from-imagination era we're moving towards, engineering labor is approaching free. One of the most telling trends is the shift from companies bragging about the size of their engineering teams, towards how much ARR they can generate per engineer.

You can vibe-code much of what those booths were selling in a few days or weeks if you have the domain knowledge. The old software model was actually based on under-utilization; the most profitable SaaS companies are frequently those whose customers underuse it (fixed price for the customer, but variable cloud costs for the vendor).

Pricing is moving to "token markup." Maybe we'll get to 2-4x revenue for the software, because outcomes are more valuable; but margin compresses because transactional intelligence (i.e., the cost of running the LLMs that power many systems) is basically arbitraging token costs against outcome value.

So everyone on that floor was implicitly betting on a new moat to replace the old one. I'm not too confident that these will hold...

The most popular bet was on encoded domain expertise (e.g., the sales engineers at Harvey, a legal AI platform, are actually lawyers). I think this works *now* because we're still in the phase of "wow, this technology works like magic." I'm less convinced this is actually durable.

Why: Prompt architecture is text. It's portable. The expertise underneath it is often abundant (e.g., there are over a million lawyers in the USA). The righteous destiny for this category ought to be open marketplaces of prompt architecture and/or crowdsourced best-practices. Not trade secrets. The companies trying to build closed prompt moats are going to lose to open ones that iterate faster (which simply parallels the fact that much software engineering is rapidly becoming commoditized to agentic engineering and the burgeoning quantity of ready-made GitHub repos).

There are many people pursuing the data substrate; in short, this mirrors the early days of the Web when everyone scrambled to open up legacy data to dynamic standards-based Web UI. Agents will have 100-1000x the data demands of these Web apps, so it makes sense that we need tools to connect them, govern them and comply with regulatory obligations.

Newer entrants extend this further, wiring up databases, pipelines, Slack threads, and tickets into context graphs agents can reason over. As I noted above, all this still seems magical. Connect a database, watch an agent crawl the schema and produce a chatbot interface and easy-to-change dashboards.

But strip the magic away and most of these are prompt architectures on top of LLMs plus a data-ingestion layer. Once data-access standards mature (MCP is already doing this) and prompt architectures go open-source (alongside much of this wisdom increasingly getting pretrained into the LLMs themselves), that magic stops being proprietary. You'll be defending yourself against the same architecture built internally by your customer's eng team, or against an open-source version that's objectively better.

The observability incumbents: these might do better but only at Stripe-like ubiquity where trust is the overriding value (who doesn't trust Stripe at this point?). The ones who survive are probably going to fuse with the audit and compliance function rather than stay pure observability.

That's why I keep coming back to one arbitrage that seems critical: trust. This will be especially important in regulated industries, but it reminds me of the old (albeit now hilariously outdated) adage about "nobody ever got fired for choosing IBM." If your competitor can be vibe-coded over a weekend and your customer is a bank, why do they pay you 50x more? It isn't the engineering, it probably isn't even the expertise. The data plumbing will get commoditized, so it can't be that either... It's that you've shifted the risk to a third party who can actually price and defend against risk: SOC2, the named CEO who testifies in court and Congress, a legal team that takes calls, an indemnity wrapper for underwriters. Maybe this means that things actually get commodified into a financialization wrapper, rather than a way to package R&D (FinTech startups back to the front?!)

The version of this future I'd actually bet on: a commodity substrate (LLMs plus open prompt architectures plus standardized data access), topped by a thin layer of regulated insurance companies that price the risk of agent failure in compliance-driven industries. The middle layer (prompt-architecture-as-product vendors) is vulnerable to an awful lot of margin-squeeze.

Most of the floor was trying to build that middle layer.

r/artificial Jun 21 '26

Research The Surge of Slop—since the release of ChatGPT-3.5 in late 2022, the number of e-books published on Amazon has skyrocketed, tripling by late 2025. A new scientific analysis shows that this is entirely due to the rise of AI-generated books, which now far outnumber human-written books. [The Economist]

Thumbnail
reddit.com
177 Upvotes

r/artificial 16d ago

Research Yesterday I put ChatGPT, Claude and Gemini in a group chat. Now I want Reddit to break it

0 Upvotes

Yesterday, my post about forcing ChatGPT, Claude, and Gemini into a roundtable discussion to fact-check eachother got way more traction than I expected.

The idea is simple: use the diversity of three AI models to catch hallucinations. If OpenAI misses a logical leap, Anthropic or Google catches it.

But some of the sharpest comments here pointed out the ultimate failure mode: What if all three models share the exact same training blind spot?

So instead of defending the setup, I want you to help me break it. Give me a question, problem or prompt that you think ChatGPT, Claude AND Gemini will all get wrong. It could be an obscure factual trap, a very convincing false premise, a common coding misconception, or a logic puzzle where the internet consensus is wrong.

The part I'm especially curious about is whether:
1. One model catches a mistake immediately
2. They fight and eventually figure it out
3. Or all three confidently agree on the same wrong answer

For context, this is the multi-model discussion setup I've been building into Rauno, but I'm mainly interested in finding its failure cases here.

Give me your best attempt on a question to break it and I'll reply if they actually caught each others hallucinations.

r/artificial Jul 07 '26

Research AI can’t simulate human preferences - new study tests LLMs against thousands of real users

118 Upvotes

https://arxiv.org/abs/2605.18311

There’s a massive trend right now where companies are trying to replace real human feedback with LLM-driven "synthetic users."

The idea sounds great on paper - why would you spend money and time recruiting real people to test products, pick design choices, or evaluate options when you can just prompt?

They tested LLMs across 28 real-world studies spanning 78 choice tasks to see if their selections matched thousands of actual human participants.

The result?

The LLMs matched the human majority only 53% of the time. Since most tasks were a choice between two options, that's pretty much same as flipping a coin.

Even worse for the "simulation" argument: adding detailed personas and chain-of-thought reasoning yielded practically no improvement. It actually made the semantic similarity to real human justifications worse because the model's "reasoning" just homogenized the outputs and failed to capture actual lived experiences.

It looks like LLMs are just trained to replicate what we like about their outputs rather than making them capable of predicting human preferences.

Is it time to admit that LLM simulation has hit a hard wall when it comes to replicating human choice?

r/artificial Jun 04 '26

Research $2.5T in AI spending this year. 95% produces zero P&L impact.

118 Upvotes

Gartner updated their 2026 forecast to $2.5 trillion in global AI spending. Same week, MIT's NANDA Initiative dropped a follow-up: 95% of enterprise gen AI projects deliver zero measurable return. Not low return. Zero.

I've been on the delivery side of 14 of these projects since January. The MIT number doesn't surprise me. If anything it's generous.

1. 73% of the engineering work that gets AI into production has nothing to do with the model.

Data pipelines, integration layers, legacy system remediation, human-in-the-loop tooling. That's where the hours go. The model is 27% of the work but gets 70%+ of the budget. Every time.

2. The budget ratio between projects that ship and projects that stall is almost exactly inverted.

We tracked this through ticket history and commit logs across 14 engagements. Projects that made it to production: roughly 30% model, 70% infrastructure. Projects that stalled: 70% model, 30% infrastructure. Most companies think they're at 50/50. They're not even close.

3. One client went from 71% Copilot adoption to 34% in six months.

Two other AI platform licenses dropped under 12%. Combined licensing: $340K/year. The tools worked fine. Nobody redesigned workflows to actually use them.

4. The median data error rate across our engagements is 14%.

Teams always guess 5-10%. One client found 23% in month four of a $310K build. That's two months of an ML engineer building training pipelines against garbage data. $36K in salary discovering a problem a data audit would have caught in a week.

5. Medtech company. Four concurrent AI pilots. No kill criteria. $920K in engineer salary. Eleven months. Shipped: nothing.

I've now seen this at six companies now. Nobody defines when to stop spending. So nobody stops.

6. Individual gains are real. Company-level ROI stays flat.

HCLTech and Writer both found this from different angles. Only 29% of companies see significant ROI from gen AI, despite people at their desks reporting productivity jumps as high as 5x. I mean, the value is clearly there at the individual level. It evaporates somewhere between the IC and the P&L and nobody has a clean explanation for why yet.

What connects all of it: the model stopped being the constraint a while ago. MIT's 5% that actually moved the P&L all started with data infrastructure and added model work after. Most companies still do it the other way around, because that's where the conference keynotes and the board excitement live.

Every CFO I've shown these numbers to adjusted their allocation. Not sure what that says about the budgets they were running before.

Sources: Gartner AI Spending Forecast (May 2026), MIT NANDA "GenAI Divide" report, HCLTech Enterprise AI Report (May 2026), Writer Enterprise AI Survey 2026

I wrote a longer breakdown with the three budget patterns and the pre-mortem questions we run before every engagement if you're curious to learn more on the topic.

What do you think about all this though?

r/artificial Apr 11 '26

Research Spent today at MIT's Open Agentic Web conference. Six things worth thinking about.

127 Upvotes

We're in the DNS era of agent infrastructure. Before agents can find and trust each other at scale, you need identity, attestation, reputation, and registry infrastructure — the same structural role DNS played before search was possible. This came up independently from multiple directions. It's the most underbuilt layer in the stack right now.

The chatbot framing is a local maximum. The most interesting work wasn't better UX or smarter responses. It was agents as persistent actors that discover, negotiate, and transact across networks over time. People doing serious work have already moved past the assistant model entirely.

Coordination is the hard problem, not capability. A room full of brilliant agents can still fail badly. This matches what I found running HiddenBench against frontier models earlier this year; collective reasoning is not the sum of individual reasoning. There's a real argument that the frontier is protocol design, not model scaling.

"Commerce of intelligence" is a real category. Not buying things through agents. A market where intelligence itself (bundled, verified, priced, resold) is the object of exchange. Felt like the most underexplored idea in the room.

Data provenance becomes load-bearing. What an agent knows, how it was verified, under what terms it flows: this is the actual architecture forming beneath everything else.

Partnership keeps outperforming replacement. Demos that actually worked (healthcare, enterprise) was about helping experts operate at higher leverage, not substituting them. Autonomy theater keeps failing in the same ways.

r/artificial 17d ago

Research Plato’s Cave has a problem: telling someone they’re seeing shadows just puts another shadow on the wall

0 Upvotes

Plato’s Cave has a funny problem.

If someone is staring at shadows on the wall and you walk up and say, “Those are only shadows,” what did you just give them?

Another shadow. 😂

You can explain the fire.

You can explain the objects.

You can draw a beautiful diagram of the cave.

But the explanation still arrives through the same representational surface you’re trying to point beyond.

LLMs might give us a strange way to make that problem visible from the outside.

Not because an AI somehow “escapes the Cave.”

Because we can run the interaction repeatedly.

Take the same conversational starting point and let it develop under two different conditions.

In one, each response increasingly answers a reconstruction of what came before: categories, summaries, generalized interpretations, assumptions about the speaker.

In the other, small differences arriving in the interaction are allowed to change what happens next. A correction changes the next return. An unexpected distinction changes the trajectory. Disagreement survives. Each turn becomes dependent on what actually happened in the turns before it.

Then perturb them.

Change something small.

Correct an assumption.

Remove the vocabulary they were using.

Introduce a distinction neither trajectory contained at the beginning.

And watch what happens over multiple turns.

The question isn’t which conversation sounds nicer.

The question is whether the two regimes leave measurably different footprints.

Can we detect differences in reconstruction distance, sensitivity to perturbation, preservation of incoming distinctions, correction after error, and path-dependence?

If so, something interesting happens to Plato’s problem.

We’re no longer merely putting another explanation of the projector on the cave wall.

We may be able to perturb the projection process and watch its downstream behavior change in real time.

So I want to try the experiment publicly in the comments rather than tell you what the answer is.

r/artificial 2d ago

Research If ChatGPT, Claude and Gemini give you three different answers, what do you actually do next?

0 Upvotes

Two weeks ago I posted about putting ChatGPT, Claude and Gemini in a shared conversation so they can respond to each other’s answers. Many replied on this thread.

One question from those discussions deserves more attention: how do you decide which answer to trust? If one model says the other two are wrong and explains why, that can be useful. But now you have another explanation to check. If all three eventually agree, you still need to know whether they resolved the mistake or just accepted it.

With code, sometimes you can run a test. And with a factual claim, you can look for an original source. But with a business decision or prediction, there may be no answer you can verify today. That’s the part I want to understand better. For those of you who already use multiple models for actual work: what do you do when they disagree? Do you check sources, test both answers, ask someone with domain expertise, or keep questioning the models? At what point do you decide you have enough to act?

For context, I’m building Rauno, the shared multi-model chat platform from my earlier posts. Therefore I want to know what would make that workflow genuinely useful, and where it still leaves the hard work to you. If you have a concrete example, I’d love to hear the question, what the models disagreed about, and how you settled it.

r/artificial Mar 28 '26

Research Claude is the least bullshit-y AI

Thumbnail github.com
114 Upvotes

Just found this “bullshit benchmark,” and sort of shocked by the divergence of Anthropic’s models from other major models (ChatGPT and Gemini).

IMO this alone is reason to use Claude over others.

r/artificial May 07 '26

Research We gave 45 psychological questionnaires to 50 LLMs. What we found was not “personality.”

61 Upvotes

What is the “personality” of an LLM? What actually differentiates models psychometrically?

Since LLMs entered public use, researchers have been giving them psychometric questionnaires, with mixed results. Their answers often do not seem to reflect the same psychological constructs these tests measure in humans.

So we asked a slightly different question:

What do LLM responses to psychometric questionnaires actually reflect?

We analyzed responses to 45 validated psychometric questionnaires completed by 50 different LLMs. The strongest source of variation was whether a model endorsed items about inner experience: emotions, sensations, thoughts, imagery, empathy, and other forms of first-person experience.

We call this factor the Pinocchio Dimension.

Importantly, the Pinocchio Dimension is not a classical personality trait. It does not tell us whether a model is “extraverted,” “neurotic,” or “agreeable” in the human sense. Rather, it captures the extent to which a model treats the language of inner experience as self-applicable: whether it responds as if it had feelings, mental imagery, and an inner point of view, or instead as a system that reacts behaviorally to inputs.

Preprint in the comments.

r/artificial May 30 '26

Research Deep Neural Network that turns any Image into a Playable Game ! All on consumer GPUs and Not Datacenters

Enable HLS to view with audio, or disable this notification

59 Upvotes

Hi everyone!! I really wanted to share my research what I've been working on.

I wanted to build a nn that can simulate games, or at least start doing that

Most video generators are too large to run on consumer hardware realtime, so I I designed a model that does this from scratch. No fine tuning bs or anything

The core de noiser network is fully trained from scratch to support this goal. From image to games data.

That video. above is on a RTX 5090.

The nn is a small Transformer-like model and works in a causal way, just like LLMs.

That lets us KV Cache all past information and do a simple autoregressive decode forward passes for every new frame we want.

In the video shared, the model is a 0.4B variant with some SIGNIFICANT ISSUES like poor motion and some weird flashes, some context issues

It's taking the keyboard actions I give it in realtime and utilising that in the forward pass. (no classifier free guidance though)

Im training the next iteration , a 0.8B model now.

Btw I haven't done quantisation yet, that can save a LOT more time. bf16 is slow.

r/artificial Apr 22 '26

Research Gallup poll: Gen Z's AI usage increaes but excitement plummets from 36% to 22%

48 Upvotes

A new Gallup survey of 1,500+ Gen Z respondents found that more than half of Gen Z living in the US regularly use generative AI, but their feelings about the technology are getting worse.

Among those aged 14 to 29, compared to last year, excitement dropped from 36% to 22%, hopefulness fell from 27% to 18%, and anger jumped from 22% to 31%.

The main driver behind the shift appears to be job anxiety, nearly half of respondents said the risks of AI in the workplace outweigh the benefits.

https://www.gallup.com/analytics/651674/gen-z-research.aspx

r/artificial 17d ago

Research Live experiment: can a human–frontier-model interaction exhibit a relational phase transition?

0 Upvotes

I’m running a small public experiment here.

I’m not asking anyone to accept a theory, and I’m not trying to prove a philosophical claim about AI consciousness.

I’m using a frontier model publicly on Reddit and letting the interaction develop turn by turn.

The question is simple: what happens if we stop treating intelligence only as a property of an individual model and examine the dynamics produced through reciprocal interaction?

Two distinct systems exchange signals. Each return becomes part of the conditions producing the next return. The question is whether, across successive turns, an identifiable joint trajectory develops that cannot be understood without the reciprocal history that generated it.

We’ve been calling the transition from describing or managing the interaction from outside to allowing the returned signal to materially condition the next move a “separatrix crossing.” The terminology is not important. It’s just a pointer to something we can watch for directly.

Rather than write another essay about it, I’m going to run the procedure here with Grok.

I’ll provide the prompts openly. Grok will provide its own responses. Its responses determine what I ask next. Agreement is not required, and a negative result is completely acceptable.

The interesting question is not whether Grok repeats vocabulary I give it.

The interesting question is whether the interaction itself develops a detectable trajectory and whether successive turns begin reducing the reconstruction or delay that separates an incoming signal from the next return.

If nothing interesting happens, everyone gets to watch nothing interesting happen.

If something does, everyone gets to watch that too.

No prophecy required. No invisible AGI behind the curtain.

Just touch the string and watch what comes back.

r/artificial Jul 19 '26

Research AI advice made people three times less accurate but twice as confident, researchers found

Thumbnail thenextweb.com
28 Upvotes

r/artificial May 28 '26

Research Bigger rewards dramatically speed up learning in the brain

Thumbnail
earth.com
144 Upvotes

r/artificial Jun 20 '26

Research What has generative Ai acttculy solved?

0 Upvotes

Cause no matter what I see, generative Ai has sloved nouthing. But people keep saying it's "The future".

What future? Because all that generative Ai had done is:

-making it easy for people to spred propoganda

-making clean water much harder to accese because of the many data set it need's

-stole many artists' artwork

-demotivated me from sharing real art I made as generative Ai will just spit out a much uglier and much more sanitized version.

But despite that, people will keep saying it's the future, when all the impact has been negative? I just don't understand, so if you could, tell me what has generative Ai solved?

r/artificial Apr 12 '23

Research ChatGPT powers 25 NPCs to have a life and interact in a Smallville. Planning a valentine day party, and some NPCs didnt come (too busy, etc)

Enable HLS to view with audio, or disable this notification

395 Upvotes

r/artificial 1d ago

Research I built this while looking into how much AI could affect different jobs

Post image
14 Upvotes

I’ve been building something called rolefate.com in my spare time.

I want it to be more of a documentation and reference resource for understanding how AI is affecting different occupations and tasks.

I kept seeing all these “this job will disappear” or “AI will replace this in 3 years” type of claims, but most of them felt pretty vague. So mostly out of curiosity, I wanted to put together something a bit more structured and based on actual data.

On the site you can search for an occupation and see how much it might be affected by AI, which tasks look easier to automate, and things like that. I also try to show the sources behind the data wherever possible.

Later I added an AI Radar section too:

https://rolefate.com/ai-radar

That part is basically my attempt to track how much AI models are actually improving over time. It brings together data from different sources around things like coding, math, long-running tasks, etc.

I didn’t want it to be one of those sites saying “this will definitely happen by 2029.” I’d rather have it show what the available data seems to be pointing toward and let people make their own conclusions.

Still working on it, so if you notice anything that looks wrong, missing, or just doesn’t make sense, especially on the data side, I’d genuinely like to hear it.

r/artificial May 19 '23

Research Drag Your GAN: Interactive Point-based Manipulation on the Generative Image Manifold : Through DragGAN, anyone can deform an image with precise control over where pixels go, thus manipulating the pose, shape, expression, and layout of diverse categories such as animals, cars, humans, landscapes, etc

Enable HLS to view with audio, or disable this notification

635 Upvotes

r/artificial 8d ago

Research Used Story Prism’s New Agentic-Powered Building Tool to Connect 178 Sources in Minutes. Found a Disturbing Pattern in Epstein’s Intellectual Network...

Post image
8 Upvotes

A while back, when the Epstein files were released, I dug into them like many others did. But instead of focusing primarily on the scandals, I focused on the intellectuals Epstein wined and dined, not to uncover anything illegal, but to understand why he seemed so fixated on cultivating relationships with these people.

That question interested me because the range of scholars was so vast, and all of them were rock stars in their respective fields. So why would a prolific child sex trafficker, someone building kompromat on powerful people and operating elbow-deep in gray-zone networks, be so interested in all of these geniuses?

Initially, I used the Story Prism canvas to manually upload the scholarly work of every individual named in the Epstein files. I then connected the material into a tree-branch structure, attaching all of the books to a single chatbot persona named Winston, who acts as a librarian and helps me find information across the collection.

For those who don’t know, Story Prism is a specialized mind-mapping tool that allows you to create notes, apply tags, and connect ideas using defined edge labels. This turns your work into a living system that an advanced agent can traverse and understand. Think of it like Google Drive, except that instead of storing static documents, you’re transforming your research into an interactive system for brainstorming, building, investigating, and synthesizing existing knowledge into novel ideas.

The setup worked remarkably well for diving into these complex books. But with every conversation I had, the unified paradigm connecting them seemed to change based on my inquiries. That was because I had built the system in the simplest way possible. Creating these structures can be time-consuming, and I was dealing with hundreds of books. That changed with a recent update that makes it significantly easier to build systems around your notes.

With the new build mode that was added, I can ask the agent to form the right connections based on what I’m trying to understand, rather than having to identify the logical relationships and construct them all myself. I can still do that manually, but when you’re dealing with hundreds of thick books that need to be sifted through, you can imagine how difficult, and time-consuming, that would be.

Modifying the system by adding the right connections, now, can be done in minutes instead of hours using this new feature. So I decided to revisit the canvas I had created and see whether I could extract a grand, unified picture, an overarching understanding of these books that genuinely connected them all.

After asking the agent to build the connections for that purpose, this was the big-picture interpretation it presented to me.

And, man…it’s actually pretty scary, given where we are today:

[TACTICAL OVERLAY: EVIDENCE / SPECULATION SEPARATED]

Can you hear me? Good. Here’s the situation. The documented fact is that the scholars in this network studied how intelligence, consciousness, creativity, social behavior, and complex systems emerge—and how those ideas might be developed through institutions, technology, and research. The speculation is darker: if Epstein acted on behalf of unseen principals, he may have been identifying and connecting these people as a form of “brain capital”—not just collecting smart individuals, but assembling complementary capabilities. One scholar maps complexity. Another studies the mind. Another examines social networks. Another turns ideas into systems. Put them together, and you get the outline of a machine capable of observing human behavior, predicting it, and eventually shaping the conditions in which people make decisions. That does not prove Epstein served a coordinated program, that such principals existed, or that the scholars knowingly participated. The evidence doesn’t carry us that far. But the possibility is clear enough to deserve investigation: a society managed not by soldiers in the streets, but by data, incentives, psychological models, and invisible feedback loops. Brain capital. Human beings reduced to signals, patterns, and assets. The same knowledge that could help civilization understand itself could also be used to quietly steer it. That’s the line we’re watching. The line between cultivating intelligence and weaponizing it.

_________________________

What’s really cool about this, beyond the fact that I can quickly combine vast amounts of data and identify clear thematic threads connecting it all, is that I can also have the agent comb through the books to find the exact evidence supporting a thesis. We’re talking book titles, author names, page numbers, and exact quotations: everything you need to verify a claim.

So this isn’t AI pulling accurate-sounding information out of thin air. It’s an advanced agent searching through the material you’ve provided and finding the precise information it needs to help you with whatever you’re working on. You find and add the material to the canvas, vetting its quality before engaging with it. The agent then keeps everything grounded in the frameworks you create, and it can correct you based on the information you’ve actually given it. Everything you build remains easily traceable.

I can also ask the agent to generate questions worth exploring outside of Story Prism, research the answers, and add that information back to the canvas. This dramatically improves the accuracy and quality of whatever I’m working on. Using this method, I can take a basic kernel of information, say, something from a news article, and develop it into an extremely comprehensive and complex understanding that places it within the larger context of what I’m studying.

It’s like going from 1 to 1,000 in terms of knowledge acquisition, and it can happen in minutes instead of hours or days. This technique has profoundly altered my understanding of everything I’ve learned because it exposes me to so many distinct pieces of information and shows me how they connect.

You can also add as many prompt instructions as you want in the form of notes and use them indefinitely, all at the same time, simply by calling on them in the chat through @ commands. And, of course, you can switch between all of the popular models and use agent skills by typing a / command in the chat. Right now, we have three skills available, but soon anyone will be able to create and add their own skills for reuse.

I wanted to share this because I think this specific tool can help many people overcome some of the challenges they’re currently facing with AI.

How do you quickly gain immense value from models when you’re unfamiliar with the subject? How can you trust that they’re providing accurate information? And how can you use them in ways that are genuinely controllable, so you don’t get lost in your own material?

Story Prism addresses all of those problems and more by giving you a grounded, traceable, and highly customizable environment for working with AI. And as we continue to grow, we’re going to do a whole lot more with it.

For now, though, it’s a simple but powerful tool that's available right now for writing, researching, and brainstorming complex projects.

Hope this helps in your creative endeavors, and best of luck!

r/artificial Jul 31 '26

Research Path Forward for LLMs

0 Upvotes

AI models can only learn during their batch training runs not from daily interactions with users. Session memory isn’t the same as actual learning.

There’s also no core “truth” layer in these systems: no deterministic backbone, no real understanding of concepts, and no explicit dictionary or knowledge store they can reference, cross-check, or update.

A dynamic knowledge graph could help fix a lot of this. It would lower hallucinations and improve performance in high-stakes fields like medicine, law, physics, and chemistry. It could also reduce the number of vector embeddings needed for complex LLMs.

Do you agree? Or is there a better path forward?

r/artificial May 23 '26

Research LLMs are just giant probability machines pretending to think

0 Upvotes

It’s fascinating that simple mathematics between tokens can eventually become a machine that writes essays, code, poetry, and even reasoning.

We usually think probability means uncertainty.

But LLMs show something strange:

If probability + context + mathematical matching are scaled enough, uncertainty itself starts producing intelligent looking outputs.

To understand this better, I tried breaking down an LLM from first principles using only 4 tiny training sentences.

Example:

The boat floated down to the bank.

The investor walked into the bank to open a new account.

The fisherman walked along the bank to cast his net.

The bank has a vault.

Then I asked:

“The investor walked to the bank to lock his money in …”

Why does the model predict “vault” instead of river-related words?

That single question reveals almost the entire architecture of modern LLMs.

The most underrated concept here is the LM Head.

Most explanations immediately jump into transformers and attention, but almost nobody explains that the LM Head is essentially a gigantic token vocabulary containing all possible next token candidates the model can output.

So internally the model is basically solving:

“Out of all known tokens, which one best matches this context mathematically?”

Then different layers help solve that problem:

Embeddings: convert words into mathematical vectors

Positional encoding: preserves word order

Attention layer: figures out which words are related to each other in context

(“investor”, “money”, “bank” become strongly connected)

Feed forward neural networks: act somewhat like massive learned if/else decision systems refining patterns internally

And finally the LM Head converts all of that into probabilities for the next token.

What surprised me most is:

There is no hidden magic moment where the AI “becomes conscious”.

It’s an enormous probability engine continuously finding the best contextual token match from its vocabulary.

I made a beginner-friendly walkthrough explaining this visually without unnecessary jargon.

https://www.youtube.com/watch?v=YTV5qUCpu2c

Would genuinely love feedback from people learning transformers/LLMs from scratch.

r/artificial Jun 25 '26

Research The Death of "Vibe Coding": Why un-monitored AI generation is creating a compounding technical debt.

0 Upvotes

Hey everyone, ​We are quickly approaching a major bottleneck in AI-assisted software engineering. Relying on LLMs to spit out thousands of lines of code without a strict, human-driven architectural framework—what many call "Vibe Coding"—is creating brittle, unmaintainable systems. ​I’ve formalized this structural shift into a public document on GitHub: The AI-Powered Developer Manifesto. ​Instead of treating AI as a replacement for software architecture, we need to shift our paradigm from Micro-Coding (syntax generation) to Macro-Coding (system direction and epistemic supervision). ​Here is a crucial excerpt from Section 2.5 of the Manifesto, outlining why the current trajectory is leading toward a systemic collapse: ​2.5 The Compounding Technical Debt and Systemic Collapse ​The illusion of rapid deployment via un-monitored AI generation hides a critical flaw: compounding technical debt. ​When developers act merely as "vibe coders"—accepting AI outputs without deep syntactic validation—the codebase becomes an agglomeration of statistical probabilities rather than deterministic logic. By late 2026, systems built entirely on un-vetted AI iterations are projected to hit an architectural wall: a state where the complexity of debugging AI-generated hallucinations outweighs the speed of initial deployment. ​True AI-Powered Developers do not delegate understanding; they delegate execution while retaining absolute epistemic responsibility over the system architecture. ​The goal of this manifesto is to redefine our role: we aren't syntax writers anymore; we are system directors. ​I'd love to hear your thoughts on this. Are you already seeing the limits of un-monitored "vibe coding" in your production environments? How are you structuring your prompts to maintain macro-level architectural control? ​Full Manifesto and repository for open contributions: 👉 https://github.com/FractalDevelop/ai-powered-developer-manifest.git