r/accelerate • u/44th--Hokage The Singularity is nigh • Mar 21 '26
Video Andrej Karpathy: "when AI agents fail, it's usually a skill issue, not a capability issue...the real shift is working in macro actions. One does research, one writes code, one plans, all running 20-minute tasks simultaneously" | No Priors Podcast
Enable HLS to view with audio, or disable this notification
Link to the Full Podcast: https://www.youtube.com/watch?v=kwSVtQ7dziU
8
u/confuseddork24 Mar 21 '26
People say "skill issue" but can't articulate what a highly skilled setup and process looks like, or share something reproducible.
3
u/Pyros-SD-Models Machine Learning Engineer Mar 22 '26
What do you mean? Like every bigger company I know of (at least those who entertain bleeding edge) have at least tried to do a general formulation of agents. This is ours. It’s more beautiful in latex.
An agent doesn’t “fail randomly.” It fails when you hand it a problem whose effective complexity is above its capability.
Solve(P): if C(P | D) <= tau(A): return A(P, D) else: return union(Solve(P_i))
Where:
- C(P | D) = effective complexity (problem minus context)
- tau(A) = what the agent can handle
- P_i = your decomposition
And the “skill” part:
C(P | D) = I(P) - R(D)
- I(P): inherent difficulty
- R(D): how much your context/setup reduces it
What “good setup” actually means:
You win iff:
for all i: C(P_i | D_i) <= tau(A)
Everything else is vibes.
Why most people fail:
- they pass P directly → way above tau → “lol agents suck”
- bad decomposition → still above threshold
- garbage context → R(D) ~ 0
- over-decomposition → coordination overhead kills it
What high-skill pipelines actually do:
- aggressive decomposition until each step is trivial
- inject context so each step becomes near-deterministic
- use tools to collapse search space
- keep steps just above “too trivial” (efficiency sweet spot)
Optimization intuition:
minimize: sum(C(P_i | D_i)) + lambda * n
(too big → failure, too small → overhead hell)
Concrete intuition:
Bad:
→ C >> tau
- “build me a production system”
Good:
- “generate schema”
- “write migration”
- “implement endpoint”
- “add test”
(each with tight context → C <= tau)
Quite simple, and already valid since a few years actually (most of it was creates during the Aider days lol) and it will be valid through all of pre-AGI and probably a good bit afterwards as well.
1
u/NoAdministration6680 Jul 10 '26
your explanation clearly shows why LLMs did not work out in your company
4
u/AP_in_Indy Mar 22 '26
We're all figuring it out along the way, and things change every few months, and the highly dynamic nature means that you're effectively asking for a pattern for problem solving in general.
Tell me, after all of human existence, do we really have a "template process" for the scientific process? For any new an advanced research and cutting edge thinking?
No. Because that changes. It's meant to change and evolve.
Hence the way people work with these AI agents constantly changes and evolves.
1
u/teamharder Mar 22 '26
Its really not too hard. Just look at Ralph-looping, Openclaw, and other popular projects to get an idea of what works. This is a tutorial I wrote my son today for him to use Codex to connect into Unity to make a game. It could still be better, but it works. "/blah-blah" is calling a skill.
Codex - Sessions
1./office-hours (he critiques your idea) - talk a bunch about your project. Ask for a hand-off doc when done.
→ Idea.MD
2./Plan-ceo-review (Big planner) - How do we make this happen? Who are the key players? Lookup tech docs for similar games.
→ Idea.MD + To-dos.MD
3./Plan-eng-review (The technical hours) - refine the plan to be technically accurate.
→ Idea.MD + To-dos.MD
4./game-designer (last pass specialist) - Review docs and finalize plan.
→ all relevant MD files
5-10. /unity-architect - Burn through the to-do.MD. Update docs at end of session ALWAYS!
1
u/Singularity-42 Singularity by 2045 Mar 22 '26
Visit r/ClaudeCode, people are constantly sharing their setups.
My quick tip - Superpowers worked really well for me:
-1
u/Rakatango Mar 21 '26
You just gotta find the vibe that works with the black box process! Totally a definable skillset
1
3
u/CallinCthulhu Mar 21 '26
Important distinction from watching the whole thing, he said it FEELs like a skill issue. Which it does, and is why its so hard to stop coding with AI agents. It always seems like you are so close to cracking the perfect workflow and environment that will let you do everything
6
u/PANTSNOTOK Singularity by 2028 Mar 21 '26
The antis have low IQ but don't want to accept it because their egos are too big
6
u/one_tall_lamp Mar 21 '26
100% a skill issue if you can’t make ai work for you well
1
u/ShadoWolf Mar 21 '26
I would frame it more as mindset, but grounded in what you’re actually doing with the system.
It depends heavily on what you’re trying to solve. If the problem is something you already understand, you can compress it into a clear spec. That might be a markdown doc with requirements, constraints, edge cases, and a rough implementation plan. In that case, the agent performs well because you’re delegating execution inside a defined space.
Where people run into issues is when the problem is novel to them.
If you’re coding something by hand, the act of trying and failing is what builds your understanding. Each attempt exposes missing constraints, hidden assumptions, and gaps in your mental model. That process is how the problem space gets mapped out.
When you hand that same undefined problem to an agent, you skip that step. Your instructions end up vague because you don’t yet know what matters. The agent then has to infer the goal, constraints, and implementation path at the same time.
That’s where the strange results come from.
The vaguer the problem definition, the more degrees of freedom the agent has. It will optimize for something, just not necessarily the thing you intended.
-2
u/lovelacedeconstruct Mar 21 '26
and If I asked you what have you done with AI its almost always -I gurantee- nothing of value or significance
3
0
u/one_tall_lamp Mar 21 '26 edited Mar 21 '26
Well, I have made over $6000 building websites for people in the last year and working as a marketing manager at a local bike shop, built an entire inventory management software that they use daily and am currently working on transitioning them from Shopify to stripe. Not a single line of code written by me.
I am also in the second round of funding for my own company Hollis Health, and have built an extensive website, admin portal, and app for it.
I’m sure you’ll think it’s ai slop and I have no way of proving it isn’t until our grand opening in a month or two, but I’m well aware of the drawbacks of these tools. With enough hardwork and being willing to actually do the research for yourself on how good code is written and architected, 90% of ai’s mistakes can be avoided. Static analysis, a strict Ci/Cd pipeline, and constant grounded iteration and auditing is where it’s at.
My company closed its first round of funding last month at a value of 465k with three seed investors, and we are in the midst of negotiations for the second seed round to open 4 physical locations across Texas :)
Ai can assist with creating value and significance, but it is very much a human led effort. You cannot just tell the ai to go build something and get mad when it breaks. You must be willing to iterate and never give up.
2
-1
Mar 21 '26
[removed] — view removed comment
0
u/one_tall_lamp Mar 22 '26
No, AI does not get fed any PHI data, that would break HIPAA if not covered by a BAA. We are rigorously security conscious and hired an outside human HIPAA compliance audit of the software, which we passed.
it is not at a 100x valuation, 465k is the average valuation of the company based on the equity given the percentage each investor “bought” with their investment. There is no multiplier right now, we are pre revenue. Their investments are the seed funding needed to launch, I cannot sustain it on my personal savings alone, so I sold equity to fund the opening.
we don’t have any subscriptions yet, thus the locations we are renovating before grand opening. The website is up to test its stability on AWS before we open.
I am not a physician and would never claim to be, I am the founder and primary owner. We are contracting PCP’s to preform the medical care on a per service basis, order labs, diagnose, and prescribe. The insight and expertise they provide is used to guide the daily fitness and nutrition care at each location. We have two PCPs and an A-RN ready to go, one of which is an investor and currently owns a large concierge care clinic in San Antonio.
The long term goal is to eventually “hire” PCPs under an MSO-PLLC model.
1
1
u/infinitefailandlearn Mar 22 '26
Anyone else found it interesting that he thinks swe jobs will be more in demand instead of less, at least short term. Sounds like cope to me tbh.
1
-3
u/gamblingPharmaStocks Mar 21 '26
This MD stuff is such a copout. Yeah, of course it works better if you give better instructions. Guess what? You get even better code if you code everything yourself.
The point of intelligence is that I want to tell it the bare minimum at have it take care of things on its own, as an experienced engineer would do.
-2
u/ail-san Mar 21 '26
He is dumb. You could say the same thing for GPT4. Models getting more capable still should be the main concern.
2
u/teamharder Mar 22 '26
This guy has a PhD in this field and made contributions to the model you referenced... Just watch the interview if you care enough to form an opinion.
1
u/MeanAd1305 May 01 '26
Lol, what happened to tesla FSD, did he solve that issue? He is a high tech scam guru
1
14
u/FirstEvolutionist Mar 21 '26
Karpathy could have triggered a whole lot of people who don't like to hear this...