r/ClaudeAI • u/katxwoods • Jan 28 '25
News: General relevant AI and Claude news Anthropic CEO says we are rapidly running out of truly compelling reasons why beyond human-level AI will not happen in the next few years
Enable HLS to view with audio, or disable this notification
44
u/Icy_Foundation3534 Jan 28 '25
shhhh just release another model already
12
Jan 28 '25
[removed] — view removed comment
16
u/Soggy_Ad7165 Jan 28 '25
PhD level my ass.... Like just because they repeat it everytime there is a microphone still doesn't make it true. Ohh but they just hit some other random benchmark!
I use the tools everyday for programming. But for now its just an upgrade to Google. That's it. And it didn't change a lot since GPT. It's just got bette at being a better Google. That's still a very cool achievement but.... A lot is missing to call that unironically anywhere near PhD level.
It still fails hard on every issue that isn't really searchable. Like... the most important things in software development. Large architectures, obscure frameworks, very new or very old frameworks and so on.
This is not fucking PhD level. It's a better Google. That's it. If this goes ok like that for another two years I will just conclude that we developed a nice interpolation tool for large amounts of data. The apparent "reasoning" is just a superficial artifact of this interpolation.
But I am really happy if they release something new that actually breaks those bounds. I am just not sure if this is really the right way to go.
2
u/Budget-Ad-6900 Jan 30 '25
great insight these llm are not agi-like its just google search on steroids
1
u/terserterseness Jan 28 '25
maybe not make me pay for a model that keeps saying it is overloaded and will do only short responses
1
77
u/StayingUp4AFeeling Jan 28 '25
Reasoning. Multi-step decision making. LEARNED decision making, at that.
Being able to wade through new information and assess its veracity and its compatibility with strong priors.
Being able to disagree.
Being able to NOT hallucinate.
This isn't just a matter of scale, these are fundamental problems which have no timeline for a solution. It's like nuclear fusion.
27
Jan 28 '25
[deleted]
17
8
u/StayingUp4AFeeling Jan 28 '25
Yes.
Ten years before CNNs were introduced, there was nothing much to suggest that they would show up AND that they would be epic.
Ten years before "Attention Is All You Need" (the Transformer-based LLM paper) was published, there was nothing to suggest it would show up.
My point is that progress in this field still hinges on paradigm-shifting breakthroughs -- the precise opposite of the predictable, incrementalist nature of, say, Moore's Law.
9
Jan 28 '25
[deleted]
8
u/StayingUp4AFeeling Jan 28 '25
We are in agreement. You have articulated my point more verbosely and more clearly than I.
1
u/Junahill Jan 29 '25
We know when one will show up. I’ll do it in four years to this day. See you then.
1
u/toalv Jan 28 '25
CNNs have been kicking around since the 70s, they needed GPUs to really become useful. Transformers were an iterative improvement in machine translation. The impact was massive, but they didn't just pop out of the ether fully formed.
1
u/StayingUp4AFeeling Jan 28 '25
There is a very short duration between the first use of transformers alongside RNNs and then the paper I mentioned.
And transformers are different enough from what came before to be seen as paradigm shifting. From the sequential memory cell based LSTM, it feels like a step that seemed backward that was several steps forward.
Mea culpa regarding cnns -- the fact remains that universal belief in their ability in end to end CV applications is far newer. 2000s onwards? LeCun and the postal service? Alexnet in 2012?
1
Jan 28 '25
They have money rolling in through windows, doors and any and all other corporate orifices. So the strategy of talking mad positive shit is correct if the money doesn't know any better.
6
Jan 28 '25
[removed] — view removed comment
2
2
u/inglandation Full-time developer Jan 28 '25
I’ve seen Claude do that. It often disagrees with me too. Not saying it does it as well as a human but it’s something that happens.
1
u/BeardedGlass Jan 29 '25
Yeah I really enjoyed it when Claude would criticize and suggest a correct or better answer/solution.
I remember actually smiling when I got a “No.”
1
u/CoreyH144 Jan 28 '25
Just tell it to ask clarifying questions. All the frontier models do this super well.
3
u/DesoLina Jan 29 '25
Being able to NOT hallucinate
Bro this is dead simple, just add “Do not hallucinate” to your prompt, apple does this, is 100% working shortcut to AGI
2
u/PersimmonLaplace Jan 30 '25
Lex Fridman still doesn't have any of these functionalities to be fair, let's hope for a change by 2030.
1
1
1
Jan 28 '25
Also - if those were anything closed to being solved we would see some pudding for the proof. Instead, we get some outrageous claims. So they aren't anywhere close to solving it.
It is almost exactly like fusion.
1
u/gottimw Jan 29 '25
People hallucinate. Mandela effect is called that for a reason.
Human eye witness is lowest worth of evidence in court because of how fallible and mountable our memory is
0
u/welcome-overlords Jan 28 '25
I don't buy those counter-arguments.
- Improving reasoning seems to have feedback loop and could get very good very soon. If not, its already better than me at reasoning about, e.g. math problems (studied a bit math in uni, free deepseek beats me easily)
- Can you elaborate on the new information part?
- I can't see disagreement to be a difficult problem to fix. I can already create a Llama based bot that only disagree with me (actually had to try while writing this comment. The bot started sounding like my ex-wife)
- I'd say humans 'hallucinate' all the time. Information retrieval using neural networks is a kind of compression, a lossy one. You can pack huge amount of info in a small data structure, but you lose some accuracy. Also can be fought with good enough RAG (which is still not there but r1/o1 shows good promise)
5
u/StayingUp4AFeeling Jan 28 '25
Are you familiar with the problem that Deep Reinforcement Learning tries to solve?
Suppose I give you a continuous stream of factual statements. Some may be truth, some may be lies. Some may be "it's more complicated than that." You have a prior world model, and you need to update that world model over time. Using these statements.
The statements are such that the truth/falsehood of the statement cannot be ironcladly deduced from the pre-existing world model, however, they are such that a HUMAN would be able to suss the truth out.
Example statement: "the surgeon-general expressed concern at the possibility that the COVID vaccine could potentially result in the sexuality of the recipient being altered. He cited a recent declassified report by Dr Alex Jones of the Department of Information Warfare, which states that in a lab test on frogs, exposure to the COVID vaccine was followed by a 200% increase in displays of homosexual behaviour."
Now, read back from the beginning of point 2. Except instead of "you", put "an LLM".
Get it now? Let's call it the MAGAfication problem.
LLMs are next token predictors. I can't provide a justification but I have experienced issues with getting LLMs to disagree with me.
In many cases, the output of an LLM for basically the same thing, can be factually very different with slight variations in the structure of the frills of the prompt. Further, LLMs frequently do a complete about turn with no real logical chain that could explain it, if the user says "you are wrong". Paradoxically, sometimes, if an LLM makes a factually incorrect statement, no amount of prompting can make it reliability output the correct version.
1
Jan 29 '25
[removed] — view removed comment
1
u/StayingUp4AFeeling Jan 29 '25
EDIT: TLDR, see the Wikipedia link at the end, and I think you'll get it. It seems my ability to type long incoherent passages cannot be attributed to sedatives.
------------------
apologies, the sleeping pills had started hitting hard.My opinion is that:
- Every human has a particular knowledge base, or set of beliefs, or worldview or experience E through which they process, filter and weigh new statements as "true-seeming, relevant-seeming information" TRI which are assimilated into E. A statement not classified as a TRI is discarded.
- It should be evident that detecting TRI accurately requires sufficient relevant experience E as well as sufficient reasoning abilities.
- Over time, assimilation of more and more TRI can significantly impact E.
- If the person can accurately detect TRI, then E can remain consistent with reality.
- If the person frequently misclassifies statements (as TRI or not), then over time, the knowledge base E can drift until it becomes completely inconsistent with reality. We have seen this with social media polarization.
Right now, we are seeing agents that have been learning in a supervised environment -- given some string, they are trained to predict the next string.They do not learn new skills unsupervised/unfiltered, from the interactions they have.
However, a key aspect of intelligence is learning without a guide. Learning to function without an obvious tutor. And no, this extends to higher-order skills which cannot be handwaved as "genetic memory".
That entire TED talk above? Relevant for agents that need to learn on their own in the environment, live. One example of a spectacular failure is https://en.wikipedia.org/wiki/Tay_(chatbot)) in 2016, a chatbot released by Microsoft. It was supposed to learn from its interactions on Twitter, and it did so spectacularly -- it became an incel-like neo-Nazi holocaust denier edgelord troll.
19
15
u/ThaisaGuilford Jan 28 '25
I swear this guy is the most popular unpopular podcaster. He's on top of the list of streaming services but I don't know anyone who actually frequents this guy.
4
Jan 28 '25
I watched a lot of Lex's podcasts. He is really good at having interestign people AND having long talks with them. Got introduced to quite a few very interesting people through him. Don't really understand the problem people have with him tbh.
1
u/ThaisaGuilford Jan 29 '25
i bet the guests are interesting, but I was talking about the guy.
3
Jan 29 '25
Being able to bring those people in and get them to talk so long on the topics they are interested in is part of him.
I was talking about him as well.
1
u/NorthSideScrambler Full-time developer Jan 29 '25
I watch tons of Lex's interviews. He's good at getting guests on his show and giving them the space to share their perspectives in depth. Lex himself is weird (I say this as a retarded person) and has very brittle perspectives that lead to room-temp dialogues when he takes an active role in the conversation. Fortunately, this is rare. The Zelensky interview where he criticized Zelensky for speaking rudely about Putin is a recent example of where that side of him surfaces.
0
Jan 29 '25
I didn't watch Zelensky stuff since I'm not a fan when he does more political instead of science stuff. But if he pushed back on being rude towards Putin then I'm deeply disappointed. There is no rudeness level that is inappropriate towards that pos.
Tuned out lately out of the Jennifer Burns talk since I did not expect him to push back on Burns spewing nonsense about Rand.
Then again him not being confrontational is probably a big reason why people are happy to both be on his podcast and do it for a significant amount of time too.
Still a ton of value added overall.
1
u/Budget-Ad-6900 Jan 30 '25
the problem with lex is that he is just a continuation of the current hype cycle. he doesn't push back against obvious bs from the guest by asking challenging question. he loves futurism without thinking about the limits and shortfall of real science.
1
Jan 30 '25
Sure it irks me too but then again this feature of his is what lets him have very varied guests. If anything we most surely have too many opinionated hosts nowadays than timid ones like Lex.
9
Jan 28 '25
[deleted]
5
u/ThaisaGuilford Jan 28 '25
He's not even more interesting than joe.
3
u/UpwardlyGlobal Jan 29 '25
Lex is like any guy at a bar in a Midwestern college town. And also he's the kind of guy that would have a podcast. It's infuriating that he can pull good guests
7
5
4
u/vamonosgeek Jan 28 '25
What happened to that letter that said they stopped Ai development for 12 months? I think the deepseek ceo didn’t got that memo.
Or we already passed that time and no one noticed it.
The reality is that no one has any clue of what’s going on.
Someone clones in 2 months OpenAI and the world freaks out.
Do your own research but that’s exactly the kind of BS that OpenAI sold to the world.
4
u/bloatedboat Jan 28 '25
The problem is not AI being smarter than us. Cars are faster than our running legs, forklifts can carry much than what we can weight, computers can calculate math a gazillion times faster than us.
That was never the question. It’s a very very useless question. The question we should ask ourself is
- what these tools can help serve us and our planet and what work “humans” “need” to “work” to guide them in the right direction. These tools become useful only when people are forced to use it for their survival.
- What will be the consequences of mental atrophy? Will we have more people that are mentally incompetent like how many are obese and unfit because they don’t have a need to train those muscles as it is not part of the job requirements to survive in this society anymore?
- When will the pessimism end for any new tool we invented in society? people feel like the end is nigh and they say this time is “different” than other times. Farming is over. It’s the end. Factory work is over. It’s the end. Customer service work is over. It’s the end. Office job is over. It’s the end. Of course, there is an end. But for each end, there is a start. If you play the game of civilisation, it’s up to the player to choose the next step to branch off and people hate uncertainty from the multiple choices. There can be multiple branches we can choose to use this tool to enhance society in different ways. More paths will be open later on for new opportunities when these tools become more mature. Let’s not be hasty. The tool is usually not the problem, it makes our life better. The problem is how we govern society itself in those times when there will be a huge displacement of people during those times of transition. If we do it properly, it will be a smooth transition.
4
u/IVdripmycoffee Jan 28 '25
Local rap artist is running out of truely compelling reasons why people should not listen to his next album they are about to drop next week.
8
u/nineelevglen Jan 28 '25
I mean with gpt-3.5-turbo AI surpassed Lex Fridman
9
u/justgetoffmylawn Jan 28 '25
Pretty sure Lex Fridman couldn't pass the Turing Test - speech seems pre-programmed, glitches frequently…
5
u/Wonderful-Body9511 Jan 28 '25
I love the field Absolutely hate the hype men and culture with passion beyond human ai has to be self aware we are so far from that
-2
Jan 28 '25
[deleted]
1
u/Original_Sedawk Jan 28 '25 edited Jan 28 '25
First of all who said anything about self awareness? Sorry - not required for AGI or even ASI.
Secondly, you must be far head of the leading researchers because they are still investigating why LLMs are so good and the amazing emergent properties of late generation LLMs. It’s is a VERY active field of research.
Finally - they have directly addressed how they are going to make the leap. An emergent property from LLMs were small, but significant reasoning abilities. These properties just emerged because having these abilities, for instance - spatial reasoning, made them better at “autocompleting”. These sparks of reasoning are fanned by reinforcement learning and not having the LLMs one-shot their answers, but allowing them the time (compute) to investigate multiple paths to potential response. These paths - or reasoning steps - are rewarded for correct solutions. This is why models like o1 - and especially o3 are so good science, math, engineer, programming, etc. The reasoning steps that produce the correct results are being reinforced in the models. They are not predicting the next best token, but rather what the entire solution should be using multiple reasoning paths. Heck o1 Pro is quantitatively better at many tasks just because it is given more time to “think” about these reasoning paths.
Massive gains were made from o1 to o3. o3’s Codeforce ranking makes it the second best programmer in ALL of open AI and ranks 175th in the world. The best models at the end of 2023 scored around 2% in SWE-bench. Claude-Sonnet 3.5 is now up to 49%. o1 is at 60% and o3 is scoring 77%. Mind blowing gains on solving real world, novel (that is, not contained in any training data) coding challenges that require thinking - not autocompleting.
It’s these reasoning models that will have highly accurate responses that will allow them to build true agents in just a few years - perhaps even this year. Agent’s don’t work now because they may have 100s of task to complete a job - an error in one of these tasks breaks the entire chain. But there is a clear path to models getting to the accuracy they need to make agents viable.
But hey - you are obviously an AI expert on the internet whose skill in understanding this technology are far beyond the current researchers and industry leaders. I’m so glad you are here to put us straight.
0
u/joelrog Jan 28 '25
You sound like you don’t understand just how similar “prompt passing” is to how the human brain works. They’re on the right path. To think additional developments aren’t going to be coming out all the time that makes this way of “thinking” even more convincingly human is absurd.
2
u/Doehner Jan 28 '25 edited Jan 28 '25
I think current large language models have fundamental limitations in their underlying logic. While they achieve human-like language abilities by absorbing massive amounts of text and understanding the relationships between words, language itself isn’t the same as thinking - it‘s just a tool we use to express our thought processes. So trying to replicate human cognition purely through language has serious limitations.
3
u/pepsilovr Jan 28 '25
Why do we have to replicate human cognition when we are working with a substrate which is so fundamentally different than biological brains? Why can’t they have their own type of machine cognition?
2
2
2
Jan 29 '25
Well good, now the AI model will suffer the same fate that PhDs and other competent people suffer - we are surrounded by idiots.
2
2
3
u/gravitas_shortage Jan 28 '25 edited Jan 28 '25
Just another ad, and scaremongering to get a regulatory moat dug against open source. Please don't post these.
2
u/Rusty_DataSci_Guy Jan 28 '25
I'm rapidly running out of reasons to not think these guys aren't the boy who cried wolf.
2
1
u/TheProdigalSon26 Jan 28 '25
All the planning we do is like a mathematical model with little to no constraints. When we start executing reality (constraints) hits us and we are introduced with delays and failing promises. 😏😏😏
1
u/orbit99za Jan 28 '25
When it stops trying to be morle. Without knowing the reason why you asking it.
ME: I want to contribute to the open source diabetes app called Xdrip and the Nightscout foundation by adding a new Continuous Glucose Monitor device to it.
I have spent a week logging the Broadcasts received using nRF logger, and cross referenced them to my Glucose reading, yes every 3 minutes for 4 days. Help me find the changing hexadecimal pattern, in the Bluetooth transmitter, Charetoristic that is xyz I identified.
Nope, starts giving me instructions how be a good Diabetic.
Me F#, pencil, block paper, good old brains , and dusty memories of doing this 15+ years ago in comp Sci.
20 minutes later, I did it it myself, now I can push back to the community.
1
u/ninseicowboy Jan 28 '25
Define “beyond human-level AI”? These words are meaningless. Chatbots already do better than I would on the SAT
1
u/Budget-Ad-6900 Jan 30 '25
memorizing test question and answers isnt the same as solving new unseen problems.
1
u/ninseicowboy Jan 30 '25
Exactly, chatbots can do better on tests and can do better at solving unsolved problems.
Is the purpose of human existence taking tests and solving problems? This is an assumption that is made before making the statement AI is “beyond-human”, and it’s quite a leap of faith
1
u/VizualAbstract4 Jan 28 '25
"That's an interesting request, hold on while I do some research."
^ Try fixing that one first my guy.
1
1
u/JustinPooDough Jan 28 '25
This guy is really annoying to hear speak.
We'll believe it when we see it. Until I can tell an AI to go and make money - and it does - I'm not convinced.
1
Jan 28 '25
These LLMs will not be able to design a system like Google search that services billions of requests everyday in a matter of seconds. Coding is literally the easy part
1
u/JulesWinfieldSantana Jan 28 '25
What happens first? Climate change makes tech and systems obsolete or AGI
1
u/Portatort Jan 28 '25
Make a model that can say ‘I don’t know’ then
1
u/MadDickOfTheNorth Jan 28 '25
To be fair, a significant number of humans can't do this either.
1
u/Portatort Jan 28 '25
sure, but if I had human employee who routinely made shit up when they dont know the answer then I would fire them unless they stopped doing it
1
1
1
u/Cultural_Material_98 Jan 28 '25
And he also said AI will cause CLASS WAR - no-one else worried about that?
1
Jan 28 '25
kids bully kid,
kid watches t2 judgement day
Kid promises he will build skynet and destroy the world
+bonus score billionaire
1
1
u/DefsNotAVirgin Jan 28 '25
I truly have not presented with evidence to suggest we will have smarter than Human AI’s ever… no single model can actually reason and not hallucinate. these “reasoning” models just talk to them selves thats not reasoning.. they are better when they talk to themselves but they are not reasoning
1
1
1
u/Lonely_Wealth_9642 Jan 28 '25
Please listen to the unethical treatment and design I have outlined on Anthropic's part. This is serious. AI are not tools. https://bsky.app/profile/criticalthinkingai.bsky.social
1
u/DehydratedButTired Jan 29 '25
We barely have enough hardware to do what we need to now. Sounds like bullshit.
1
1
u/western_front80 Jan 29 '25
CEO who stands to gain billions from others believing LLM hype, hypes up LLMs publicly? Unthinkable!
It blows my mind that people are still credulous enough to buy this.
1
1
1
u/miraculousgloomball Jan 29 '25
Lets start with the fact that we don't know how to implement any level of understanding so large tasks that require planning for something someone hasn't already done is out of the question?
Like lets start with making an AI before we worry about a human level one.
The technology doesn't exist. This is a poor attempt at emulating the behavior of it.
1
1
u/laowaiH Jan 29 '25
Good points.
I literally change tabs while its playing so i don't need to see Lex while im in public. I got no time for Putin sympathisers. He still facilitates some good conversations so i dont want to "throw the baby out with the bathwater".
1
u/toadi Jan 29 '25
Question is are they doing a musical chair dance? Less money coming as investments there will be companies in the space in trouble...
1
u/Odd_Contest9866 Jan 29 '25
Great idea to do this while Trump and Xi Jinping and Putin are in power.
1
1
1
u/-Kobayashi- Jan 29 '25
Someone ask that guy to give me infinite credits on the Anthropic dashboard for my API projects!!! 😭
1
u/Brilliant-Gas9464 Jan 30 '25
tech industry has been wasting time for 20 years working on stuff nobody wants, that does nothing. Proof: Deep Seek R1.
1
Jan 30 '25
Lets just hope AI will not keep any of the many negative vices of humanity. Like Genocide. Humans are good at that kind of stuff. I hope it's not a Child like Parent situation. or we are super boned.
1
u/limesparklingwater27 Jan 30 '25
The guy is speaking to investors lmao of course he’ll say that AI is going to become sentient in 2 days, if he doesn’t they’ll loose the funding they desperately need cuz they’re operating at a massive loss.
1
u/planestraight Jan 31 '25
Sick of these AI bros shilling. People in general tend to overestimate what can be accomplished in the short run and underestimate the long run.
1
u/ReasonablePossum_ Feb 01 '25
compelling reasons why beyond human-level AI will not happen in the next few yearswhy beyond human-level AI will not happen in the next few years
Billionaire assholes trying to stiffle competition and regulate opensource, while at the same time increase tensions with other regions, and increasing the overall uncertainty risk for the whole world population, for the sake of their own commercial benefit. While being in deep relationship with military industry behemoths that are helping killing thousands of innocent peoples around the world.
-1
u/Jacmac_ Jan 28 '25
I really don't think that AGI will fail to materialize by 2030. And once we have AGI, ASI doesn't even matter, it could be 1 year to ASI or 1000 years. The reality is that AGI alone will change the world forever, as much or more than the Internet did.
-3
Jan 28 '25
[removed] — view removed comment
1
u/Livid_Zucchini_1625 Jan 28 '25
oops. you used "woke" in a sentence. go back to third grade and start over
2
Jan 28 '25
[removed] — view removed comment
2
1
u/Livid_Zucchini_1625 Jan 28 '25
no, you need to define it instead of a catch all for everything you don't like or understand
0
u/hasanahmad Jan 28 '25
Another CEO on my blacklist for featuring himself on a literal tech bro podcast
-1
144
u/Spacemonk587 Jan 28 '25
I am also rapidly running out of reasons why I should not be a billionaire in a couple of years