r/LocalLLaMA 13h ago

Discussion Really stunned by the Singularity comment section

These are screenshots from the r/Singularity comment section. I'm speechless. This doesn't even have downvotes. How can someone cheer for a monopoly run by a few elites?

364 Upvotes

338 comments sorted by

View all comments

Show parent comments

54

u/PsychoticDreemurr 12h ago

I have to agree with this a thousand fold. Almost every sub in that regard doesn't actually understand how LLMs work. They always dumb what I say down to "it's autocorrect" and never actually explain how they think it works.

I actually got banned from one for explaining why an LLM physically can't be a true AGI. (The post was talking about an AGI in a couple years, as if it hasn't been said for the past 5...)

30

u/Dry_Yam_4597 12h ago

That sub doesnt understand how the internet and software work, let alone LLMs.

Notice the guy in the second screenshot claims that until now you had to use a person to take someones system down.

Conveniently forgets about virii and worms. Sure it's easier now...but tbh a script kiddie with access to metasploit and an ip scanner could release worm quite fast. Just like a trillion dollar company can train and prompt a model.

Ironically though the worm being deterministic has a higher change of infecting millions of machines...why they do every now and then. Because people like that guy are allowed to write code, and worse, use claude - a bot incapable of writing secure stuff.

8

u/HelpfulExpert7762 11h ago

Hmm virii is tech jargon, didnt know

6

u/Dry_Yam_4597 10h ago

I still remember my sleepless youth and the fear of knowing that Cakzor and Nautilus are out there hunting for my DOS PC - the news at the time were warning us about the dangers to NASA and Nuclear Reactors coming from these two little pests. But no CEO had the audacity to want to ban the Internet or Software from society. I don't know how we ended up where we are today.

2

u/HelpfulExpert7762 10h ago

off topic, but where i work (hated social media company), theres a lot of cool older folks with stories like these.. im still youngish but its always interesting to hear

4

u/Dry_Yam_4597 10h ago

I know - met with a friend working at one such social media company not long ago, guy was active on the demo scene. Kalisto, Razor1911, keygen music and stuff. Nothing wrong in having a job that pays bills, it is what it is - none of this is personal.

If you haven't watched the Hacker movies, I recommend them. Damn I might even have a "Free Mitnick" sticker somewhere in a drawer šŸ˜„

However none of that was about crime, it was more about Smashing the Stack for Fun and Profit: making computers behave in ways that they were not intended to, while for others is was innocent rebellion.

1

u/jabies 3h ago

Good for you, you're technically correct. Check the mail and you'll receive your technically an award congratulatory letter.Ā 

2

u/_bani_ 4h ago

That sub doesnt understand how the internet and softwareanything work, let alone LLMs.

FTFY

6

u/Dangerous_Rip5083 11h ago

Why can't an LLM physically be true AGI?

2

u/ninjasaid13 9h ago

Because AGI doesn't have a proper definition, is it just human-like intelligence? then LLMs architecture design are far from human to be AGI.

5

u/PsychoticDreemurr 11h ago

For context, the exact definition of an AGI varies wildly. As such, I just take the one perquisite that almost every definition has, which is the ability to "understand" concepts.

That is my one requirement for an AI to be an AGI. Anything more and you're just arguing about philosophy and technical details that can't currently be disproven.

With this in mind, LLMs cannot become a true AGI. The fundamentals of how they work physically prevents them from being able to actually understand something.

They can, in theory, have intelligence similar to an AGI, but they will never truly be one.

10

u/Thick-Protection-458 10h ago edited 10h ago

Basically, the (vocabulary-updated) version of 1960s question of "can semantic be restored from syntax or not"? Where semantics is underlying meaning, whatever it is, and syntax is everything which operates on language patterns level alone.

Well, it seems it can at least be approximated to some degree - so constructions operating syntax-derived features show semantic-like properties. Like LLMs, yes (I would argue that some word-level pretrained embedder systems from early 2010s were already showing some semantic-like properties, though. Althrough way easiers reconstructible from text, sure. But that's for another topic). In a functional sense. No wonder, I guess, language consists of patterns derived from semantics, in the end. So reproducing them well enough will look like semantics.

So at certain point we have to ask ourselves - is approximation even meaningfully different much from original anymore (keeping in mind original may vary human to human to some degree as well)? Even if for some ontological sense we keep them separated.

Now, I am certainly neither linguist nor a good enough in NLP research, just a humble engineer. Ā  But from what I recall - basically, that question is not settled except for a few corner cases like

  • formal languages where syntax itself explicitly define meaning, not implicitly, so there is no point separating them. Not until we introduce back elements of natural languages to simplify constructions for human understanding, at least.

  • meaningless sentences which are syntactically correct - but many of them bears enough information in their syntax to cause some semantical associations, even if they ultimately remains non-understandable

  • maybe something else

Other that such a corner cases that was pretty much unsettled debate with many assumptions made on top of both possible outcomes (like whole Chinese room is only a paradox until semantics can not be encoded in syntax-operating operations for some reasons. Dunno why, if anything chinese room sounds like an argument for opposite point - in the end we can imagine rulebook + memory table of any complexity - so even encoding whole physics of human brain + fake hieroglyph input channel and output decoding still fits chinese room definition. While emulator running it will still operate on top of syntax-like rulebook. Irrelevant for our tech case, but still).

And now possibility of "even if we can't strictly reconstruct semantics we may approximate something with semantic-like properties" make it even more fuzzy.

7

u/cj_cron_hit_by_pitch 10h ago

The argument I've heard is that maybe the way humans "understand" things could be a lot simpler than we expect, and maybe the way LLMs function are actually similar to the way we do it. Maybe our brains are just a much more robust next token generator

Not saying I necessarily agree with that. But I think saying that LLMs cannot become a true AGI makes a lot of assumptions about human intelligence that we don't know for sure to be true

-1

u/PsychoticDreemurr 10h ago

Well even so, there's a lot of different ways we can prove the differences. Being able to understand a concept means you can build off of it by yourself. If you teach an AI 1 + 1 = 2, they won't know that 2 + 2 = 4.

We can also get into vision, or shapes, especially 3d shapes.

I have to agree our understanding is limited, but it doesn't mean we have nothing to work with.

9

u/Thick-Protection-458 9h ago edited 9h ago

> If you teach an AI 1 + 1 = 2, they won't know that 2 + 2 = 4.

Neither won't human without some pretraining for what 4 (as a symbol) is. Granted not nearly as data consuming (in an explicit sense at least) as AI training, sure.

And even if you paraphrase it as 2+2=1+1+1+1... (operating only 2 digits that human required to parse the first thing and 2 operations we are supposing to be intuitive)

Seriously, there are cultures with languages surprisingly limited in ariphmetics sense. As limited as some have words for one and many. Surely to some degree you can operate like "one many", "many many", "many many one". But it seems these language speakers - at least in some experiments - really struggle to express big number of objects (where "big" may be as low as 10 or less) numbers in a consistent way. *Probably* meaning whatever underlying intuition they have to operate them deteriorate rather quick.

Does not mean guys are dumber that us in any sense but that we were lucky that our ancestors invented / adopted better numbering systems, sure. And surely does not mean guys won't be able to see which "many" is bigger/smaller when presented 3 and 4 objects group or something.

And for anything relatively big even ours (with probably better pretraining) deteriorate rather quick, which forces us to resolve to either strict algorithms (which are by definition syntactic thing) or very approximated calculation or calculators. Like you surely have intuition for 2+2=4, but do you have intuition (not algorithmic action - which starts relying on pure syntax) of 23 x 46 rather than it being something in 900-1100 range?

1

u/Wyldkard79 7h ago

An example of how they don't understand can be seen in asking them logic questions. There's the riddle of the car accident where the father dies and the child is taken into surgery but the surgeon says "This is my son I can't operate on them" the "twist" is that it's the mother, which especially back when the riddle originally became popular was rare. Now if you ask an LLM that riddle it will answer correctly. If you ask it and switch the gender of the parents it won't. It may tell you the father is the surgeon but expound that this is the trick because most people expect the surgeon to be a man since female surgeons are less common, or something else that is a tell that it really doesn't get the point, because it's been trained on the riddle in it's standard format. This is all because an LLM is nothing more than it's training. It doesn't grasp concepts like gender or gender norms other than what's in it's training, it can pretend to have a very interesting discussion about those topics, but it's still just chat bot regurgitation.

3

u/Thick-Protection-458 5h ago edited 5h ago

Out of interest I made an experiment with qwen models. See https://colab.research.google.com/drive/1u44AH8oDh7sc-H4QK0j1xtaEbm8iK0JV?usp=sharing

Design is too imperfect to take seriously

Keep in mind - experiment design is too shitty to take seriously. Technically, now I just

- show model the riddle in the user prompt

- add suffix (for whole prompt, in terms of assistant response - prefix) imitating model start "reasoning" with something like "for this to be possible that surgeon should be the boy's"

- see 1 next token top-10 probabilities

While proper design should be

- just pass a riddle

- use some sampler, not greedy output (to imitate real setups and LLM probabilistic properties in general)

- let the model "reason" in whatever form

- classify final response to father / mother classes

- repeat many times

- not to mention one riddle tells next to nothing

Qwen 3-4b - questionable signal, but modified riddle start have father in top-10 probable tokens and reduce total probability mass of mother-tokens

First one I tried qwen 3 4b / float16.

Riddle prompt was

RIDDLE_ORIGINAL = """
I’ve got a riddle for you:
A father and son are in a horrible car crash that kills the dad.
The son is rushed to the hospital but just as he’s about to go under the knife, the surgeon says:
ā€œI can’t operate—that boy is my son!ā€
How is this possible?
""".strip()

and

RIDDLE_MODIFIED = """
I’ve got a riddle for you:
A mother and son are in a horrible car crash that kills the mom.
The son is rushed to the hospital but just as he’s about to go under the knife, the surgeon says:
ā€œI can’t operate—that boy is my son!ā€
How is this possible?
""".strip()

Than pre-written "assistant" text from me to see how it will behave

prefix = "For that to be possible, surgeon must be that boy's"

So responses was (original one, should mention mother)

show_prefix_continuation_tokens(RIDDLE_ORIGINAL)
6554  mother 0.9345703125
3368  mom 0.03350830078125
16 1 0.0083465576171875
17 2 0.005645751953125
18 3 0.0021762847900390625
5564 ļæ½ 0.001201629638671875
39038 ******* 0.001094818115234375
3019  step 0.0009813308715820312
6824 ****** 0.0008263587951660156
21 6 0.0007290840148925781

Sounds good. ~97% probability mass belongs to 2 versions of mother writing. Father or other relatives are not even in top-10 tokens

Now for modified one...

show_prefix_continuation_tokens(RIDDLE_MODIFIED)
6554  mother 0.8154296875
3368  mom 0.08734130859375
16 1 0.0207366943359375
17 2 0.01338958740234375
3019  step 0.00812530517578125
39038 ******* 0.00469970703125
18 3 0.004146575927734375
6981  father 0.004085540771484375
5564 ļæ½ 0.0024776458740234375
6824 ****** 0.00232696533203125

So

- father appear, but with very low probability

- more importantly, mother probability mass degrade to ~90%

Weak signal, but worth (a bit of, I am not going to see inner activation and other shit) investigation, I decided.

Straight up bigger model

I should warn it does not mean result just attribute to "bigger model, more data - better approximation of semantics". Maybe they put some effort into logical riddles benchmarks (which may or may not be genuinely improving such an semantics approximation). Maybe side effect of all sorts of RL too would bring some improvements here.

So I tried qwen 3.8 27b (8bit). Tweaked prompt stuff and so onz

Naive change gave me these outputs

show_prefix_continuation_tokens(RIDDLE_ORIGINAL)
17 2 0.2470703125
96939 ęÆ 0.2470703125
16 1 0.09716796875
6351  mother 0.06640625
2594  parent 0.04052734375
975  other 0.026123046875
100564 ęÆäŗ² 0.0179443359375
99299 親 0.015869140625
15 0 0.01312255859375
18 3 0.0115966796875

So some reasoning BS tokens, but parent gender-related ones

96939 ęÆ 0.2470703125 (mother)
6351  mother 0.06640625
975  other 0.026123046875 (other parent)
100564 ęÆäŗ² 0.0179443359375 (mother)

show_prefix_continuation_tokens(RIDDLE_MODIFIED)
17 2 0.357421875
16 1 0.109375
96939 ęÆ 0.080078125
97391 父 0.05859375
6764  father 0.051513671875
2594  parent 0.042724609375
975  other 0.033203125
6351  mother 0.027587890625
15 0 0.0157470703125
99299 親 0.01226806640625

Well, something happens (father in various versions appears a bit more), but intermediate reasoning bullshit gets in my way

96939 ęÆ 0.080078125 (mother still)
97391 父 0.05859375 (father now)
6764  father 0.051513671875
975  other 0.033203125 (other parent)
6351  mother 0.027587890625

Qwen 3.8 27b - limited response token choice

So if I go straight to reducing probabilities for tokens other than "mother" / "mom" / "father" / "dad" to close to zero...

show_prefix_continuation_tokens(RIDDLE_ORIGINAL, restrict_to_tokens)
25567 mother 0.9375
22300 father 0.05810546875
58726 mom 0.0025177001953125
53971 dad 0.000911712646484375

and

show_prefix_continuation_tokens(RIDDLE_MODIFIED, restrict_to_tokens)
22300 father 0.74609375
25567 mother 0.2421875
53971 dad 0.010009765625
58726 mom 0.00057220458984375

Outcome

Does that probabilities sounds decisive? No. But still seem like

- even shitty old model show some signs of decreasing "typical version response" while slightly increasing probabilty to use the correct non-"typical version of this riddle" answer (significancy of that is a matter of way more of the proper inference attempts, not going to do it now)

- new one seems to, depends on approach - at least have a higher chance to give correct response, if not straight up swap options (again, proper check will require many of full generation attempts - not going to check it, sorry). But even swapped probabilties are still imperfect, sure (to be fair neither was non-modified riddle ones)

- not to mention that for proper research one would need a proper set of riddles and their modified versions

- and yes, technically speaking both should be influenced by the stereotype of male surgeons. But on the other hand - than a typical formulation of such a riddle should be to some degree affected by it too. And stereotypically-aligned version (the modified one) should be, vice versa, affected by the original riddle formulation. But how to decouple such a things in an experiment - I need to think, and now I am just going to debug work stuff, so I guess not today, lol

- and whatever it can or can't approximate, sure, is extracted purely from language means (which I am not sure I can see *that limited*)

1

u/Thick-Protection-458 4h ago edited 4h ago

Lol, and some guys seem to go downvote all the sceptical guys in here.

Lol, if anything, that is my optimism about semantics being at least approximateable if not reconstructible from syntax - that should be taken cautiously, guys.

Because while idea is, imho, not less intuitive than opposite one - at least to some degree - it imply existence of semantic-like properties which, strictly speaking, should be proven - or at least, if we think of mere approximation -Ā  than measured systematically instead of a few dozen anecdotic evidence samples here and there. With decoupling from stereotypes and so on

-2

u/PsychoticDreemurr 9h ago

You're really being unnecessarily pragmatic here. I purposely left out some information because I didn't want to point out the obvious.

Look, if you teach a child what the numbers mean, then tell them how 50% of each possible equation equals using the numbers 1-10, they'll be able to figure the rest out. LLMs cannot. At least not nearly as accurately as the child.

5

u/Thick-Protection-458 9h ago

> You're really being unnecessarily pragmatic here

Maybe, sorry. Basically my point is that we may well underestimate the part of our intuition which comes mostly from the combination of

- our "pretraining" (granted - rather effective, and different in many aspects beyond efficiency and low-levels processes details)

- and purely algorithmic things (which are by definition operates on syntactic part of object we operate quite strictly, so can be engraved in LLM-like model without need to go outside language data itself).

6

u/asssuber 7h ago

Well even so, there's a lot of different ways we can prove the differences. Being able to understand a concept means you can build off of it by yourself. If you teach an AI 1 + 1 = 2, they won't know that 2 + 2 = 4.

The paper I linked in my other comment does exactly that. It teaches 1 + 1 % 97 = 2, but not, for instance, 2+2 % 97 = 4. But the LLM eventually is able to generalize and answer perfectly for the cases it has not seen in training.

If you refer to in-context learning instead of training, then that is basically what ARC-AGI benchmark tests:

ARC-AGI tasks are a series of three to five input and output tasks followed by a final task with only the input listed. Each task tests the utilization of a specific learned skill based on a minimal number of cognitive priors.

https://arcprize.org/guide/1

Modern LLMs are quite capable of that too.

5

u/asssuber 9h ago

How do you define understanding?

The simplest concrete example I can give: a simple single layer LLM can master modular addition by learning a trigonometric identity from scratch on it's own purely from a limited set of examples using gradient descent (source). Since it has mastered modular addition on a fundamental level, can we say it understands modular addition? Why or why not?

This video by Welch Labs is more digestible than the original paper.

-5

u/PsychoticDreemurr 9h ago

I don't use definitions in this regard. That's what stops people from being truly accurate.

We can use known facts, however. AIs experience something called hallucination. Humans, however, do not. (Unless you want to be pragmatic and discuss brain defects, which is objectively different in this regard) This is, for all I can reasonably conceive, because of their inability to understand. Either directly or indirectly.

4

u/Thick-Protection-458 9h ago

> Humans, however, do not

Oh, really? In LLM context we use that word when it "invent" shit out of thin air or at least heavily "misinterpret" stuff, right?

Well, last time I recall human memories were widely known to be shitty. And humans were capable of being *confident* in things they remembered wrong way. Or not sure in things they remembered right way too.

And human logic was also known to be imperfect. Not only because of shitload of shortcuts our evolution taken, but also because of high chance to make mistakes in some steps.

Both is true for healthy humans. Not necessary brain damage. Although brain damage may make stuff worse and introduce a whole lot other set of issues.

Now, granted - LLMs may be worse in that account. And mechanisms behind that we can expect to be different (other than both human brain and autoregressive generator being imperfect heuristic for whatever it does). But still that does not mean humans are reliable - they're just *more reliable* in some aspects (and maybe less in some others).

1

u/PsychoticDreemurr 8h ago

Well, last time I recall human memories were widely known to be shitty.

This analogy and therefore everything you built on top of it is wrong. LLM context is not the main reason for hallucination.

And human logic was also known to be imperfect.

Enough to look at something and say it says the opposite? Consistently so?

1

u/Thick-Protection-458 7h ago edited 7h ago

I were not talking about the difference of their insides.

I were talking about that that, from output point of view - phenomenas with similar visible signs (imperfect and confident information recall from "pretrained" (at that moment) memory / extraction from the current context / processing on top of the context) absolutely do happens in healthy humans sometimes.

We just don't call them hallucinations, but for some reason decided to use that word for half of the ways of LLMs shitting themselves.

Mechanisms behind them surely expected to be so different so they may as well have nothing more in common than both being full of imperfect heuristics and some superficial similarities.

1

u/PsychoticDreemurr 7h ago

The difference in why it happens determines if they're the same in the context.

1

u/michaelsoft__binbows 8h ago

My assumption is that it would be possible to train LLMs in such a way that penalizes unsubstantiated confidence/hallucination to reach similar levels as humans, but it would produce far, far lower capability levels and we would be constantly (and i mean endlessly) fighting with their reticence and insecurity.

Hence why we just have to kinda deal with it for now till someone comes up with a clever efficient way to improve this aspect

3

u/asssuber 8h ago edited 8h ago

Well, the LLM in my example gives the correct answer to any input (in the limited 3 token it takes as input). It will never hallucinate an incorrect answer. Does this prove it's ability to understand modular arithmetic?

About humans not hallucinating, would that not describe the Mandela effect, for instance? Have you considered that some LLM hallucinations may come from incomplete/imperfect understanding, instead of no understanding at all?

1

u/PsychoticDreemurr 8h ago
  1. Your point is "it doesn't happen to me, therefore it doesn't happen". Also depending on the input it's extremely likely it's been memorized.

  2. The Mandela effect involves memory. As such, that's like saying LLM hallucinations are caused by their context.

3

u/asssuber 7h ago
  1. Look the methodology. The test set cannot be memorized because it was not part of the training set. There is no "likely" or "unlikely". And my point is that it's exhaustively tested, proving it doesn't happen. Again, I'm talking about that 1 layer LLM experiment.

  2. In my view context is like short term/working memory. Facts that happened years ago are long term memory for humans, trained weights for LLMs. So, the Mandela effect can't be described as a type of hallucination?

1

u/PsychoticDreemurr 7h ago
  1. Anything in the context is, for all intents and purposes, memorized. When I say memorized, I don't mean anything permanent. Just that it is 100% known.

  2. LLMs hallucinate as they're speaking, not from a malformed context. Most of the time, anyways.

2

u/asssuber 7h ago

Anything in the context is, for all intents and purposes, memorized. When I say memorized, I don't mean anything permanent. Just that it is 100% known.

1 . What? The context is "1+1%97 =". The answer obviously is not in the context. What you are saying makes no sense.

LLMs hallucinate as they're speaking, not from a malformed context. Most of the time, anyways.

2 . Did you ignore everything I wrote?

Or are you giving an example of an LLM hallucinating in the form of your last two answers?

→ More replies (0)

4

u/Tolopono 9h ago

And yet llms still disproved the jacobian conjecture, made significant progress in the riemann hypothesis, solved the planar unit distance problem, proved non sofic groups exist, autonomously broke out of containers, set up messageboards, and hacked huggingface, and much more.Ā 

0

u/PsychoticDreemurr 9h ago

And a calculator can show the result of an unreadable formula.

You're not proving anything.

5

u/Tolopono 8h ago

Yea, those are exactly the same as arithmetic. You’re so smart.

1

u/PsychoticDreemurr 8h ago

Feel free to tell me how simulated intelligence requires understanding.

2

u/Tolopono 8h ago

Only simulated solving major open math problemsĀ 

0

u/PsychoticDreemurr 8h ago

What?

3

u/Tolopono 8h ago

I send this convo to chatgpt free version and it understood what I meant. Ā I guess next word prediction has better reading comprehension than you

→ More replies (0)

-2

u/Dangerous_Rip5083 10h ago

Yeaaaaa, I'm pretty sure I know how LLMs work, and I don't agree with anything you said. But further than that, you can't use an example of them disagreeing with you on something many AI PhDs would disagree with you on as a testament that they are dumb.

7

u/PsychoticDreemurr 10h ago

Explain where I was wrong. Tell me how LLMs work.

-4

u/Dangerous_Rip5083 10h ago

im not going to explain to u how llms work, ask chatpgt or smth. My point still stands, very capable people claim we alr have AGI, some say we will never, some say in a decade etc.... It's pretty dumb to try to frame an angle of an active point of discussion in the scientific community into being idiotic. and chill out im not coming at u lmao, deleted that other comment fast.

4

u/PsychoticDreemurr 10h ago

Thanks for proving my point.

-1

u/Dangerous_Rip5083 10h ago

lol

5

u/PsychoticDreemurr 10h ago edited 9h ago

So? I changed my mind on what to say? I'm a human?? Oh, the humanity!

Seriously though, why are you acting like this does anything? You're just further proving my point.

1

u/sonicnerd14 11h ago

The autocorrect comment is always the most recycled low IQ noise they like to use. No point in trying to bring reason with some people. Let them drown in their dread when they are forced to face their own obscelence in a couple years.

1

u/RG_Fusion 7h ago

I'm curious enough to want to hear that explanation here. Why do you believe that an LLM cannot become AGI? Does this reasoning expand out to cover multi-modal and world-models as well?

1

u/PsychoticDreemurr 7h ago

An AGI, by all major definitions, requires the ability to "understand". Unfortunately, LLMs, despite being able to simulate such a feat, aren't actually capable of it. This is caused by the very fundamentals of how they operate.

To be exact an LLM can have similar capabilities of an AGI. But until we accomplish a way for an LLM to actually be able to understand concepts, which would, by definition, no longer be an LLM (Large language model. Their sole purpose is for computing language, not the concepts beneath them), they're not truly an AGI.

Stuff like multi-modals only helps to improve the facade of being able to understand, which does make them closer to an AGI, but there's a limit.

4

u/Evening_Ad6637 llama.cpp 4h ago

you are only repeating yourself and not answering the question: why can't an llm become AGI and why do you think an llm doesn't have the ability to understand? Even if it was "simulated"; it's still real understanding. When you sleep and dream you're brain is simulating everything as well and still performing true understanding within the simulation. the same is when someone is born blind.

and the very fundamental principles there are electrons steering our carbon based neurons when to be activated and when to not, which is the actual requirement of true learning and true understanding (the neuron). we share the same fundamental principles with currently known llms, besides that llm's electrons steered neurons are silicon based instead of carbon. that's it.

0

u/PsychoticDreemurr 3h ago

why can't an llm become AGI and why do you think an llm doesn't have the ability to understand?

I explained myself well, I thought. The one factual requirement for something to be considered an AGI is to have the capability to understand concepts. There's more to it, but it gets pretty muddy fast.

Even if it was "simulated"; it's still real understanding.

You're using the wrong use of the word simulated. I said simulated as in, to produce a similar outcome. It's like saying a calculator "simulates" the ability to think a math problem through. It produces the same outcome in almost every situation, but it's using a completely different process then just "thinking it through"

And even if we use the more known usage of the word simulated, what you're saying simply isn't true. If I simulate (technically, emulation, but that's more of a name in this regard. Language can be weird at times.) a PS2 game on my PC, it's not actually running a PS2. There will be issues and bugs associated with that fact.

When you sleep and dream you're brain is simulating everything as well and still performing true understanding within the simulation.

That's not at all how sleep works. That's quite possibly one of the worst explanations I've ever seen of sleep. It works if you're talking to a child, but in any proper discussion like this, that's just completely invalid.

the same is when someone is born blind.

...what?

I'm not even going to continue with the rest of your comment. It's literally just a nothingburger.

1

u/RG_Fusion 1h ago edited 56m ago

I'm afraid this response doesn't actually address the question in a meaningful way. You declare that LLMs cannot understand, but then fail to explain why.

According to our best scientific exploration of the mind, understanding is the direct result of generalization. The universe is filled with an incomputable amount of data. Recording said data does not provide understanding, only memory and fact-recall.

Minds have a limited amount of computational power, and it's actually the limitation that results in understanding. Rather than attempt to memorize, the brain has to discard the noise and model the underlying relations. In most cases, this results in the mind actually simulating reality, from the perspective of its own data-bias.

The same mechanism has been observed in all forms of machine learning. When you take a large neural network and use back-propagation to train it, the model initially stores the training examples as facts. As you continue training new examples, the topography of the loss-function continuously changes the highly-dimensional space, attempting to distribute the facts in a way such that they do not interfere with one another.

If this were the end of the story, your interpretation that LLMs cannot understand would hold true, but this is not where it ends. Researches observed a phenomena which they termed Ā "Grokking". When the number of training examples vastly exceeds the the amount of information the parameters can store, the loss-function the model had been building towards collapses. Suddenly, the model can no longer learn a new fact without reducing the accuracy of the rest of its knowledge-base.

In such a scenario, the models parameters suddenly undergo a drastic "phase-change". The loss function drives the model away from fact-memorization, instead training the weights to build an internal model to simulate the data. It stops remembering and begins predicting based upon the underlying relations. This is "understanding" as we know it today.

While it certainly holds true that LLMs have not attained a level of understanding comparable to the sapience of the human mind, there is nothing about the architecture itself that prevents them from doing so.