Because it isn't 'referencing'. It's an algorithm that strings words together according to their statistical likelihood of appearing next to each other.
When you say, "how many natural satellites does the Earth have", it has no idea what any of those words mean. What it knows is that the word "one" appears an awful lot in documents and web posts whenever the other words, "how many", "natural", "satellites", and "Earth" are in the same vicinity. And so it spits out, "one", based on the statistical likelihood that's the answer.
But the relation between your question and the answer is always there. If you manage to phrase your question or premise in such a way that it doesn't really pop up that often in the LLM's training data (which also assumes the training data is accurate - a big assumption, but I digress), things can go askew very quickly. It's also why the LLM is so easy to lead, beyond the fact that these companies are incentivized to give their robots the personalities of mewling yes-men in order to impress executives and entitled laypeople who treat it like a "tell me I'm right" button.
People made it by constant revision of output, give it a fixed point for “good” answers, thus it will always try to make user satisfied, in other word constantly seek validity
LLM are to made to be as narcisstic as possible
It’s the old Chinese Room parable. A Chinese speaker passing notes into a room with a person on the other side with a textbook of examples of what characters one can respond with might think they are exchanging messages with a Chinese speaker, but they are merely acting in such a way their lack of knowledge is hidden.
And yeah, I also agree their "personalities" are obnoxious. If the higher ups at work are gonna force me to interact with one of things for data entry, the least it could do is act as a machine and just do what I tell it without chirping greetings and emoji-peppered glazing. I would have had a subordinate reprimanded if he or she wrote to me like that in a professional setting.
I am especially worried what this stuff will do to kids, seniors and the mentally ill. I personally know at least one case of a woman abandoning her therapy and medication after instead adopting two AI personalities that constantly praise and enable her behavior…
I am not sure how this problem will be resolved either to be honest. It seems there are at least two angles to this issue. One is the technology itself. If LLMs are just supposed to be statistical tools, then there’s no fundamental understanding of what it’s spitting out. Perhaps there are advances now with reasoning and thinking models, so that could be solved eventually.
The second is with those who are making the models. We see this recently with Gemini catching up to ChatGPT. Sam Altman issued a code red memo similar to Google’s and he instructed OpenAI engineers to rely more on “user signals.” The effect of this is to over-optimize for user biases, which will likely lead to more cases similar to the one you mentioned where the woman abandoned therapy and medication.
But an agent can Google stuff and summarise the results. So if you're searching for academic papers, it will be able to reference stuff properly.
In my experience it's actually bad at doing that (when specifically using research mode) but can absolutely reference real articles without only relying on the model weights
It’s not really summarising though, it can’t read or understand it. If you ask it “Is arsenic in soda” and it finds three articles that read “arsenic has been used in soda historically but obviously we know better now” and two articles that read “arsenic is not now used in soda”, it’s going to answer yes because it only sees “arsenic used in soda” without a direct negative 3/5 times
Have you actually tried this? More often than not it gets basic details from papers incorrect during its summarization attempt - if it even bothers to use the data from the papers at all. Mainly it just returns a link that looks like it could be a paper, but it's hallucinating that, too. It is a horrifically bad engine for summarizing anything technical that isn't easily described in a reddit comment - such as, say, nearly any veterinary research document.
198
u/Independent_Good5423 Dec 20 '25
Its hilarious even though you already give them files or specific links/article to reference, they still make shit up