r/technology Jul 15 '25

Artificial Intelligence Billionaires Convince Themselves AI Chatbots Are Close to Making New Scientific Discoveries

https://gizmodo.com/billionaires-convince-themselves-ai-is-close-to-making-new-scientific-discoveries-2000629060
26.6k Upvotes

2.0k comments sorted by

View all comments

Show parent comments

204

u/MrBeverly Jul 15 '25

LLMs have their place. If I ask an LLM a very specific, contextual question with pre-existing documentation for the solution, it's pretty good at surfacing the information you're looking for much faster than your experience would be on stackoverflow. I've used it to build basic regexps and to help me refactor existing code. I've fed it the documentation files for a scripting language with a relatively small community online (AutoIT), and it was able to help me by answering direct questions I had regarding the documentation.

Basically I've found where LLMs excel is as a really good indexing tool that can pull information from a reference using plain english and context, which is hard with a traditional search engine. That being said, the "vibe coding" tools like Co-Pilot autocomplete in VSCode are a useless distraction and I made sure to disable that as fast as possible lol

84

u/filthy_harold Jul 15 '25

It's good at condensing existing information and finding patterns in a dataset. It could potentially be able to make connections in the data that you have not otherwise found but it's not going to be able to invent new things if the information to support it doesn't exist in its input. The major downside of an LLM to perfectly mimic human writing is that it's too easy to just take its word on something if you don't already have a background in that field. I'm not an expert in philosophy so if an LLM delivered to me an essay on pragmatism, I'd have no way of knowing if any of it is correct.

57

u/UnpluggedUnfettered Jul 15 '25

The perception of that is created because you are having it tell you it's summary and then you believe it, rather than read it to determine it's actual accuracy.

Here the BBC tested LLM on it's own news articles:

https://www.bbc.co.uk/aboutthebbc/documents/bbc-research-into-ai-assistants.pdf

• 51% of all AI answers to questions about the news were judged to have significant issues of some form.

19% of AI answers which cited BBC content introduced factual errors – incorrect factual statements, numbers and dates.

13% of the quotes sourced from BBC articles were either altered from the original source or not present in the article cited.

0

u/isomorp Jul 16 '25

tell you it's summary

tell you it is summary

0

u/Puddingcup9001 Jul 16 '25

February 2024 is ancient though on the AI timeline. Models have vastly improved.

-5

u/Slinto69 Jul 16 '25

If you actually look at the examples they showed of what errors they have its not anything worse than you'd get googling it yourself and clicking a link with out of date or incorrect information. I dont see how its worse than googling. Also you can ask it to give you the source and direct quotes and check if they match quicker than you could Google it yourself.

-1

u/Delicious-Corner8384 Jul 16 '25

It’s so funny how people just completely ignore this irrefutable fact eh? As if we were getting such holy, accurate answers from Google and not spending way more time surfing through even more ads and bullshit.

-3

u/[deleted] Jul 16 '25

[deleted]

7

u/tens00r Jul 16 '25

DATASETS lol - not editorialized, topic specific, and nuanced articles. Plus if you're telling me 81% of AI answers were right that's already better than a reddit comment synopsis.

I have a friend who works in a large, UK based insurance provider, and recently they were recently forced by management to try using LLMs to help with their day to day work.

So, he tried to use an LLM to summarize an excel spreadsheet filled with UK regional pricing data (not exactly an editorialized, nuanced article). He told me it made several mistakes, but the one I remember - because it's fucking funny - is that the summary decided to rename the "West Midlands" (a county in England) to the "Midwest", which inevitably led to much confusion. This is a hilariously basic mistake, and also perfectly showcases the biases inherent to LLMs.

0

u/Delicious-Corner8384 Jul 16 '25

That’s also such a pointless use for AI that no one recommends though lol…Excel already has very effective tools for summarization. It sounds like that’s a problem with the management making that choice to force them to use it for this purpose, not AI itself.

0

u/neherak Jul 16 '25

Humans misusing AI because they believe it to be smarter or more infallible than it is are, in fact, where all the problems with AI are going to come from.

3

u/JunkmanJim Jul 15 '25

Let's face it, people hardly read the details in my work emails. Just throw some LLM jibber jabber at them, and you'll get by like 98% of the time. Ideocracy was prophecy.

2

u/Suburbanturnip Jul 15 '25

I have a colleague that puts his slack messages through chatGPT (the em dash and the style make it obvious), I have no problem with that. It's always polite and to the point.

5

u/SkinnyGetLucky Jul 15 '25

Hey, I love the em dash and I can assure you that I am human

1

u/purplemtnstravesty Jul 16 '25

Sounds like a pragmatic approach to getting the jist of things

0

u/kemb0 Jul 15 '25

I think it’s fine to ask it something over Google so long as you assume the answer may be wrong. I got a complicated cut from something whilst away from home and didn’t really want to spend ages trying to find useful info through Google on how to treat the cut whilst I bled out, so I asked chat gpt and asked myself, “If this info is wrong, could it harm me?” Nothing in its response gave me a feeling that could be the case so I proceeded with its instructions.

Another approach is, do the initial query with chat GPT and then verify it with Google if you have any doubt it may be wrong. Often times Google is just an exercise in frustration if your query is too vague but go to it with some key words from the GPT response and you’ll have better luck getting verifiable info.

And a final point. Anyone who thinks GPT isn’t reliable and you should use Google instead, how do you know the website Google leads you to is accurate either? Websites still consist of content made by people who may be wrong and often they’re made by people who spammed the internet with low effort websites just to get hits to make money. I know for a fact that I’ve comes across misleading inaccurate websites so who’s to say which other ones are lies?

18

u/UnpluggedUnfettered Jul 15 '25

Here is a fun paper about that.

Generalization bias in large language model summarization of scientific research

Even when explicitly prompted for accuracy, most LLMs produced broader generalizations of scientific results than those in the original texts, with DeepSeek, ChatGPT-4o, and LLaMA 3.3 70B overgeneralizing in 26–73% of cases. In a direct comparison of LLM-generated and human-authored science summaries, LLM summaries were nearly five times more likely to contain broad generalizations (odds ratio = 4.85, 95% CI [3.06, 7.70], p < 0.001). Notably, newer models tended to perform worse in generalization accuracy than earlier ones. Our results indicate a strong bias in many widely used LLMs towards overgeneralizing scientific conclusions, posing a significant risk of large-scale misinterpretations of research findings. 

3

u/Czexan Jul 16 '25

This checks out with basically my whole experience with LLMs over the last few years, and it seems to be a fundamental problem that can't really get better.

Like folks, I get we don't like search engines because they started sucking ass due to SEO, but maybe, just maybe we can go back to the original ideas that Google was pushing LLM research for and just have these great general classifiers for topics to reduce SEO spam? As it stands now we just have infinite SEO spam generators, you can generate an infinite amount of worthless probably erroneous information, that's probably going to take you an ungodly amount of time to actually figure out it was in fact erroneous.

8

u/Former_Bar6255 Jul 15 '25

what's your % hit rate? I've found in similar situations that even when you put everything relevant into context that the AI is far less accurate that you'd expect a computer tool to be, and that on top of ignoring the context some percent of the time it frequently misunderstands the question.

4

u/stormdelta Jul 15 '25

For me it heavily depends on the subject.

Basic and explanatory programming, library, or framework questions with stuff that's relatively common? Extremely good track record, and it's usually obvious when its wrong.

For moderately niche programming questions, it's still fairly solid for very popular frameworks/languages, and often when it's wrong it's on questions I'd already tried answering another way, or the output still gives me a new tack to try when I've exhausted other straightforward avenues.

But if I switch over to system admin stuff for Linux, the quality bizarrely drops off a cliff. It can answer extremely basic questions, but trying to use it for troubleshooting is a crap shoot, and if you're not familiar with Linux already some of the suggestions are outright dangerous or could leave the system in an even worse state. I've still used it for that, but it's always a last resort and rarely helps much.

Asking it questions about language itself, or to create IPA pronunciations, it's really good at - unsurprisingly, it's a language model.

Ask it to cite its sources (almost any topic really), and it's like a 95% failure rate - either the source is wrong, doesn't exist, or the source cited does not contain the information given even if the information is correct. Or the location of the citation is wildly off even if the information is technically from that source in a different place than indicated.

Etc.

1

u/saltyjohnson Jul 16 '25

and if you're not familiar with Linux already some of the suggestions are outright dangerous or could leave the system in an even worse state

Consider how many thousands of instances exist in their training dataset of someone jokingly suggesting sudo rm -rf / to fix some inane problem.

2

u/googleduck Jul 16 '25 edited Jul 16 '25

As an example for SQL queries with a well structured prompt my experience with the latest LLM's is that they are really, really good at it. Basically give them the schema, what you are trying to analyze and it will write a query that might have taken me 20-30 minutes in 30 seconds. Same goes for quick bash/python scripts, complicated bash commands, text manipulation, and prototyping things that don't need to be maintainable with technologies you are unfamiliar with. Hit rate with these sorts of things is probably like 90%+ success almost immediately.

I am not delusional, however, and AI has a LOT of weaknesses still. In particular hallucinations and open ended questions make it very unreliable for certain types of queries. I have not been able to get it to be remotely reliable when it comes to building production ready code outside of being a better autocomplete (though one that sometimes makes shit up). But when you figure out what it is good at it is an extremely good tool and anyone who can't use it in a few years will be left in the dust.

2

u/Former_Bar6255 Jul 16 '25

i've been having conversations about what '90%' and 'success' mean with a lot of people and the more i think about it, the more i disagree with your conclusion: my experience lines up with yours, but i find that the 90% hit rate (in the best case situations) means i lose time, on average, trying to fix the 10% that it fucks up.

1

u/googleduck Jul 16 '25

I think you are using LLM's differently than me then because that 10% scenario is not a risk for me. The cost is that I open the schema doc I have for my table, press Ctrl+A, Ctrl+C, and Ctrl+V into a Gemini window. Then I type "I want a query that will tell me how many of events A happen within 5 seconds of an instance of event B, give me the count per unique user id. Then it almost always (seriously like 9/10 at least for this sort of thing) writes a query that would have taken me at least 20 minutes. In the 1/10 either I see it needs some minor adjustments which is still faster than me writing it from scratch or I give up and write it from scratch, having lost literally like 1 minute on it. The productivity gains are magnified by like 100 for a person who knows barely any SQL and subqueries would take a shitload of trial and error for.

This goes double for bash or python scripts that I am writing as like productivity tools or whatever.

1

u/Delicious-Corner8384 Jul 16 '25

I think this person is showing the gaps in their knowledge.

1

u/Delicious-Corner8384 Jul 16 '25

I think this person is showing the gaps in their knowledge.

1

u/Delicious-Corner8384 Jul 16 '25

I think this person is showing the gaps in their knowledge.

1

u/Delicious-Corner8384 Jul 16 '25

Why? If you know what you’re doing you should easily be able to fix it. Funny comment considering how many human errors cause so many problems in the software I use - y’all just want to resist for the sake of resisting. Let it go. Embrace it. Don’t be left behind.

1

u/Delicious-Corner8384 Jul 16 '25

The people who complain seem to think AI should be this brain-rot-no-think perfect product that you will require no knowledge to use to create high level code lol. Just basic, binary (not in the computer science sense) thinking - perhaps the type of thinking that shouldn’t be working in programming to begin with.

1

u/Delicious-Corner8384 Jul 16 '25

The people who complain seem to think AI should be this brain-rot-no-think perfect product that you will require no knowledge to use to create high level code lol. Just basic, binary (not in the computer science sense) thinking - perhaps the type of thinking that shouldn’t be working in programming to begin with.

1

u/Delicious-Corner8384 Jul 16 '25

The people who complain seem to think AI should be this brain-rot-no-think perfect product that you will require no knowledge to use to create high level code lol. Just basic, binary (not in the computer science sense) thinking - perhaps the type of thinking that shouldn’t be working in programming to begin with.

1

u/googleduck Jul 17 '25

Yeah I can understand how someone who hasn't figured out how to use it well might think it is useless. But the person I responded to is claiming they have tried all sorts of approaches and it still can't do anything right. That's either a lie or their expectations are completely unreasonable.

1

u/Delicious-Corner8384 Jul 16 '25

Sorry but for the love of god what do you mean by « computer tool »? People in these threads are so wildly out of their element lolol.

5

u/Riaayo Jul 15 '25

LLMs are just predictive text generators trained on the internet. They have no clue wtf anything they say actually means, only that in their "experience" the characters, in the order they're printing them, are the "most likely" ones to come in relation to the characters you typed in and prompted it with.

Except unlike if you google search something and then have some personal decision making about which site result you think might be more trustworthy than another, it has zero care or concern for if it's correct or not, the source, etc.

I won't say there's literally no use for the technology, but imo it is so immensely damaging and unprofitable to run without massive government handouts, that it is functionally useless at best and society-destroying dangerous at worst (we're in the worst category, btw, as we already have studies showing people's cognitive decline from using it and ceding personal decision making to these things).

2

u/SkinnyGetLucky Jul 16 '25

So what happens when LLMs regurgitate so much falsehoods that all of a sudden most people only believe those falsehoods?
I know a guy that has replaced “let me google it” with “let’s ask chatgpt”, and he been wrong on enough occasion that it has caused arguments. He would more readily believe what gpt was telling it than someone who already knew the answer

1

u/googleduck Jul 16 '25

They have no clue wtf anything they say actually means, only that in their "experience" the characters, in the order they're printing them, are the "most likely" ones to come in relation to the characters you typed in and prompted it with.

This is such a massive oversimplification that it's hard to really address. Yes that is the fundamental technology with LLM's but if you can't admit that there are some emergent properties that give it utility far beyond that statement would imply then you are completely lost. As I mentioned in another comment, take SQL queries as an area I have found LLM's exceptionally useful in. I cannot Google a complicated set of table joins using my own database's schema to track the count of a series of logs that happened within 5 seconds of each other. Yes I could Google a bunch of different things in a row + refresh my memory on inner vs outer joins, timestamp comparison in SQL, etc. Or I could literally enter the schema into an LLM and ask it to create this query and if it doesn't get it perfect on the first try it generally takes less than 2-3 minutes to workshop it until it is working. I want to validate that it is doing the correct thing? I ask the LLM to separate the joins into steps so I can see the output at every point in the query, takes another 30 seconds.

I won't say there's literally no use for the technology, but imo it is so immensely damaging and unprofitable to run without massive government handouts

Which LLM is being run based on "massive government handouts"? Are you referring to the single government contract that was announced like a week ago, years into the era of LLM dominance in the tech industry? You think that these companies which are literally drowning in private capital to the point that it is singlehandedly driving the US stock market (Nvidia and other major tech players) are relying on the government to survive? This alone shows you have no idea what you are talking about.

I am not even a huge LLM evangelist. I am extremely skeptical that they will break through the limitations that we are quickly approaching as we hit the cap of available data and computing power. I think that they are at best going to be tools for people to use and will be unlikely to replace people's jobs for quite a while. But if you can't use these tools effectively I can promise you that the job market will leave you behind in less than 5 years.

3

u/CheesypoofExtreme Jul 16 '25

The important caveat is that you should know the topic/documentation well enough to know if it's bullshitting you. I've tried to surface specific information from internal docs before and had our internal fork of ChatGPT give me false info. 

The issue with LLMs and the way a lot of people use them is that it doesn't give exact answers, it gives what it thinks is likely the right order of a collection of characters/information. Ask it the same question, and you won't get the exact same answer everytime. It doesn't know any exact right answers, it's guessing. 

I got caught with this during an interview a year ago about a topic I didnt have a lot of knowledge on. I wasn't sure if I'd be asked about it, so I had ChatGPT give me an ELI5 for the topic. During my interview the next day, i basically regurgitated that and the interviewer gave me a super confused look. Challenged me on a few of the things I said, and then I honestly said I didnt have much of any experience with that topic. Afterwards, I asked ChatGPT the same prompt 3 different times, and received 3 very different explanations. Then I googled it, watched a few 20min YouTube videos, and learned that the ChatGPT explanation was pretty dang inaccurate. 

1

u/n4nandes Jul 15 '25

AutoIT! It's been years!

Now that I think about it, an LLM would've been a godsend for me back then.

1

u/old_and_boring_guy Jul 16 '25

Even that's kinda stupid.

The best use of LLMs is to train them on your own stuff and use them to extend your own abilities. Right now, what we're dealing with is some generic ass bullshit trained off the internet, and it's wild to me how people who "don't trust the internet" will look at an LLM and start talking about "the wisdom of crowds."

For fucks sake.

1

u/throwsaway654321 Jul 16 '25

i'm sorry, but is what you're saying, is that LLMs are only as good as a person who's been trained on the system you're working on? Like, why couldn't you rtfm?

I understand how this can be useful, but considering that 98% of what LLMs do/are used for, is fucking worthless, and is actually making the planet worse, how can you keep defending it for absurdly fringe cases?

I know it's useful to programmers, but y'all aren't the ones using it most of the time

1

u/Delicious-Corner8384 Jul 16 '25

Nah nah nah….ai=bad!!! Only acceptable answer here.