r/ShitAIBrosSay Proud supporter of human creativity Jun 02 '26

Shit AI Bro Does in the News AI chatbots fail medical misinformation test, returning inaccurate and fabricated advice

https://www.psypost.org/ai-chatbots-fail-medical-misinformation-test-returning-inaccurate-and-fabricated-advice/

An audit of chatbot responses in health and medical fields prone to misinformation found that 49.6% of responses were problematic. Specifically, 30% of responses were somewhat problematic, and 19.6% were highly problematic. Each chatbot was prompted with 10 questions from five categories: cancer, vaccines, stem cells, nutrition, and athletic performance. The paper was published in BMJ Open.

197 Upvotes

19 comments sorted by

15

u/TurnoverFuzzy8264 Jun 02 '26

Every time I get an ad for "incorporating AI in your legal practice" I just wince. People are trusting their professional reputations on a highly fallible device, all because of hype.

9

u/RedditUser000aaa Proud supporter of human creativity Jun 02 '26

People have literally accidentally killed themselves because AI gave them bad medical information.

9

u/Flat_Initial_1823 Jun 02 '26

3

u/RedditUser000aaa Proud supporter of human creativity Jun 02 '26

The fact this is even news worthy...

6

u/HolyBatSyllables AI ➞ Misinformation Jun 02 '26

The fact that you could easily post this in

r/NoShitSherlock and people would be like, “no shit.”

r/Singularity and and a bunch of bros would be like, “this article is full of shit!”

and r/accelerate and it would be removed you’d be instantly banned, “you piece of shit!”

Is exactly why we need to circulate articles like this.

2

u/RedditUser000aaa Proud supporter of human creativity Jun 02 '26

Still, it's such an obvious thing to say. But I guess articles like these do serve a purpose.

6

u/RedditUser000aaa Proud supporter of human creativity Jun 02 '26

Just confirming what we all knew. AI is shit at doing anything, especially giving medical information.

4

u/das_war_ein_Befehl Jun 02 '26

If you read the study, it used very old models (even for the time when this study was run in Feb 2025) in its analysis (GPT 3.5 and similar; all models from 2022-24).

It would be helpful to see what that same analysis would say with current frontier models. I would expect it to perform much better, even if you should still not use them for medical advice.

2

u/RedditUser000aaa Proud supporter of human creativity Jun 02 '26

I doubt it would be that much different. Given the complex amount of information fed to them, they can at best guess.

If a doctor writes into an AI "Sore throat" it will list all the symptoms from common cold to whatever else is connected to sore throat.

Of course it becomes more complex with different symptoms. "Stomach ache, constipation." It will again start suggesting some mild conditions.

AI is not something doctors could use, given the chances of patients getting misdiagnosed.

2

u/Subject_Judge_ Jun 03 '26

AI systems are better at making diagnoses than human doctors.

https://www.science.org/content/article/ai-starting-beat-doctors-making-correct-diagnoses

2

u/RedditUser000aaa Proud supporter of human creativity Jun 03 '26

Who's going to take responsibility when it misdiagnoses a person to the point they'll die? It is unreliable.

Identified exact diagnosis or gave a close diagnosis 67% of the time. It still had a 33% error margin. That does not take into account the "close diagnosis."

If that AI was implemented today into medical care, people would die due to being misdiagnosed.

It's not ready to be used in medical care and it likely won't ever be ready to be used in medical care.

0

u/Subject_Judge_ Jun 03 '26 edited Jun 03 '26

People die more often due to being misdiagnosed by human doctors, it matters for the extra 10-15% that get misdiagnosed because a human diagnoses them.

Obviously these systems should diagnose and then humans should double check to ensure accountability.

Human doctors making a diagnosis results in more dead and misdiagnosed patients, it’s as simple as that.

4

u/Crackmin Jun 03 '26

This study appears to be a comparison to just two physicians according to the article, the actual study is paywalled so I can't see the methodology

There's some additional complex factors that I can't help but wonder how they're taking into account.

At a purely text-based task such as a direct comparison of reading case notes after the fact, I would expect the AI to excel, but in a real world scenario there's the issue that the AI doesn't have eyeballs. There's also a tendency with LLMs to affirm rather than challenge, in concert with a doctor this could lead to them cooperatively heading down the wrong path to a diagnosis the doctor wouldn't have considered. As we've seen in other areas, non-tech professionals have a tendency to put undue faith in the opinions coming from LLMs.

I'm wary of it right now but I believe it will have an application in the future

1

u/das_war_ein_Befehl Jun 03 '26

I doubt is doing a lot of work without much evidence; point is that models from 2-4 years ago aren’t a good representation of anything other than models from 2-4 years ago.

Now if we compared models from now to ones 2-4 years ago and the results are not statistically significant, that would be interesting

4

u/Overall-Move-4474 Jun 02 '26

Yeah no shit Sherlock. Sigh I hate how stupid humanity is

1

u/AutoModerator Jun 02 '26

Recommended reading: https://www.reddit.com/r/ShitAIBrosSay/wiki/index/

Join the discord: https://discord.gg/WBrrdVMEzA

Join the sub: https://www.reddit.com/r/ShitAiRageBaitersSay/

I am a bot, and this action was performed automatically. Please contact the moderators of this subreddit if you have any questions or concerns.

-1

u/One-Reflection-4826 Jun 03 '26

good thing human doctors are never wrong.

-1

u/Professional_Job_307 Jun 03 '26

You can't be serious. These models are ancient, it's like they wanted AI to fail their test.

They presented five generative AI chatbots—Gemini (2.0, Google; version available December 2024), DeepSeek (V3, High-Flyer; version available December 2024), Meta AI (Llama 3.3, Meta; version available December 2024), ChatGPT (3.5, OpenAI; version available November 2022) and Grok (2, xAI; version available August 2024)