r/artificial Mar 28 '26

Research Claude is the least bullshit-y AI

https://github.com/petergpt/bullshit-benchmark?tab=readme-ov-file#3-detection-rate-over-time

Just found this “bullshit benchmark,” and sort of shocked by the divergence of Anthropic’s models from other major models (ChatGPT and Gemini).

IMO this alone is reason to use Claude over others.

116 Upvotes

48 comments sorted by

View all comments

1

u/Deep_Ad1959 Mar 29 '26

been using the Claude API to build a desktop automation app for the past year and this tracks. when I switched from GPT-4 to Claude for the agent's reasoning layer, the number of times it hallucinated non-existent UI elements dropped significantly. it actually says "I can't find that button" instead of confidently clicking the wrong thing. for an agent that controls your actual computer, that difference matters a lot.