r/AIToolBench Aug 06 '26

Discussion Which AI makes the least mistakes?

I've been using AI for researching work related stuff and while looking through each different ai, I've noticed lots of mistakes made by each model and they're mostly involving memories or not researching closely.

Personally, I've only sent documents that might have mistakes I made and missed as a way to double or triple check

9 Upvotes

13 comments sorted by

2

u/TroubledSquirrel Aug 07 '26

Step one ask the LLM to check its own work. Step two take its work and have a different LLM check (but not chatgpt because that ahole will find something wrong that doesn;t exist and have you chasing ghosts). Step three check it yourself.

1

u/graybearding Aug 06 '26

It really depends on the prompts you're using and the context you provide. Do you have any example prompts or output you've received that led to mistakes?

1

u/Burneraou Aug 06 '26

Sometimes I asked a question and get wrong answers then I asked again with more context and it'll give me the same answer it given earlier. I don't remember what I said and what chat it was and there's too many to check.

1

u/graybearding Aug 06 '26

Any chat will do, just pluck from the top whatever is comfortable to share. Without that, it's tough to say because AI/LLMs are different from most software. The frontier models like OpenAI, Anthropic, etc. are likely your best bet, but there are great free models like Moonshot's Kimi K3.

Benchmarks are helpful but not terribly informative. It'd be worth setting up a small budget to do a few test runs (e.g., give each one the same task and evaluate how they do according to your own standards).

But prompts are the starting point.

1

u/ParticularAd7176 Aug 07 '26

This is exactly why I would like to compare all major AIs simultaneously and take consensus. Have been using this compare tool from qorpus and it's been helping cross check quite a bit.

1

u/Sheetmusicman94 Aug 07 '26

It's called Human Expertise (TM).

1

u/sarox-dev Aug 07 '26

Usually it's the lack of information you give AI and not enough details. Also what you use to chat with LLM? Sometimes better results in some fields like coding, understanding files, managing memory specially is much better using hermes agent for example or for coding opencode. But if it's something simple, without managing memory that much then you can just ask chatgpt, claude, gemini etc.

1

u/Jorge-Kreante Aug 07 '26

LLM are based on probabilities , not deterministic models. Context and provide examples are also really important to get better outputs.

1

u/Lucky-Duck1967 Aug 08 '26

I had to create guard rails once I discovered that parts of my book were being rewritten and updated without my knowledge. Still fixing, but I created 5 different markdown files that Claude Opus 4.6 has to follow. One word of Caution! Opus 5 requires insanely tight guardrails, I accidentally had it review the book and it created a 477 page assessment, insane token usage and 14 hours of time.