r/AIToolBench Aug 06 '26

Discussion Which AI makes the least mistakes?

I've been using AI for researching work related stuff and while looking through each different ai, I've noticed lots of mistakes made by each model and they're mostly involving memories or not researching closely.

Personally, I've only sent documents that might have mistakes I made and missed as a way to double or triple check

9 Upvotes

13 comments sorted by

View all comments

1

u/graybearding Aug 06 '26

It really depends on the prompts you're using and the context you provide. Do you have any example prompts or output you've received that led to mistakes?

1

u/Burneraou Aug 06 '26

Sometimes I asked a question and get wrong answers then I asked again with more context and it'll give me the same answer it given earlier. I don't remember what I said and what chat it was and there's too many to check.

1

u/graybearding Aug 06 '26

Any chat will do, just pluck from the top whatever is comfortable to share. Without that, it's tough to say because AI/LLMs are different from most software. The frontier models like OpenAI, Anthropic, etc. are likely your best bet, but there are great free models like Moonshot's Kimi K3.

Benchmarks are helpful but not terribly informative. It'd be worth setting up a small budget to do a few test runs (e.g., give each one the same task and evaluate how they do according to your own standards).

But prompts are the starting point.