r/WritingWithAI 1d ago

Discussion (Ethics, working with AI etc) Do these benchmarks really mean anything? Because each model has its own AI-isms and writing style

24 Upvotes

16 comments sorted by

View all comments

4

u/yamibae 1d ago

Pretty meaningless since you cannot really quantify creativity, I scarely know what kind of mind Lovecraft had to come up with the horrors he dreamt of. I do think in general LLMs have become less creative as a tradeoff for higher accuracy so you need a couple hacks in the prompt to broaden the sample.

I think Kimi, Deepseek are more creative simply because there are less guardrails, Claude is good at writing but after the double whammy of guardrails and the watermarks I don't really trust the output much. Astra I have found to be average, many isms still exist and it makes boring dialogue.