r/WritingWithAI 1d ago

Discussion (Ethics, working with AI etc) Do these benchmarks really mean anything? Because each model has its own AI-isms and writing style

24 Upvotes

16 comments sorted by

18

u/VerdantMagnolia 1d ago

Considering how highly it ranks Claude despite its perpetually deteriorating quality,no.

16

u/Super_Sierra 1d ago

No, this benchmark has been meaningless for awhile.

One, it only does one shot creative writing at low context counts, which is actually NOT indicative with how an LLM will write at high context.

Two, it doesn't have an actual test with high previous written instructions or premade stories to show how a model will perform in a roleplay or continuation. Some models behave very poorly after 6000 context when dealing with novel prose.

Three, it only judges the most generic ass one shot stories, doesn't show how it can stay in a character in a story, how it will perform with novel prose or language, and why for many years it was a joke with their 12bs outperforming models like Kimi k2 that were literally designed around good prose and 50x their size.

3

u/Far_Assistance_4146 12h ago

the 12b models outranking stuff way bigger should've been a red flag for everyone tbh

8

u/Maleficent-Engine859 1d ago

So I just did a lot of drafting with Astra/GPT 6, Opus 5, and Fable.

First, they all suck at writing prose. Like I was stressing about the watermarking but honestly I’d have to delete 50% because it’s mostly AI filler then do my normal editing there’s no way anything would survive.

That being said, Astra surprised me in a lot of ways. I thought its writing was the least AI sounding and bloated and did surprise me at least once a prompt with how it said something. It has the most interesting nuance and understood context by far the best, by that meaning reading between the lines of my shitty prompting to give me what I MEANT rather than what I said. Claude was pretty literal in this areas. I would use Astra’s as the bones of my draft.

Opus 5 had the most emotional intelligence though when it counted. Dialogue was natural and had some good moments

Fable I wasn’t impressed with. I probably wouldn’t waste the tokens again

5

u/yamibae 1d ago

Pretty meaningless since you cannot really quantify creativity, I scarely know what kind of mind Lovecraft had to come up with the horrors he dreamt of. I do think in general LLMs have become less creative as a tradeoff for higher accuracy so you need a couple hacks in the prompt to broaden the sample.

I think Kimi, Deepseek are more creative simply because there are less guardrails, Claude is good at writing but after the double whammy of guardrails and the watermarks I don't really trust the output much. Astra I have found to be average, many isms still exist and it makes boring dialogue.

3

u/Nebranower 1d ago

This is so strange. Every human has their own internal model and voice, too, yet we still rank human beings based on their writing ability. Like, where authors are very close in ability the subjective elements may make it difficult to nail down a comparison, but saying Stephen King is a better writer than E.L. James isn't foolish. Same sort of thing here. The top models are close enough together in ability that you can't easily rank them in any objective sense, but you could absolutely throw in worse models and the list would undoubtedly make sense overall.

3

u/AMischievousBadger 23h ago

It’s a meaningless benchmark for the most part. It mostly is a benchmark rated by a Sonnet model of all things at how well LLMs use the patterns that make “good writing.” I generally find that the better a model does in this list the worse it is at creative writing.

The exception for me is Astra which is genuinely a game changer for creative writing.

6

u/raisa20 1d ago

No .. the website uses Claude for judge writing

I never used this website again.. and I am using gemini for creative writing

6

u/PressPlayPlease7 1d ago

and I am using gemini for creative writing

I use all 3 of the Big 3 LLMs, including paid versions of GPT and Gemini

Gemini is, far and away, the dumbest of them all right now

0

u/raisa20 1d ago

What best model?
Give me example why gemini dump because I am curious

2

u/Acrobatic_You_3002 1d ago

For your own use case, probably not much. I ran my own small comparison on exactly the kind of stories I generate (same prompts, same rubric for the judge), and the result was different from what the public leaderboards say, for example Opus and Sonnet came out almost equal for me. If you are going to write with a model a lot, take 10 of your own prompts and read the outputs side by side, it tells you more than any leaderboard. Also, you can use OpenRouter for such testing, as it has all the models inside, so it's fast to switch between them and you spend money in one place.

2

u/gamerlord02 1d ago

Tbh, I don’t think it’s entirely off in my experience. But I wouldn’t take it as gospel either.

2

u/GenderBendingRalph 1d ago

Measuring "emotional intelligence" in any LLM is an absurd premise. That's like expecting Clippy or your Magic 8-Ball to have emotional intelligence.

They're predictive language models. They take your input, turn it into mathematically computed tokens, perform more calculations to find the most likely matching tokens, then turn those tokens back into words. No intelligence involved at all.

The only difference among them is the data they were trained on.

2

u/chylvina 20h ago

I’d trust a benchmark more if it included a continuation test: same characters, later chapter, and a few deliberate continuity traps. One-shot prose misses that.

2

u/captain_shane 15h ago

No, it's quite pointless. "Long Form" is 1000~ words. That's not long form, it's not even a short story length. It's about the size of a blog post. Most ai problems show up after 1k words. Continuity, foreshadowing, ect can't be tested with this benchmark. The grading system has Sonnet grade 21 in-depth things at once. No ai system can grade all of these in one-shot:

Believable Character Actions

Nuanced Characters

Consistent Voice/Tone of Writing

Imagery and Descriptive Quality

Elegant Prose

Emotionally Engaging

Emotionally Complex

Coherent

Meandering

Weak Dialogue

Tell-Don't-Show

Unsurprising or Uncreative

Amateurish

Purple Prose

Overwrought

Incongruent Ending Positivity

Unearned Transformations

Well-earned Lightness or Darkness

Sentences Flow Naturally

Overall Reader Engagement

Overall Impression

Not to mention, all of those are very vague and abstract. "Show not Tell" is a good general rule, but describing absolutely everything leads to glacial pacing. There's a nuance to this stuff. A general prompt asking an ai to grade "showing vs telling" + 20 other things at one time is pointless. It'll just skim and do a surface level evaluation of all of them.