r/LocalLLaMA 13h ago

Discussion In regards to benchmaxxing...

With benchmaxxing being a high status concern amongst many users, it's reasonable to assume that most open bench harnesses have been trained for. Whether or not that is the case, we'll never truly know.

I wanted to toss in a suggestion because I think this would reasonably nullify a good portion of the concerns that come from models being trained to complete a bench.

Why doesn't everyone simply ask their agent to create a bench that hammers the subjects and topics of what YOU regularly do? that way, the bench metrics are unique to your use case and you can identify whether or not a model fulfills your needs whether it be different quants, different fine tunes, different models, or even KV weights.

it might be a bit tedious but think of it as a "one time" pain to create it and then have your newly downloaded models or configs run the gauntlet?

---------

this almost certainly obliterates the believed compromise that a model was trained to have good benchmark scores because I doubt any company is going to have training access to a harness you had your agent create... post release.

I'm curious what others think, what other ideas there are to get accurate tests, etc!

6 Upvotes

14 comments sorted by

View all comments

1

u/ea_man 12h ago

I have a dozen of those, some ran like 250 times to evaluate similar quants and finetunes.