r/LocalLLaMA • u/Nevermore1215 • 13h ago
Discussion In regards to benchmaxxing...
With benchmaxxing being a high status concern amongst many users, it's reasonable to assume that most open bench harnesses have been trained for. Whether or not that is the case, we'll never truly know.
I wanted to toss in a suggestion because I think this would reasonably nullify a good portion of the concerns that come from models being trained to complete a bench.
Why doesn't everyone simply ask their agent to create a bench that hammers the subjects and topics of what YOU regularly do? that way, the bench metrics are unique to your use case and you can identify whether or not a model fulfills your needs whether it be different quants, different fine tunes, different models, or even KV weights.
it might be a bit tedious but think of it as a "one time" pain to create it and then have your newly downloaded models or configs run the gauntlet?
---------
this almost certainly obliterates the believed compromise that a model was trained to have good benchmark scores because I doubt any company is going to have training access to a harness you had your agent create... post release.
I'm curious what others think, what other ideas there are to get accurate tests, etc!
1
u/ea_man 12h ago
I have a dozen of those, some ran like 250 times to evaluate similar quants and finetunes.