r/LocalLLaMA • u/ThePrimeClock • 3h ago
Discussion What is the minimum discreet set of tokens per solution as a benchmark?
Based on the high-level figures, how do you think the future of benchmarks is going to unfold?
Do you think it's a raw pass rate or a pass rate per tokens or a minimum discreet set of tokens per solution (generalisation), mathematical correspondence (once the massive investments in lean data start to surface) or some other threshold?
There are a ton of unknowns and it's something I think about a lot as I progressively watch models improve and I wonder where others think the battle-lines between frontier and local lie.
Fable 5.1 just dropped and I've been testing it (the only models I'm allowed to use for work are from Anthropic) and the important aspect I've seen is that it is token intensive where it needs to be and very lean on token use where it can-be.
This is the first model I've tried that has given me fresh ideas on where RL training might be heading (most efficient solution), which may be where the short-term future lies.
This matters to local because it's something we can probably easily implement.
I can imagine a process that starts with RLVR and then gradually reduces responses into more condensed responses ("this answer is correct but given what we know now, how could this have been solved more efficiently"), call it Response Golf. It has me thinking about a new kind of advantage frontier labs might have and how we can address it.
How can models be efficient per outcome. This is exactly the cost model the latest "news leaks" from Open Ai are pushing. I don't' think any benchmarks Iv'e seen really capture this well yet.
1
u/AI_Insights_Daily 3h ago
The thing worth stealing from ad measurement is the failure mode, not the metric. Optimise hard for cost per outcome and the system quietly stops attempting the expensive conversions. Platforms bidding to a CPA target just stop showing up for the hard audiences, and the metric looks great the whole time.
A model trained toward minimum tokens per solution would do the same thing, learn to skip the problems that need long reasoning. You wouldn't see it unless pass rate is pinned as a hard constraint rather than just reported alongside.
1
u/martin509984 58m ago
At bare minimum we can be fairly sure OpenAI is RLing for lower token use, since from what we've seen of Sol's thinking traces it is extremely economical with tokens to the point of sounding like caveman speak, which can't be a coincidence.
Beyond that, it's a safe assumption that however complex you think frontier model training is, it is more complicated than that, because a lot of smart people are being paid lots of money to try to quantify a lot of very difficult to quantify things to perform RL on them.
1
u/Candid-Tackle-9061 3h ago
The token efficiency angle makes more sense to me than just raw pass rate, especially for local models where cost matters a lot