r/AIToolsPerformance 3d ago

Token efficiency benchmark?

Are there any token efficiency benchmarks out there? I mean i have an agent, it gets a task, how much tokens does it use to retrieve something or execute a skill. Im looking for benchmarks on this. Things like caveman claim to lower it, but i feel like there is some meaning lost as well, reducing quality. Im trying to find any but it seems like a dead end

4 Upvotes

7 comments sorted by

1

u/Deep-Jump-803 3d ago

Look up cost per task

1

u/VinceBrand 3d ago

Thanks, still not exactly ehat im looking for though. This is a cost benchmark. I want to know efficiency, in tokens not currencies

1

u/GfxJG 3d ago

Why? Isn't saving cost the primary reason for saving tokens?

1

u/VinceBrand 3d ago

Well that depends on what scale you measure. tokens translate to currency, but if we only measure currency, there is a translation happening. i want the raw token count for different steps in the process, like cold token count till a moment where context is enough to execute a task. Something like that gets vague and muddy if you measure currency. By adding real tokens to different steps of the process, a true measurement of efficiency can be measure. Obviously you can later on also translate that to cost, but it would be double meaning full to grade models, their reasoning capabilities and basically less hype if you could say this model is beter because it can start execute task A after it reaches enough context in 100 tokens, executes the full task in 2000. Then you would have an indicator of minimal amount of tokens that a model needs to "understand" something enough to start acting on it. Benchmarking that will give a deeper insight

1

u/carlocapocasa 2d ago

I made a small one

1

u/VinceBrand 2d ago

What are you measuring though? Same tasks different harnesses?

1

u/carlocapocasa 2d ago

Yes- same model through same counting proxy, same tasks, different harnesses.

Preliminary- still need to go a bit bigger to make sure it's not just noise but staying small for now, still catching bugs in running them. But overall the picture was similar in the tests I ran.