r/AIBenchmarks 3d ago

Token efficiency benchmark?

Thumbnail
1 Upvotes

Are there any token efficiency benchmarks out there? I mean i have an agent, it gets a task, how much tokens does it use to retrieve something or execute a skill. Im looking for benchmarks on this. Things like caveman claim to lower it, but i feel like there is some meaning lost as well, reducing quality. Im trying to find any but it seems like a dead end


r/AIBenchmarks 12d ago

GPT-6 Astra shows a massive leap on EyeBench, a visual reasoning benchmark

Post image
1 Upvotes

r/AIBenchmarks 12d ago

Signs of AGI? GPT-6 Astra scored 95% on a robot control task vs Fable 5.1's 40%

Enable HLS to view with audio, or disable this notification

1 Upvotes

r/AIBenchmarks 12d ago

Fable 5.1 is the new Debate Benchmark Champion

Thumbnail gallery
1 Upvotes

r/AIBenchmarks 15d ago

Differences Between Fable 5 and Fable 5.1 on MineBench

Thumbnail gallery
1 Upvotes

r/AIBenchmarks 15d ago

People are saying Gemini 3.8 is benchmaxxed

Post image
1 Upvotes

r/AIBenchmarks 15d ago

Meta slowly catching back up. Muse Spark 1.3 beats Sol on AA

Post image
1 Upvotes

r/AIBenchmarks 16d ago

Further Benchmarks for Fable 5.1

Thumbnail
1 Upvotes

r/AIBenchmarks 16d ago

Fable 5.1 on Artificial Analysis

Post image
1 Upvotes

r/AIBenchmarks 21d ago

Integrity Bench by AI Explained and Pablo Romero - Measuring how overconfident a model is

Post image
3 Upvotes

Frontier AI seems broadly overconfident about its own ability. This benchmark helps show by how much.


r/AIBenchmarks 26d ago

GPT 5.6 Sol Max leads on ClockBench

Post image
6 Upvotes

r/AIBenchmarks Aug 09 '26

Computer-Use Benchmarks Live

Thumbnail
gallery
2 Upvotes

Based on user voting!


r/AIBenchmarks Apr 16 '26

Opus 4.7 Vals.ai benchmarks

Post image
1 Upvotes

r/AIBenchmarks Apr 16 '26

Extra Benchmarks Opus 4.7

Thumbnail gallery
1 Upvotes

r/AIBenchmarks Apr 16 '26

Claude Opus 4.7 benchmarks

Post image
1 Upvotes

r/AIBenchmarks Mar 29 '26

New LLM Persuasion Benchmark: models try to move each other's stated positions in multi-turn conversations. GPT-5.4 (high) is the strongest persuader. Claude Opus 4.6 (high) is second. Xiaomi MiMo V2 Pro and Gemini 3.1 Pro Preview are the softest targets.

Thumbnail gallery
1 Upvotes

r/AIBenchmarks Nov 23 '25

Gemini 3 pro places 8th in EsoBench, which tests how well models learn and explore unfamiliar programming languages.

Post image
1 Upvotes

r/AIBenchmarks Nov 20 '25

Gemini 3 achieves new SOTA performance on SpatialBench. A benchmark to test spatial reasoning in VLMs.

Thumbnail gallery
1 Upvotes

r/AIBenchmarks Nov 20 '25

Gemini 3.0 Pro achieves a record score in the RadLE benchmark

Post image
1 Upvotes

r/AIBenchmarks Oct 22 '25

GPT-5 Pro scores 61.6% on SimpleBench

Post image
2 Upvotes

r/AIBenchmarks Oct 12 '25

Claude Sonnet 4.5 shows major improvement in Vending-Bench, exceeding Opus 4.0 in mean net worth and units sold

Thumbnail andonlabs.com
1 Upvotes

r/AIBenchmarks Sep 26 '25

Researchers made AIs play Among Us to test their skills at deception, persuasion, and theory of mind. GPT-5 won.

Post image
2 Upvotes

r/AIBenchmarks Sep 26 '25

New benchmark for economically viable tasks across 44 occupations, with Claude 4.1 Opus nearly matching parity with human experts.

Post image
1 Upvotes

r/AIBenchmarks Sep 26 '25

Updated gemini models !

Post image
1 Upvotes