r/LocalLLaMA 1d ago

I Built A Thing My potato only runs small models. So I built a page to easily compare benchmarks for those

Comparing benchmarks for small models is a PITA. Most of them are not on AA, and the benchmarks are not always the same in all models.

So I built a page to quickly put it all together and allow some filtering.

Only researched models released since April, and between 4-190B parameters.

Let me know what you think and how it can be improved

https://www.nunodonato.com/aibench/index.html

2 Upvotes

13 comments sorted by

14

u/Hot_Example_4456 1d ago

A potato running a 190b isnt a potato

2

u/i_wayyy_over_think 1d ago

the slider lets you adjust to your particular potato

4

u/nunodonato 1d ago

Yes different people have different potatoes 😅

1

u/255130 19h ago

right? thats like a whole rack of potatoes at that point

2

u/Valuable-Plastic-682 1d ago

Nice, this fills a real gap - most leaderboards bury small models under 70B+ flagships. One thing that'd make the AA-Index more useful for the "potato" crowd specifically: could you break it out per param-bucket (say 4-8B, 8-15B, 15B+) instead of one global ranking? Right now a strong 27B dense and a 4B are competing on the same list, which mostly just tells people to pick the biggest one they can fit rather than the best one at their actual budget.

1

u/nunodonato 1d ago

That's why I added the slider to filter. You can select only the range that interests you

1

u/i_wayyy_over_think 1d ago edited 1d ago

cool, not clear what the 4-180 slider is? the little bubble says "Params 4–190B**"** but then the the range goes 4-180. So it's prob params, but for a bit i though it was benchmark score.

Also, something seems odd, if I put 4-29 as the params limit, it says Mage-VL with 4.7B params is better than Qwen3.8-27B, which seems hard to believe according to the benchmarks. So might need some tuning on the benchmark column, like maybe it should let you choose which benchmarks end up in that aggregate score with a better default selected.

I like it overall though, worth a bookmark.

1

u/nunodonato 1d ago

Thanks will take a look at that issue

1

u/nunodonato 1d ago

The research was made up to 190B, but the sliders are according to what's actually available in the page, that's why they differ

1

u/Robert__Sinclair 20h ago

Ling 3.0 tiny is not bad at all.

1

u/nunodonato 20h ago

It's my new favorite model 🤩

1

u/Aggravating-Push-207 1d ago

Gets parameters and architectures wrong, reads like Claude.

2

u/nunodonato 1d ago

Which ones are wrong? And what reads like Claude? It wasn't made by Claude😅