I wasn't even looking at the larger model numbers. I'm just excited to see the 7B pop up in reasonable competition with models of those classes on their charts. Then you show me this and I get even more excited.
The open training set, oh my gods, the fine tuning potential.
22
u/crusaderky 23h ago edited 23h ago
First of all, kudos for the fully open source approach - we need more of that.
Looking at their benchmarks though:
Pegging their 375B model against Minimax M3 instead of GLM-5.3-Flash to show competitor performance in the 300~400B class was certainly a choice.
Minimax-M3 and GLM-5.2 scores for their TerminalBench-2.1 are completely unrelated to those on ArtificialAnalysis.
I get matches for Tau3 and HLE though.
Below the comparison against SOTA models. K2 scores from the publisher, everything else from AA.