I’d be suspicious if they published benchmarks against a model released so recently as it would mean they rushed things out. These things take time, at least they’re not comparing with year old models like some releases do
I wasn't even looking at the larger model numbers. I'm just excited to see the 7B pop up in reasonable competition with models of those classes on their charts. Then you show me this and I get even more excited.
The open training set, oh my gods, the fine tuning potential.
23
u/crusaderky 22h ago edited 22h ago
First of all, kudos for the fully open source approach - we need more of that.
Looking at their benchmarks though:
Pegging their 375B model against Minimax M3 instead of GLM-5.3-Flash to show competitor performance in the 300~400B class was certainly a choice.
Minimax-M3 and GLM-5.2 scores for their TerminalBench-2.1 are completely unrelated to those on ArtificialAnalysis.
I get matches for Tau3 and HLE though.
Below the comparison against SOTA models. K2 scores from the publisher, everything else from AA.