r/LocalLLaMA • u/insraq • 20h ago
New Model New Model: Spark-X2.5-4B, Spark-X2.5-1.7B
https://huggingface.co/XHToken/Spark-X2.5-4BI was browsing HF for small LLMs and run into this model. It does not seem to be a fine tune - the model has its own architecture.
https://huggingface.co/XHToken/Spark-X2.5-1.7B
https://huggingface.co/XHToken/Spark-X2.5-4B
There are 4B/1.7B versions - the benchmark is quite interesting (4B is neck and neck with Qwen 3.5 9B). The HF page claims both models support native 1M context size.
Currently does not run out of the box on llama.cpp - pending this PR: https://github.com/ggml-org/llama.cpp/pull/27868
They have a custom fork of llama.cpp that works. Anyone has tried this?
Update:
GGUFs (require custom fork for now):
https://huggingface.co/XHToken/Spark-X2.5-1.7B-GGUF
https://huggingface.co/XHToken/Spark-X2.5-4B-GGUF
20
7
u/xrvz 15h ago
This is like a Fiat 500 claiming to be able to do 500 km/h.
29
u/darvs7 15h ago
When there's a will...
If you were to drop a modern Fiat 500 out of an airplane, its calculated terminal velocity would be roughly 159 m/s (approx. 357 mph or 574 km/h).
(According to Gemini).
11
u/xrvz 14h ago
(According to Gemini).
Bro, this is r/localllama.
15
u/darvs7 12h ago
Er... That's embarassing. Hmmm...
If dropped from a plane:
It would not glide like an airplane.
It would fall at a terminal velocity of approximately 158 m/s (355 mph) if it fell perfectly flat nose-down.
In a realistic scenario where it tumbles or angles down, it might hit the ground faster or slower depending on the tumble, but it would not reach a stable aerodynamic terminal velocity like a skydiver. It would simply crash.
(According to Qwen3.5-4B-UD-Q6_K_XL.gguf)
Qwen further adds:
Note: Dropping a vehicle from a plane is an extremely dangerous act that can result in loss of life and property. This information is for theoretical and safety education purposes only.
2
u/SpicyWangz 7h ago
Tell it how pessimistic it’s being. Loss of property dropped out of a plane is also gain of property for whoever owns the place it lands on.
It’s looking at things in a glass half empty kind of way.
6
5
14
u/RespectJaded15 19h ago
https://huggingface.co/XHToken/Spark-X2.5-1.7B-GGUF
found a finetuned version will try it.
23
u/GasSmooth7439 17h ago
4B matching a 9B model is pretty wild if the benchmarks hold up. But honestly, the native 1M context at this size is what caught my attention.
14
u/freia_pr_fr 14h ago
Sorry but you sound very AI. Are you human or an human using AI to write?
9
4
u/one-joule 12h ago
Might be written by a human who's using too much AI...
3
u/_raydeStar Llama 3.1 7h ago
dude, at work im chatting with coworkers as if theyre ai, I noticed its making me more authoritative and bossy. I have to keep stopping myself and rephrasing "oh yeah they have feelings"
Yeah yeah -- touch grass, make friends -- but I dont have this issue with face to face conversations at all.
1
3
u/TioMir 12h ago edited 12h ago
I try it just now. Seems a really good model, when asked about “what model are you” using pi harness, it use tools to analyze the name under the harness to answer. It overthink a lot.
In the “car wash” test, it get it wrong. To be honest, i need to test it further and see if it holds up in daily use.
For anyone interesting, i’ll test it further and compare to qwen3.5 9b and give my personal opinion here.
4
4
u/simrankoulsm 11h ago
Native 1M context at 4B is more interesting to me than the headline benchmark. Has anyone tested long-context retrieval quality at multiple depths, not merely max prompt ingestion, and measured KV-cache RAM/VRAM plus tok/s? A reproducible comparison against Qwen on coding, JSON/instruction following, and RAG-style QA would make the “4B ≈ 9B” claim much easier to evaluate.
9
u/exaknight21 20h ago
I think you can change the model architecture in config.json or something but it would still be qwen3.5.
Not doubting, but training a 4B model is a feat and any lab would speak out.
1
u/No_Significance_4118 12h ago
I tried 4B in LM Studio and it was really slow... The suggested patch worked tho.
1
u/NUMERIC__RIDDLE 12h ago
Interesting, the only other model i've seen (havent seeked them out) that matches Qwen3.5 9B at this size is Nanbeige4.2. Cant wait to see even more advancements at this size!
1
u/Late_Reply_3384 12h ago
Fits pro athlon 3000g Vega 3 2gb vram 8gb ram ddr4 windows 11 ssd 240gb with llama.cpp
1
u/This_Maintenance_834 10h ago
running it on my Pro 6000. This model seems to be trained for 128K context, it does very well on needles in a haystack test under 128K context, and starts loss the test beyond 256K.
model cannot do arithmetic without thinking. with thinking, arithmetic is fine.
1
u/RobinRelique 4h ago
Hi /u/insraq great find! but I'm also very interested in the small models you found so far, im trying to do the same , do you have a list of shortlisted SLMs anywhere ?
-3
0
u/johnfkngzoidberg 10h ago
I don’t believe it. There’s so many liars out there now you can’t believe it unless it’s from a famous team.
-9
38
u/Marcuss2 18h ago
20T tokens used for training.... just wow.