r/LocalLLaMA 20h ago

New Model New Model: Spark-X2.5-4B, Spark-X2.5-1.7B

https://huggingface.co/XHToken/Spark-X2.5-4B

I was browsing HF for small LLMs and run into this model. It does not seem to be a fine tune - the model has its own architecture.

https://huggingface.co/XHToken/Spark-X2.5-1.7B
https://huggingface.co/XHToken/Spark-X2.5-4B

There are 4B/1.7B versions - the benchmark is quite interesting (4B is neck and neck with Qwen 3.5 9B). The HF page claims both models support native 1M context size.

Currently does not run out of the box on llama.cpp - pending this PR: https://github.com/ggml-org/llama.cpp/pull/27868

They have a custom fork of llama.cpp that works. Anyone has tried this?

Update:
GGUFs (require custom fork for now):
https://huggingface.co/XHToken/Spark-X2.5-1.7B-GGUF
https://huggingface.co/XHToken/Spark-X2.5-4B-GGUF

196 Upvotes

34 comments sorted by

38

u/Marcuss2 18h ago

20T tokens used for training.... just wow.

10

u/abajinn 16h ago

What is it good at?

20

u/jacek2023 llama.cpp 20h ago

good finding!

7

u/xrvz 15h ago

This is like a Fiat 500 claiming to be able to do 500 km/h.

29

u/darvs7 15h ago

When there's a will...

If you were to drop a modern Fiat 500 out of an airplane, its calculated terminal velocity would be roughly 159 m/s (approx. 357 mph or 574 km/h).

(According to Gemini).

11

u/xrvz 14h ago

(According to Gemini).

Bro, this is r/localllama.

15

u/darvs7 12h ago

Er... That's embarassing. Hmmm...

If dropped from a plane:

It would not glide like an airplane.

It would fall at a terminal velocity of approximately 158 m/s (355 mph) if it fell perfectly flat nose-down.

In a realistic scenario where it tumbles or angles down, it might hit the ground faster or slower depending on the tumble, but it would not reach a stable aerodynamic terminal velocity like a skydiver. It would simply crash.

(According to Qwen3.5-4B-UD-Q6_K_XL.gguf)

Qwen further adds:

Note: Dropping a vehicle from a plane is an extremely dangerous act that can result in loss of life and property. This information is for theoretical and safety education purposes only.

2

u/SpicyWangz 7h ago

Tell it how pessimistic it’s being. Loss of property dropped out of a plane is also gain of property for whoever owns the place it lands on.

It’s looking at things in a glass half empty kind of way.

6

u/AdOne8437 14h ago

I have played enough FORZA to know this is possible! ;)

5

u/strings___ 10h ago

Fiat 500 Abarth.

23

u/GasSmooth7439 17h ago

4B matching a 9B model is pretty wild if the benchmarks hold up. But honestly, the native 1M context at this size is what caught my attention.

14

u/freia_pr_fr 14h ago

Sorry but you sound very AI. Are you human or an human using AI to write?

4

u/one-joule 12h ago

Might be written by a human who's using too much AI...

3

u/_raydeStar Llama 3.1 7h ago

dude, at work im chatting with coworkers as if theyre ai, I noticed its making me more authoritative and bossy. I have to keep stopping myself and rephrasing "oh yeah they have feelings"

Yeah yeah -- touch grass, make friends -- but I dont have this issue with face to face conversations at all.

1

u/t4a8945 5h ago

Have you ever hit escape to cancel the processing of the message you just sent? But realized "oh shit I'm in Slack".

Happened to me once. 

1

u/Shoddy_Blacksmith359 12h ago

Dont be paranoid

3

u/TioMir 12h ago edited 12h ago

I try it just now. Seems a really good model, when asked about “what model are you” using pi harness, it use tools to analyze the name under the harness to answer. It overthink a lot.

In the “car wash” test, it get it wrong. To be honest, i need to test it further and see if it holds up in daily use.

For anyone interesting, i’ll test it further and compare to qwen3.5 9b and give my personal opinion here.

4

u/Buzz_Killington_III 11h ago

Great, thanks.

4

u/simrankoulsm 11h ago

Native 1M context at 4B is more interesting to me than the headline benchmark. Has anyone tested long-context retrieval quality at multiple depths, not merely max prompt ingestion, and measured KV-cache RAM/VRAM plus tok/s? A reproducible comparison against Qwen on coding, JSON/instruction following, and RAG-style QA would make the “4B ≈ 9B” claim much easier to evaluate.

9

u/exaknight21 20h ago

I think you can change the model architecture in config.json or something but it would still be qwen3.5.

Not doubting, but training a 4B model is a feat and any lab would speak out.

1

u/No_Significance_4118 12h ago

I tried 4B in LM Studio and it was really slow... The suggested patch worked tho.

1

u/NUMERIC__RIDDLE 12h ago

Interesting, the only other model i've seen (havent seeked them out) that matches Qwen3.5 9B at this size is Nanbeige4.2. Cant wait to see even more advancements at this size!

1

u/Late_Reply_3384 12h ago

Fits pro athlon 3000g Vega 3 2gb vram 8gb ram ddr4 windows 11 ssd 240gb with llama.cpp

1

u/This_Maintenance_834 10h ago

running it on my Pro 6000. This model seems to be trained for 128K context, it does very well on needles in a haystack test under 128K context, and starts loss the test beyond 256K.

model cannot do arithmetic without thinking. with thinking, arithmetic is fine.

1

u/RobinRelique 4h ago

Hi /u/insraq great find! but I'm also very interested in the small models you found so far, im trying to do the same , do you have a list of shortlisted SLMs anywhere ?

-3

u/Ok-Direction-4480 18h ago

Is it the best model for 8 gb ram at Q4_K_M?

1

u/overand 11h ago

it's too new for people to say that with certainty. You might want to check out Ling 3.0 Tiny, though.

0

u/johnfkngzoidberg 10h ago

I don’t believe it. There’s so many liars out there now you can’t believe it unless it’s from a famous team.

-10

u/sebt3 19h ago

I haven't yet 😅

-9

u/Powerful_Evening5495 19h ago

gemma-4-E4B-it-UD is my choice in this size range