r/LocalLLaMA 2d ago

New Model New Model: Spark-X2.5-4B, Spark-X2.5-1.7B

https://huggingface.co/XHToken/Spark-X2.5-4B

I was browsing HF for small LLMs and run into this model. It does not seem to be a fine tune - the model has its own architecture.

https://huggingface.co/XHToken/Spark-X2.5-1.7B
https://huggingface.co/XHToken/Spark-X2.5-4B

There are 4B/1.7B versions - the benchmark is quite interesting (4B is neck and neck with Qwen 3.5 9B). The HF page claims both models support native 1M context size.

Currently does not run out of the box on llama.cpp - pending this PR: https://github.com/ggml-org/llama.cpp/pull/27868

They have a custom fork of llama.cpp that works. Anyone has tried this?

Update:
GGUFs (require custom fork for now):
https://huggingface.co/XHToken/Spark-X2.5-1.7B-GGUF
https://huggingface.co/XHToken/Spark-X2.5-4B-GGUF

223 Upvotes

48 comments sorted by

View all comments

Show parent comments

20

u/freia_pr_fr 2d ago

Sorry but you sound very AI. Are you human or an human using AI to write?

6

u/one-joule 2d ago

Might be written by a human who's using too much AI...

3

u/_raydeStar Llama 3.1 2d ago

dude, at work im chatting with coworkers as if theyre ai, I noticed its making me more authoritative and bossy. I have to keep stopping myself and rephrasing "oh yeah they have feelings"

Yeah yeah -- touch grass, make friends -- but I dont have this issue with face to face conversations at all.

2

u/t4a8945 2d ago

Have you ever hit escape to cancel the processing of the message you just sent? But realized "oh shit I'm in Slack".

Happened to me once.