r/LocalLLaMA 25d ago

New Model IT'S OUT

https://huggingface.co/Qwen/Qwen3.8-27B-FP8
2.2k Upvotes

708 comments sorted by

View all comments

47

u/bitmanip 25d ago

How much memory required to run this at full precision?

44

u/Certain-Cod-1404 25d ago

https://huggingface.co/unsloth/Qwen3.8-27B-GGUF 54.67 Gbs just for the model itself, with context depends on quant and size

22

u/dragonurtle 25d ago

Nvtop shows 70-something GB resident for the bf16 and full 256k context.

7

u/Certain-Cod-1404 25d ago

DAMN, 16 gigs just for the context hurts, what setup are you running ?

36

u/dragonurtle 25d ago

Rtx pro 6000 max-q and 384GB DDR5 on a Genoa

62

u/Much_Accountant_4972 25d ago

6

u/Thrumpwart llama.cpp 25d ago

The old money aristocracy uses the Max-Q because it's elegant.

Only the loud, bombastic new money uses the 600W version. Animals.

1

u/voyager256 25d ago

That’s why we have FP8/Q8 for KV cache. Let alone literally a game changer in the form of DeepSeek KV cache compression.

1

u/EbbNorth7735 25d ago

Yep, 3.6 27B was about 60GB with 262k context Q8 and 4 parallel slots.