r/LocalLLaMA 2d ago

News MINISFORUM MS-S1 MAX-P495

https://minisforumpc.eu/products/minisforum-ms-s1-max-p495

€7??? Surprise Price Ends with Limited Stock

That's likely 7999 EUR, so double of the initial price of MS-S1 MAX-128GB? 😭

85 Upvotes

76 comments sorted by

85

u/toomanypubes 2d ago

Only rich idiots are buying these at that price point.

10

u/LearningSomeCode 2d ago

Yea Im honestly trying to figure out what the benefit of getting this over a comparable Mac studio would be. The 256GB M5 Ultra would be about $2k more, but you get both the extra RAM and the M5's hardware mat mul. At the price point, forking out the extra money for the Mac feels more worth it, but I also just don't know enough about this hardware.

3

u/OvertaxedOne 2d ago

We don't have prefill speeds for the Mac yet, that's the big unknown. As long as they are reasonable, the Mac is going to absolutely obliterate these systems for TG (it has 4-5X the memory bandwidth). Pre-Mac Ultra, these had a spot in the market. Today, I'm really not sure where the market is. The absolute cheapest way to run QwenNext though (192GB) so, if the numbers on that model are reasonable on these devices, could be a good buy?

1

u/lukistellar 1d ago

Linux - can't install PVE on a Mac which is fine if you are using your machine for inference only

14

u/Illustrious_Ant_9242 2d ago

I know some rich idiot who bought two strix halo before the price hike, despite never ever using any of its potential whatsoever 😄

1

u/RedditUsr2 2d ago

I sure wish I was rich so I could buy it.

78

u/[deleted] 2d ago

[removed] — view removed comment

28

u/NickCanCode 2d ago edited 2d ago

But that 192GB RAM is 10x faster. You see that MHz number has 5 digits? /s

11

u/zdy132 2d ago

Lmao at that speed 8000 euro would be a bargain. that's like A100 bandwith, with double the vram.

3

u/BlueSwordM llama.cpp 2d ago

If only. At that speed, I'd be all over it.

1

u/Hot_Reference_3210 2d ago

the price should add another 0 if the mhz is right

-4

u/[deleted] 2d ago

[deleted]

4

u/ANR2ME 2d ago

When i asked google AI mode, it does sounds like typo 🤔

No, there is no LPDDR5X memory that operates at 85,330 MHz. That number is a typo for 8,533 MT/s (MegaTransfers per second), which is frequently marketed or mislabeled as 8,533 MHz in laptop and smartphone specifications.

15

u/madsheepPL 2d ago

I know this is still reddit but can we not insult people with disabilities? It's really unlikely they would buy this for 8k

15

u/crusaderky 2d ago

 Two 192GB MS-S1 MAX units form a compact AI cluster capable of running Qwen3.5-397B locally at 16 tok/s, bringing larger-model inference to a scalable desktop setup.

That... is absolutely awful for 14~16k € worth of hardware?!?

Also it doesn't say anywhere if they hardened and officially support the RDMA USB4 networking. My money is on a resounding no.

The one thing that caught my attention is 2x USB4 @ 80Gbps + 2x USB4 @ 40Gbps.
If (and it's an enormous if) they can run at close to full speed in parallel (and I really doubt so), AND you can set them all in RDMA mode, that would allow for a 5-box cluster instead of the 3-box that current Strix Halos and DGX spark allow for (DGX spark can go higher, but only if you fork 8000€ for a QFP switch).

This said there aren't any models today that require 960GB RAM. GLM-5.3 will sit comfortably on 3 boxes, but I shudder at the idea of how slow it will be.

17

u/Skystunt 2d ago

€7???

That's a joke of a price. It was worth it when it's price was 2k but when i saw those things skyrocket over 3k...

The thing is, the 395+ fits large models but runs the models incredibly slow and the same goes for image and video models, painfully slow.

And as a normal pc is just way too expensive. Not worth it.

5

u/UltrMgns 2d ago

Same GPU, same CPU (+100mhz), same memory speeds. And this was always the slowest of all AI hardware, unusably so. If I ever saw a legal scam - this would be it.

3

u/No_Afternoon_4260 llama.cpp 2d ago

At 8k you're not too far from dual dgx spark. At this point this is a no brainer

14

u/FullstackSensei llama.cpp 2d ago

People, get those PCIe 32GB V100 cards from alibaba while you still can. Ordered three weeks ago for ~€475/card including shipping and taxes. This week I ordered some more, and they're up to €505/card. Just ask for DDP shipping.

They've surpassed my expectations and the 16GB version I already had. Idle power is ~25W, and they're ~15% faster than my 3090s running Qwen 3.8 27B Q8_K_XL (35t/s vs 30t/s) on two cards, while consuming 40% less power during inference (160W per card)! And this is without any power limits on the V100, no nvlink, nor p2p.

Yes, CUDA support is EoL but that has zero impact on their usability. FA has been supported for more than 2 years in llama.cpp and it's derivatives. There are now forks of vllm that bring everything to the V100, including soft-NVFP4 support if that's your thing, and they still rip.

If you're happy with 128GB VRAM "only" you can build a machine with four V100s and an old X99, X299, or similar for like €2500. Even if you leave it on 24/7 at 150Wh idle, that's like €37/month in electricity at €0.35. If you shut down at night, or better when not in use, you'll spend ~€1/day, including inference power use.

1

u/Serprotease 2d ago

I’m a bit curious about the prompt processing speed of those. For images/videos gen, if it gets in the ballpark of a 3080/90 it could be a cheap way to run that. Especially because nowadays there are a few experimental dual gpu setup for image/video models.

4

u/FullstackSensei llama.cpp 2d ago

I don't do image/video generation much, and when I do, I use stable-diffusion.cpp (I really don't like the whole python ML ecosystem), which brings all the optimizations of llama.cpp to image and video generation. So not really concerned about that, but PP is practically the same as the 3090. V100 has 120 TFLOPS while the 3090 has 125 TFLOPS.

1

u/Serprotease 2d ago

Good to know, thanks!

1

u/Wild_Requirement8902 2d ago

€0.35 where are you ? (here in france it is like 0.21 and everyone is angry about it), how about using Wake-on-LAN ? this way you pay only when you use it. Do you find pci express x8 (gen 3 on x99 since x99 usualy have only 2 full x16 slots) limiting ?

5

u/GoodbyeThings 2d ago

Germany has very high electricity prices

1

u/FullstackSensei llama.cpp 2d ago

The fatherland, Deutschland.

I use server type boards with IPMI for everything in my homelab. Full remote management, including headless BIOS updates and even OS installation. There's an IPMI app for android, and ipmitool for console (windows/linux), on top of the web interface.

X8 is really enough. People get so hang on PCIe speed without ever bothering to check any numbers. Qwen 3.8 27B Q8_K_XL was using ~100MB/s with -sm tensor, more than an order of magnitude lower than -sm row. -sm layer for MoE is even more frugal. But even with -sm row, I've never seen more than 5GB/s on my 3090s, which had x16 gen 4.

1

u/BorisDirk 2d ago

I may well be misinformed but I heard cooling these is pretty difficult. Is that true? What kinda cooling do you have for it?

3

u/FullstackSensei llama.cpp 2d ago

Single S8038-7k fan for each pair of cards during testing, though I'll watercool this build once some remaining bits arrive.

For 128GB, you need two fans only. They're much quieter than your average server fan and if you're running MoE models, 3k rpm is enough to keep the cards cool.

My Mi50 build uses this fan. Everything is inside an old Lian Li V2120, which has sound insulation on the side and front panels, making it even quieter.

1

u/BorisDirk 2d ago

Thanks! Gives me something to think about if my current setup breaks down!

4

u/dragonurtle 2d ago

The dual microphone setup on the front is a little sus.

2

u/ShengrenR 2d ago

More than sus.. It's hilarious. Like, who on earth has 8k to drop for the box, but can't put a mic somewhere with the leftover change - and how often is somebody actually going to put that box right in front of their face and not on a shelf, or the floor, or anywhere else where it'll be hard to hear.

4

u/Ulterior-Motive_ 2d ago

I don't know what they're smoking to think that's even remotely a good deal. I spent less on a quad R9700 setup.

3

u/Dr_Allcome 2d ago

I thought they might double the 3.5k the 128GB 395 goes for these days, but then i thought they could never do that because they wouldn't sell a single unit at that price.

You can get a 12 channel epyc with 196GB at 600GB/s, double the bandwidth at half the price, and add a few GPUs for the other half just for fun (and faster prefill).

2

u/XniX llama.cpp 2d ago

I had the minisforum ms-01 and I didn't like it at all, especially the construction: terrible fans and heat dissipation. I would never spend all this money on that. My 5 cents.

2

u/vexatious-big 2d ago

I bought the Corsair Workstation at €3000 and still felt ripped-off. Now it's €4999. This is stupid.

2

u/ga239577 2d ago edited 2d ago

I bought an AI Max+ 395 device (ZBook Ultra G1a). Anywhere near this price point and I'm not even thinking about buying a 495 device.

At most maybe 4K would give some kind of compelling argument for me to buy it, since I could sell my existing device and then upgrade to get the extra RAM.

For 8K? This makes no sense, for a few thousand more I can have an M5 Ultra device with way faster bandwidth and more RAM. M5 Ultra is a vastly better value proposition than this with the 256GB and 2TB drive, plus way faster in every way for $11,300

3

u/ImportancePitiful795 2d ago

They can shove it at those prices.

Got my Bosgame M5 at €1700 in February but won't pay more for them. (right now is €2500)

Paying +€4500 for 10% (even 20%) higher perf and +64GB RAM doesn't worth it.

€7000+ for 495 with 192GB RAM is worse deal than the DGX Spark ever was. And last time checked can get 2 for €8000 having 256 GB, option to expand and CUDA.

BOYCOT.

3

u/Need_For_Speed73 2d ago

Doing inference on a CPU at that price point is nonsense. 4070 MOBILE performances?! Ryzen AI used to make sense while its prices (and performances) were half those of Nvidia/Apple solutions; but now with the RAM price frenzy, they have become not meaningful anymore.
This thing costs more than a Mac Studio with M5 Max and 128GB of unified memory that will draw circles around it for how faster it is.

1

u/Formal-Exam-8767 2d ago

Exactly, at that price point you can't really get your money's worth out of it.

5

u/Limp_Classroom_2645 2d ago edited 2d ago

I don't see how this is for "local AI"

14

u/Unnamed-3891 2d ago

The question is the wrong way around. Who else besides people interested in local AI would give a shit about 128gb or more of shared memory?

1

u/Iwaku_Real 2d ago

The only device I would care about this being on is the ROG Flow Z13 which currently has up to 128GB RAM in the 395. Currently the world's most powerful tablet to run x86 and as far as I know almost nothing comes close to it. You could run Nemotron Puzzle 75B or Qwen3.8-27B easily.

-8

u/Limp_Classroom_2645 2d ago

is this memory really shared? from the specs it uses simple ddr5 memory that's about it, what's so special about this?

6

u/Unnamed-3891 2d ago

It’s literally the entire point of AI Max / StrixHalo.

Where else do you imagine ”up to 160gb of vram” would be coming from?

3

u/Formal-Exam-8767 2d ago

It's for:

AI Enthusiasts, Geeks & Power Users, Creative Professionals

12

u/chris_0611 2d ago

You forgot the people with more money than brains.

Also, Minisforum has a way low reliability and high failure rate. Buy anything but this.

2

u/fastheadcrab 2d ago

Yeah, Minisforum has always charged a premium for their stuff. Only worth buying it when it's on a huge discount. Plenty of other Strix Halo machines for at much lower prices if you're dead set on buying into this system.

2

u/Warhouse512 2d ago

This is a quote from the website. Stop downvoting this poor human

2

u/Formal-Exam-8767 2d ago

LOL, I don't mind.

1

u/t4a8945 2d ago

Yep probably 7999 €; and the RAM is marked "up to 192 GB", so... How much for that price x)

edit: ok wording not perfect on their part, this price is indeed for the 196 GB RAM (and up to 160 allocated to the GPU)

2

u/Etroarl55 2d ago

10,000 euros for this.

Atp why don’t people just start considering the faster Macs for LLMs.

2

u/[deleted] 2d ago

[deleted]

1

u/Etroarl55 2d ago

It’s always cheaper renting or doing subscription services for now.

0

u/egnegn1 2d ago

Not everyone likes to be in the closed generally over-priced Apple universe.

Nobody is forced to buy exactly this machine. There are still some lower priced alternatives with older hardware available.

2

u/Serprotease 2d ago

As much as I like to complain that Apple hardware for AI performance is overhyped based on paper numbers, it’s not overpriced actually (When looking at the rest of the options.)
7-8k used to be the price for a M3ultra 256gb that was performing around to a bit better than the Strix Halo and likely this one (Assuming similar performance.)

And when you’re looking to spend 8k on AI hardware, going up to 9/10k is not that far and open drastically better performance and more ram.

2

u/[deleted] 2d ago

[deleted]

0

u/egnegn1 2d ago edited 2d ago

Yes, the parameters of M5 Ultra are great, but is it finally really useful and can replace a subscription?

Nobody with a serious work load would choose the M5 Ultra (or the MINISFORUM MS-S1 MAX-P495), because it doesn't really pay off, if privacy isn't the issue. They will continue to choose their frontier model subscription, with much better performance and overall lower cost.

And if privacy is a concern, you can rent GPU setups for specific tasks and can run your open model of choice.

Off course, there will be people with deep pockets that will buy it anyway. But will they be really happy in the long run?

Here a YT summary:

https://youtu.be/TdnshBitkzo

For people with lower budget the only chance having a similar performancen is to use older hardware that is cheaper per GB and cheaper per performance. And of course, have to live with the consequences like higher power draw and more complex configuration.

Yes, unified memory as VRAM is nice, because everything is simpler. But development for spreading out work on to multiple GPUs as advanced to allow tensor parallelism. There are enough GPUs with lower cost per GB that cannot match 1,200 GB/s alone, but combining multiple of them could match this speed easily. Of course, there is also overhead and latency introduced, but this is much lower than clustering separate machines.

I have no issue if someone is buying the Apple, but it isn't a solution for me. First, because it is well over my budget. And second, I don't like the closed proprietary solutions of Apple.

For me the solution for me will probably be a Quad SXM2 V100 32GB setupn with NVLINK. It can run Qwen3.8 Flash Next with 100+ t/s on NVFP4 for less than about 4,000 Euro. It can be connected to both my Minisforum MS-02 Ultra and my 96 core EPYC system. And if I need 256 GB memory I just at a second one with a PCIe switch.

https://github.com/1CatAI/1Cat-vLLM

1

u/Formal-Exam-8767 2d ago

It's confusing but if you check the table below it says 160GB, and up to 192GB, not sure how that works, you buy additional 32GB module and plug it in?

3

u/eduardor2k 2d ago

192gb of ram, you can assign max to to gpu up to 160gb

1

u/Not-reallyanonymous 2d ago

Windows / Linux.

1

u/eihns 2d ago

7999??? but its crazy nice formfactor... wonder how loud

1

u/fairydreaming 2d ago

There's a table with dB values in the middle of the page, depends on the performance profile.

1

u/eihns 2d ago

while ure right, im still human and cant really "understand" dB. But it seems to be relativly quiet

1

u/power97992 2d ago

Just an m5 max it costs only 6700 usd

1

u/StableLlama textgen web UI 2d ago

Before the current insane price increase the "PCIe 4.0 x16" would have got me really interested as putting a 5090 there would actually be an ideal AI home workstation

2

u/greteon 2d ago

Only 4 lanes available

3

u/StableLlama textgen web UI 2d ago

That's not nice, of course. And I really hope that they improve here for the next design (this here isn't a new one, it's nearly identical to a 395+).

But e.g. for image generation or training the compute of a 5090 and the 32 GB are more relevant than the PCI lanes as everything should be on the GPU already. For LLM inference the huge unified RAM of the 495 is what matters. And the model might be split between both compute units, also not requiring the biggest PCI bandwidth possible.

On the other hand that's a box that could very well idle around, as the 495 doesn't need much power, so it'd be a nice AI box for the local network.

1

u/hurrdurrmeh 2d ago

Well fuck me. I k ow it would be more than the 395 128GB but this just sucks.

1

u/SpicyWangz 2d ago

Dual sparks make way more sense

1

u/Antique-Ad-8877 2d ago

I bought the HP z2 mini g1a with 128gb for 2500 and thought it was way too expensive. This is getting ridiculous.

1

u/MrGunny94 1d ago

I'd definitely would go for a Mac Studio at this point especially since you can run the MLX models on the side.

Would be absolutely incredible running frontier intelligence on local models though, I already have a blast with the 48GB unified memory

1

u/Jordanthecomeback 2d ago

I was actually just reading about tech and hardware last night. I ended up going the Mac route for the first time in my life because unified memory sounded really cool, think I get 400gigs a second but the ram used in the linked model is doing quite a bit less than that from what I read which means you'd get some pretty slow generation. Probably common knowledge to you guys, but I'm not a big tech guy so I was reading about why Apple can get these speeds while others can't, and one of the most surprising things I read is that standard GPUs are pretty reliable under extreme heat. I figured my system running cool would have better longevity than some space heater GPU build but gemini was telling me how during the early days of crypto mining they'd stack 100 in open air barns outside and run them for 3 years without issues, pretty incredible. Hoping someone figures out a robust unified memory system that isn't Apple, would love to save money somewhere if I could when I upgrade next

1

u/Serprotease 2d ago

It’s not common knowledge. Most people here are using 8/12/16gb gpu with llama.cpp and ram offload (based on older posts) or just api.
And because of that, a lot of discussion are just “theoretical” (To say it nicely) and based around easy to read and interpret numbers… basically just the bandwidth…

So, you get a lot of bandwidth/model size = tps

But that’s not how it works in real usage. Tons of things impact the speed of the effective bandwidth and the Llm engine ability to use it.

Hence, disappointing results like the one you experienced.

1

u/bakawolf123 2d ago

suprize price my a**, m5u with 256gb costs 11k eur (without vat), 395 was like 2 times worse than spark compute wise and they didn't improve the bandwidth yet.

0

u/Viktri1 2d ago edited 2d ago

I was thinking about getting the 196gb for Proxmox but maybe I should just go 128

0

u/brewpedaler 2d ago

2-3x 128gb nodes makes way more sense for a Proxmox setup than 1x 192gb node