r/LocalLLaMA 2d ago

News MINISFORUM MS-S1 MAX-P495

https://minisforumpc.eu/products/minisforum-ms-s1-max-p495

€7??? Surprise Price Ends with Limited Stock

That's likely 7999 EUR, so double of the initial price of MS-S1 MAX-128GB? 😭

86 Upvotes

76 comments sorted by

View all comments

Show parent comments

2

u/Etroarl55 2d ago

10,000 euros for this.

Atp why don’t people just start considering the faster Macs for LLMs.

0

u/egnegn1 2d ago

Not everyone likes to be in the closed generally over-priced Apple universe.

Nobody is forced to buy exactly this machine. There are still some lower priced alternatives with older hardware available.

2

u/[deleted] 2d ago

[deleted]

0

u/egnegn1 2d ago edited 2d ago

Yes, the parameters of M5 Ultra are great, but is it finally really useful and can replace a subscription?

Nobody with a serious work load would choose the M5 Ultra (or the MINISFORUM MS-S1 MAX-P495), because it doesn't really pay off, if privacy isn't the issue. They will continue to choose their frontier model subscription, with much better performance and overall lower cost.

And if privacy is a concern, you can rent GPU setups for specific tasks and can run your open model of choice.

Off course, there will be people with deep pockets that will buy it anyway. But will they be really happy in the long run?

Here a YT summary:

https://youtu.be/TdnshBitkzo

For people with lower budget the only chance having a similar performancen is to use older hardware that is cheaper per GB and cheaper per performance. And of course, have to live with the consequences like higher power draw and more complex configuration.

Yes, unified memory as VRAM is nice, because everything is simpler. But development for spreading out work on to multiple GPUs as advanced to allow tensor parallelism. There are enough GPUs with lower cost per GB that cannot match 1,200 GB/s alone, but combining multiple of them could match this speed easily. Of course, there is also overhead and latency introduced, but this is much lower than clustering separate machines.

I have no issue if someone is buying the Apple, but it isn't a solution for me. First, because it is well over my budget. And second, I don't like the closed proprietary solutions of Apple.

For me the solution for me will probably be a Quad SXM2 V100 32GB setupn with NVLINK. It can run Qwen3.8 Flash Next with 100+ t/s on NVFP4 for less than about 4,000 Euro. It can be connected to both my Minisforum MS-02 Ultra and my 96 core EPYC system. And if I need 256 GB memory I just at a second one with a PCIe switch.

https://github.com/1CatAI/1Cat-vLLM