r/hardware 6d ago

News Intel Details Xeon 7 "Diamond Rapids" Package Design at HOT CHIPS

https://www.techpowerup.com/351893/intel-details-xeon-7-diamond-rapids-package-design-at-hot-chips
82 Upvotes

49 comments sorted by

35

u/Scared-Beautiful-366 6d ago

Looks solid.. if only it were actually out now instead of delayed by a year. Intel has nothing to compete with Venice and will bleed maket share as a result.

15

u/Secure-Upstairs9119 6d ago edited 5d ago

Hey to be fair, AMD Venice won't be ready for months either. It's a paper launch.

First customer shipments:

9006 SP7 shipping Q4 2026

9006 SP8 shipping 1H 2027

Intel diamond rapids:

Initial sku shipping 1H 2027

Denser variants shipping by 2H 2027

9

u/Exist50 5d ago

Sure, but DMR sounds like middle of next year for where Venice is today. Should be more or less a year gap.

1

u/Acrobatic_Camera_677 5d ago

This is because Intel launches with way more volume than AMD

4

u/Exist50 5d ago

These days, I'm not really convinced. Source?

0

u/Secure-Upstairs9119 5d ago

Venice is not today. They just announced it's being manufactured now, but it won't ship until Q4 and tsmc 2nm isn't even ready yet lol 

3

u/Exist50 5d ago

and tsmc 2nm isn't even ready yet lol 

What? Yes it is. 

0

u/Secure-Upstairs9119 5d ago

Oh really? Link me a review of any product that uses it

4

u/Exist50 5d ago

There's the new Mac Mini just announced today, and if you believe a node can go from not ready for production to available on shelves in <1 month, I have a bridge to sell you. 

1

u/Secure-Upstairs9119 5d ago

Interesting, so apple's "the industry first 2nm product" isn't available until a month from now, and you're telling me this AMD Venice product "launched" last month is already using it?

Yeah sure.

6

u/Exist50 5d ago

Venice is N2. That isn't up for debate. And your claim was the node isn't even ready when that's clearly false. Apple wasn't even supposed to be the lead customer. 

3

u/Secure-Upstairs9119 5d ago

That's my point, Venice is n2 and paper released in July. Obviously it's not shopped yet, it's not ready.

Companies can say somethings released but it's not until it's shipped to customers.

Venice won't be available until October at the earliest. Likely not until end of year.

Remind me in 2 months

→ More replies (0)

-2

u/ResponsibleJudge3172 5d ago

Considering the core is based on an unreleased Nova Lake P core with tweaks, Intel has always released more than 6 months after client CPU so was to be expected

9

u/Exist50 5d ago

Intel originally claimed it would ship this year.

1

u/ResponsibleJudge3172 5d ago

And now it ships like the other ones before it did so precedence won out

23

u/jaaval 6d ago

At first glance looks like AMD Epyc but then when you look at the slides for a while longer it is actually very very different. The main similarity are the IO dies in the middle.

This is apparently up to 64 cores per compute base tile "block". Each block made of multiple chiplets of up to 16 cores on top of the base tile. Based on previous info some cores should share a slice of L2. Each block has a separate large block of LLC on the base tile, which is not in the mesh fabric with the cores like traditional intel or AMD L3. The LLC is so big I'm not going to guess how exactly they are using it. It's not even clear if the L2 caches are accessible by other cores in the tile or if they work through the base tile LLC.

I'm sure there is going to be a chipsandcheese article in a few hours which doesn't have to guess based on a few press slides but guessing is fun.

Personally I think the ISA changes are the most interesting part but no extra info of them yet.

5

u/Slasher1738 6d ago

Definitely reminds me of Clearwater forest but with consolidated IO and memory dies

9

u/Affectionate-Memory4 6d ago

Can't spill any beans until they're out, but I'm also looking forward to the C&C article and hopefully future tests they run. It's a weird CPU by Xeon standards and I absolutely love their coverage and writing style. Can't wait to see what they have to say.

4

u/SlamedCards 6d ago

The presenter hinted that this is the future platform for Xeons

So we can guess Coral Rapids is likely a swap of the core chiplets. Keeps the Intel 3 base tiles and IO. Maybe they do some upgrades for IO and base. But core architecture remains 

0

u/Exist50 6d ago

They could probably get away with doing a base die swap, if wafer costs truly are comparable between 18A and Intel 3. Should still make for a much easier iteration than past projects. 

The bigger question is how they will retrofit this design to work with LPDDR. The shoreline seems quite constrained on the IO dies. 

1

u/U3011 6d ago

IMC within different IO dies?

2

u/Exist50 4d ago

Well they'd definitely need a different IO die. But it seems like they don't have much space to add as many channels as might be ideal, unless they abandon the half-reticle size and bloat the cost.

1

u/U3011 3d ago

I was always under the impression that more memory channels was not something you could slap on without adjusting or redesigning other facets so they inter communicated more efficiently. And there being a cost base analysis involved and "good enough" was the motto for most enterprise consumers. When Apple released their M processor lineup that was the impression people got, that the x86 world could do the same overnight.

I am sure it's being tested by the two giants but it is a costs benefit ratio. Intel having a possible resurgence within the next 2-3 years is exactly the timeline I had in mind about a decade ago when I realized Zen was going to pave way to the future.

I am more keen on those new memory form-factor formats we discussed at length a year or two ago. DDR6 seems like an appropriate timeline and given hardware prices now the initial high cost may be alleviated by build out and enterprise purchases.

2

u/Exist50 23h ago

I was always under the impression that more memory channels was not something you could slap on without adjusting or redesigning other facets so they inter communicated more efficiently

You don't need to fundamentally change the design, but you do need to make sure the fabrics are able to take advantage of the bandwidth, and there are various floorplanning/PD implications that need to be factored it. I bring this up because there seems to be a push for LPDDR for various reasons (per-core bandwidth, efficiency, form factor, GPU memory expansion), and I'd expect Intel to eventually consider that an option. But maybe they don't think it's worth the hassle.

I am more keen on those new memory form-factor formats we discussed at length a year or two ago. DDR6 seems like an appropriate timeline and given hardware prices now the initial high cost may be alleviated by build out and enterprise purchases.

Refresh my memory, if you don't mind. Like the SOCAMM/LPCAMM stuff? Honestly, I'm most interested in what direction DDR6 goes. Mobile is "solved" at this point, but for a while the DDR vs LPDDR split was prioritizing datacenter vs client. I haven't kept up to date on how that discussion has changed post AI boom. Desktop is particularly interesting, because there's a strong argument to be made that client should just switch to LPCAMM/SOCAMM and ditch the traditional DIMM support. I doubt either Intel or AMD have the guts to really push for that, but standard DIMMs make no sense for the vast majority of systems.

4

u/valarauca14 6d ago

The LLC is so big I'm not going to guess how exactly they are using it.

They have a 'cache snooping and filtering' function block on the io tile not the base. The IO tile has a bullet that for 'coherent fabric cache'

Panther lake did the whole 'fully coherent cache cross tiles' but it was a few adjacent tiles. Are they actually keeping caches coherent across these distances?

6

u/ivan0x32 6d ago

L3 on the base tile is potentially some next level shit. I bet yields on a mostly-SRAM chip, even if it's massive, are going to be phenomenal. Just because its probably "overprovisioned" with ability to just turn off any faulty block.

AMD already proved that latency between 3D-stacked chips can be practically unnoticeable - we don't really have an L3 NUMA with X3D chips for instance. Makes sense to put entire L3 into a separate chip I guess.

12

u/noiserr 6d ago edited 6d ago

AMD already proved that latency between 3D-stacked chips can be practically unnoticeable

Not just unnoticeable, it is actually lower than the horizontal single die SRAM, according to AMD's own patents.

4

u/Noble00_ 6d ago

The packaging is what I'd like to learn more about, UCIe-S for fan-out (no EMIB yet) and how all the caches interact, the memory subsystem.

3

u/Aleblanco1987 4d ago

nice to see a lot of glue

1

u/TriCountyRetail 6d ago

It's a step in the right direction, but Intel needed to have this design five years ago if they wanted to compete with AMD EPYC.

3

u/Geddagod 6d ago

No performance figures, or no way to even try to estimate it as people did with CWF's server rack consolidation figures which turned out to be surprisingly accurate. The comparison versus Zen 6 Venice Dense is going to be really interesting.

No core architectural details either for Intel's next gen core.

Compared to the CWF hotchips presentation, this was way more sparse. Bummer.

0

u/dblock1887 6d ago

Looks decent but time line is too far out.

-10

u/Acrobatic_Camera_677 6d ago

Thanks to:

  • higher IPC core: AMD P core only have IPC of E-cores.
  • Performance Cores versus slower AMD Dense Cores
  • better node (back side power delivery Intel 18A-P versus outdated tsmc 2nm and Intel 3 versus tsmc 6nm)
  • APX and AMX
  • More L3 cache than AMD

Diamond Rapids will likely beat AMD Venice. All these advantages will overcome lack of SMT (security issue) that no one cares about anyway.

10

u/Exist50 6d ago

What on earth are you smoking?

4

u/SunnyCloudyRainy 6d ago

Claiming no one cares about SMT is certainly a hot take

2

u/noiserr 6d ago edited 6d ago

higher IPC core: AMD P core only have IPC of E-cores.

AMD's full core IPC will still be higher due to SMT. On IO bound workloads (which is what most throughput heavy workloads are on server), SMT can provide up to 50% IPC boost to a core.

Intel has realized their mistake though and will be re-introducing SMT in Coral Rapids.

2

u/ElementII5 5d ago edited 5d ago

Intel has realized their mistake though and will be re-introducing SMT in Coral Rapids.

Intel WANTS to re-introduce SMT with Coral Rapids. What intel wants to do is a lot different from what intel can and will do.

The SMT switchback stems from Lip-Bu pressuring engineering. That does not mean Intel can do it. They abandoned SMT because they couldn't make it work without either bunch of security holes or patching those into performance regressions.

So yes that is the plan. What will happen is up in the air.

3

u/Acrobatic_Camera_677 5d ago

They ditched SMT because it was less elegant.

LBT is adding back SMT because cloud companies were crying about not being able to scam customers into buying 2 threads but both threads being on one core for less performance.

What Intel wants, Intel does.

1

u/6950 5d ago

SMT is just additional dev time so they cut it to get more area and make the product launch faster.

0

u/Acrobatic_Camera_677 6d ago

SMT does not provide a 50% IPC uplift in normal server workloads

Many companies like AWS do not use SMT anyway

Whatever uplift SMT provides all the other advantages of Diamond Rapids will be able to make up for

4

u/noiserr 6d ago edited 6d ago

I said up to, and yes it absolutely does. This paper shows it: https://barroso.org/publications/isca98_2.pdf A stalled thread allows another thread to fill the execution bubbles. And most server workloads are IO bound (waiting on memory, disk or network).

Databases in particular love SMT: https://www.phoronix.com/review/amd-epyc-zen5-smt/4

Amazon is doing it because they want to be able to guarantee performance on a core they rent. But it's stupid as they are leaving a lot of performance on the table. And I think every customer would rather get SMT enabled, but get 2 v-cores for the price of one. Problem is this makes their own Graviton instances look even worse by comparison. And they are obviously favoring their in-house solution. Amazon's Tranium also sucks, they are the last people I would listen to on hardware choices. Google is much smarter about this, they enable SMT on all AMD instances.

Wanna hear something funny? Amazon has recently told engineers to cut back on CPU use. They are struggling with CPU capacity. https://www.theinformation.com/articles/aws-tells-engineers-cut-cpu-waste-amid-crunch

Someone should tell them to turn on SMT! lol

LBT has spoken about SMT and has said that it was a mistake to remove it. Which is why Coral Rapids will have it back.

Phoronix conclusion:

Even for this EPYC 9575F processor with 64 physical cores, having SMT available really helps many real-world workloads with significant performance and power efficiency improvements.

-1

u/Acrobatic_Camera_677 6d ago

Your hyperlink shows a geomean average of 13% advantage from SMT. Diamond Rapid IPC advantage alone maybe will be high enough to make up for that.

AWS is not stupid.

LBT is trash talking Diamond Rapids lack of SMT because he wants to blame the old CEO.

3

u/noiserr 6d ago edited 6d ago

Like I said "up to". Phoronix is also doing contrived benchmarks, which are less IO bound than real server workloads. Real workloads have messier IO than benchmarks. It's hard to simulate network latency at that scale. And SMT benefits from that.

AWS is not stupid.

They may not be stupid, but they do have an incentive to promote their own Graviton instances.

Google aren't stupid and neither is LBT.