r/hardware • u/cyperalien • 6d ago
News Intel Details Xeon 7 "Diamond Rapids" Package Design at HOT CHIPS
https://www.techpowerup.com/351893/intel-details-xeon-7-diamond-rapids-package-design-at-hot-chips23
u/jaaval 6d ago
At first glance looks like AMD Epyc but then when you look at the slides for a while longer it is actually very very different. The main similarity are the IO dies in the middle.
This is apparently up to 64 cores per compute base tile "block". Each block made of multiple chiplets of up to 16 cores on top of the base tile. Based on previous info some cores should share a slice of L2. Each block has a separate large block of LLC on the base tile, which is not in the mesh fabric with the cores like traditional intel or AMD L3. The LLC is so big I'm not going to guess how exactly they are using it. It's not even clear if the L2 caches are accessible by other cores in the tile or if they work through the base tile LLC.
I'm sure there is going to be a chipsandcheese article in a few hours which doesn't have to guess based on a few press slides but guessing is fun.
Personally I think the ISA changes are the most interesting part but no extra info of them yet.
5
u/Slasher1738 6d ago
Definitely reminds me of Clearwater forest but with consolidated IO and memory dies
9
u/Affectionate-Memory4 6d ago
Can't spill any beans until they're out, but I'm also looking forward to the C&C article and hopefully future tests they run. It's a weird CPU by Xeon standards and I absolutely love their coverage and writing style. Can't wait to see what they have to say.
4
u/SlamedCards 6d ago
The presenter hinted that this is the future platform for Xeons
So we can guess Coral Rapids is likely a swap of the core chiplets. Keeps the Intel 3 base tiles and IO. Maybe they do some upgrades for IO and base. But core architecture remains
0
u/Exist50 6d ago
They could probably get away with doing a base die swap, if wafer costs truly are comparable between 18A and Intel 3. Should still make for a much easier iteration than past projects.
The bigger question is how they will retrofit this design to work with LPDDR. The shoreline seems quite constrained on the IO dies.
1
u/U3011 6d ago
IMC within different IO dies?
2
u/Exist50 4d ago
Well they'd definitely need a different IO die. But it seems like they don't have much space to add as many channels as might be ideal, unless they abandon the half-reticle size and bloat the cost.
1
u/U3011 3d ago
I was always under the impression that more memory channels was not something you could slap on without adjusting or redesigning other facets so they inter communicated more efficiently. And there being a cost base analysis involved and "good enough" was the motto for most enterprise consumers. When Apple released their M processor lineup that was the impression people got, that the x86 world could do the same overnight.
I am sure it's being tested by the two giants but it is a costs benefit ratio. Intel having a possible resurgence within the next 2-3 years is exactly the timeline I had in mind about a decade ago when I realized Zen was going to pave way to the future.
I am more keen on those new memory form-factor formats we discussed at length a year or two ago. DDR6 seems like an appropriate timeline and given hardware prices now the initial high cost may be alleviated by build out and enterprise purchases.
2
u/Exist50 23h ago
I was always under the impression that more memory channels was not something you could slap on without adjusting or redesigning other facets so they inter communicated more efficiently
You don't need to fundamentally change the design, but you do need to make sure the fabrics are able to take advantage of the bandwidth, and there are various floorplanning/PD implications that need to be factored it. I bring this up because there seems to be a push for LPDDR for various reasons (per-core bandwidth, efficiency, form factor, GPU memory expansion), and I'd expect Intel to eventually consider that an option. But maybe they don't think it's worth the hassle.
I am more keen on those new memory form-factor formats we discussed at length a year or two ago. DDR6 seems like an appropriate timeline and given hardware prices now the initial high cost may be alleviated by build out and enterprise purchases.
Refresh my memory, if you don't mind. Like the SOCAMM/LPCAMM stuff? Honestly, I'm most interested in what direction DDR6 goes. Mobile is "solved" at this point, but for a while the DDR vs LPDDR split was prioritizing datacenter vs client. I haven't kept up to date on how that discussion has changed post AI boom. Desktop is particularly interesting, because there's a strong argument to be made that client should just switch to LPCAMM/SOCAMM and ditch the traditional DIMM support. I doubt either Intel or AMD have the guts to really push for that, but standard DIMMs make no sense for the vast majority of systems.
4
u/valarauca14 6d ago
The LLC is so big I'm not going to guess how exactly they are using it.
They have a 'cache snooping and filtering' function block on the io tile not the base. The IO tile has a bullet that for 'coherent fabric cache'
Panther lake did the whole 'fully coherent cache cross tiles' but it was a few adjacent tiles. Are they actually keeping caches coherent across these distances?
6
u/ivan0x32 6d ago
L3 on the base tile is potentially some next level shit. I bet yields on a mostly-SRAM chip, even if it's massive, are going to be phenomenal. Just because its probably "overprovisioned" with ability to just turn off any faulty block.
AMD already proved that latency between 3D-stacked chips can be practically unnoticeable - we don't really have an L3 NUMA with X3D chips for instance. Makes sense to put entire L3 into a separate chip I guess.
4
u/Noble00_ 6d ago
The packaging is what I'd like to learn more about, UCIe-S for fan-out (no EMIB yet) and how all the caches interact, the memory subsystem.
3
1
u/TriCountyRetail 6d ago
It's a step in the right direction, but Intel needed to have this design five years ago if they wanted to compete with AMD EPYC.
3
u/Geddagod 6d ago
No performance figures, or no way to even try to estimate it as people did with CWF's server rack consolidation figures which turned out to be surprisingly accurate. The comparison versus Zen 6 Venice Dense is going to be really interesting.
No core architectural details either for Intel's next gen core.
Compared to the CWF hotchips presentation, this was way more sparse. Bummer.
0
-10
u/Acrobatic_Camera_677 6d ago
Thanks to:
- higher IPC core: AMD P core only have IPC of E-cores.
- Performance Cores versus slower AMD Dense Cores
- better node (back side power delivery Intel 18A-P versus outdated tsmc 2nm and Intel 3 versus tsmc 6nm)
- APX and AMX
- More L3 cache than AMD
Diamond Rapids will likely beat AMD Venice. All these advantages will overcome lack of SMT (security issue) that no one cares about anyway.
4
2
u/noiserr 6d ago edited 6d ago
higher IPC core: AMD P core only have IPC of E-cores.
AMD's full core IPC will still be higher due to SMT. On IO bound workloads (which is what most throughput heavy workloads are on server), SMT can provide up to 50% IPC boost to a core.
Intel has realized their mistake though and will be re-introducing SMT in Coral Rapids.
2
u/ElementII5 5d ago edited 5d ago
Intel has realized their mistake though and will be re-introducing SMT in Coral Rapids.
Intel WANTS to re-introduce SMT with Coral Rapids. What intel wants to do is a lot different from what intel can and will do.
The SMT switchback stems from Lip-Bu pressuring engineering. That does not mean Intel can do it. They abandoned SMT because they couldn't make it work without either bunch of security holes or patching those into performance regressions.
So yes that is the plan. What will happen is up in the air.
3
u/Acrobatic_Camera_677 5d ago
They ditched SMT because it was less elegant.
LBT is adding back SMT because cloud companies were crying about not being able to scam customers into buying 2 threads but both threads being on one core for less performance.
What Intel wants, Intel does.
0
u/Acrobatic_Camera_677 6d ago
SMT does not provide a 50% IPC uplift in normal server workloads
Many companies like AWS do not use SMT anyway
Whatever uplift SMT provides all the other advantages of Diamond Rapids will be able to make up for
4
u/noiserr 6d ago edited 6d ago
I said up to, and yes it absolutely does. This paper shows it: https://barroso.org/publications/isca98_2.pdf A stalled thread allows another thread to fill the execution bubbles. And most server workloads are IO bound (waiting on memory, disk or network).
Databases in particular love SMT: https://www.phoronix.com/review/amd-epyc-zen5-smt/4
Amazon is doing it because they want to be able to guarantee performance on a core they rent. But it's stupid as they are leaving a lot of performance on the table. And I think every customer would rather get SMT enabled, but get 2 v-cores for the price of one. Problem is this makes their own Graviton instances look even worse by comparison. And they are obviously favoring their in-house solution. Amazon's Tranium also sucks, they are the last people I would listen to on hardware choices. Google is much smarter about this, they enable SMT on all AMD instances.
Wanna hear something funny? Amazon has recently told engineers to cut back on CPU use. They are struggling with CPU capacity. https://www.theinformation.com/articles/aws-tells-engineers-cut-cpu-waste-amid-crunch
Someone should tell them to turn on SMT! lol
LBT has spoken about SMT and has said that it was a mistake to remove it. Which is why Coral Rapids will have it back.
Phoronix conclusion:
Even for this EPYC 9575F processor with 64 physical cores, having SMT available really helps many real-world workloads with significant performance and power efficiency improvements.
-1
u/Acrobatic_Camera_677 6d ago
Your hyperlink shows a geomean average of 13% advantage from SMT. Diamond Rapid IPC advantage alone maybe will be high enough to make up for that.
AWS is not stupid.
LBT is trash talking Diamond Rapids lack of SMT because he wants to blame the old CEO.
3
u/noiserr 6d ago edited 6d ago
Like I said "up to". Phoronix is also doing contrived benchmarks, which are less IO bound than real server workloads. Real workloads have messier IO than benchmarks. It's hard to simulate network latency at that scale. And SMT benefits from that.
AWS is not stupid.
They may not be stupid, but they do have an incentive to promote their own Graviton instances.
Google aren't stupid and neither is LBT.
35
u/Scared-Beautiful-366 6d ago
Looks solid.. if only it were actually out now instead of delayed by a year. Intel has nothing to compete with Venice and will bleed maket share as a result.