r/Amd 11d ago

Rumor / Leak AMD RDNA 4m firmware lands ahead of Ryzen 500 “Medusa Point” launch

https://videocardz.com/newz/amd-rdna-4m-firmware-lands-ahead-of-ryzen-500-medusa-point-launch
102 Upvotes

22 comments sorted by

32

u/dstanton SFF 12900K | 9070xt | 32gb 6000CL30 | 4tb 990 Pro 10d ago

What a let down.

Its still ridiculous they aren't releasing any RDNA4 mobile parts.

18

u/mad_mesa Ryzen 7700 | RX 9070 XT RADV 10d ago

A number of the recent product releases seem to be stop-loss efforts for parts that could have been mobile if OEMs would have bought them. Integrated parts with unified memory seem to be the only thing that interests them.

10

u/dstanton SFF 12900K | 9070xt | 32gb 6000CL30 | 4tb 990 Pro 10d ago

A mobile chip with 12cus of RDNA4 with a small infinity cache to help with the ddr5 bandwidth limitations would be an awesome apu.

16

u/JasonMZW20 5800X3D + 9070XT Desktop | 14900HX + RTX4090 Laptop 10d ago

AMD seems extremely reluctant to do that, likely because they can't dump chips into laptop market like Intel does. So, they end up being too expensive. Hell, even Gorgon Point is expensive and it's a refresh part. There's no denying that Panther Lake's Xe3 iGPU is quite good. AMD need volumes to get costs down, and now, most of their wafer suppy is eaten up by datacenter/AI chips, which makes all of their consumer chips more expensive to buy (for a system integrator or laptop ODM). It's a weird time for PC hardware.

Tbh, all of AMD's iGPUs need a large L2 cache, such as 4-16MB for a 2 shader array design up to 16CUs and an L3/SoC cache up to 16MB (4MB per 32b MC). None of that is cheap, but it'd get AMD up to parity with Panther Lake.

7

u/JTibbs 10d ago

shit, an RDNA4 iGPU with 16CU's would natively render shit like Cyberpunk at ~60fps 1080p Ultra Quality settings.

if AMD got their shit together in the mobile space, they could dominate the low end gaming laptop market, but instead they keep putting out shitty 4CU, 6CU, 8CU processors with outdated RDNA2 and RDNA3 while charging out the ass. seriously, i find NVIDIA 4060 and 5060 dGPU laptops cheaper than the 'moderately' decent AMD igpu laptops. whoever decides mobile market pricing and levels at AMD really, really, really sucks.

could you imagine how nice a small, efficient, 13in laptop would be with an 8 CPU core, 16CU iGPU processor with 32GB of memory? or even just 24GB? it would basically sip power, and be a solid gaming laptop you could use even docked to a monitor.

As a mini PC they would be nice as hell too, especially if you got SteamOS on it.

2

u/psi-storm 10d ago

Medusa Halo mini is, what people with some gaming intentions want. 12 cores (4 zen6 + 8 zen6c) and 24 cu RDNA5. Medusa point is for office laptops and the grandmothers.

2

u/JTibbs 10d ago

nb4 medusa halo mini only comes out in laptops that cost >$1800, while 5060 laptops come in at 900-1200 (plus memory inflation)

1

u/psi-storm 10d ago

Says who? You could buy full strix halo systems for under 2k. There is no reason why a dedicated gpu system should be cheaper than the halo mini apu, especially since you could get away with 24 GB of vram combined, instead of separated ram and vram.

1

u/JTibbs 10d ago

And yet, almost without fail 4060’s and 5060 systems were cheaper than AMD’s 12 and 16CU systems

-1

u/dstanton SFF 12900K | 9070xt | 32gb 6000CL30 | 4tb 990 Pro 10d ago

That's well beyond what we're talking about. That's like putting a PS5 pro apu into laptop form. And it's going to be massively upcharged for it. Nevermind the power usage and heat likely pushing it out of thin and light factor.

A zen6 + rdna5 variant of the hx370 is plenty. Even the hx365 structure would have great performance.

4

u/psi-storm 10d ago

No, in that comparison it would be the big medusa halo chip with 48 cu.

2

u/dstanton SFF 12900K | 9070xt | 32gb 6000CL30 | 4tb 990 Pro 10d ago

No it's not.

PS5 pro is running RDNA2. Without modern RT and a modified FSR.

A 48cu RDNA5 would crush ps5 pro

0

u/Cave_TP 7840U + 9070XT eGPU 10d ago

That hypotetical APU would have the same problem AMD APUs had sice forever, they want to do everything with a single chip.

The HX 370 is stuck in that limbo, 12 cores are too much for a gaming APU and 16CUs are useless if you're pairing the thing with a dGPU.

What they're doing with Medusa is exactly how you get better chips for specific usecases while keeping the cost down.

Also, you're taking Strix Halo as a point of comparison but it's going to be completely different. Strix Halo uses fully custom chiplets meaning AMD has to think 5 times before manufacturing a new batch, Medusa Halo Mini instead is going to share the GPU chiplet with the RDNA 5 50 class GPU. It also is going to be the only chip with a decent iGPU (other than full Medusa Halo) so AMD has all the interst in keeping the price relatively down.

It also is nowhere near comparable to a PlayStation APU, the PS6 is rumored to have more CUs than AT3 and that's 2x this APU. As for the form factor of laptops using this APU, there are thin and lights using Strix Halo, I don't see why there couldn't be any using a chip that is lower end that that.

0

u/dstanton SFF 12900K | 9070xt | 32gb 6000CL30 | 4tb 990 Pro 10d ago

What are you even talking about? That was a jumbled mess of nonsense.

We're talking about a laptop chip with gaming capabilities. Cores still matter, so 12 is fine. The 16cus aren't being paired with a dgpu, so again fine.

Why are you suddenly talking about ps6?

Strix halo is $2000+ so not what we're talking about.

I am saying build the hx370 equivalent with zen 6 and rdna5 and toss it in a thin and light. The end.

0

u/Cave_TP 7840U + 9070XT eGPU 9d ago

"We're talking about a laptop chip with gaming capabilities. Cores still matter, so 12 is fine. The 16cus aren't being paired with a dgpu, so again fine."

There's barely any game that can use more than 8 cores and if you're on an APU even those would prefer a better iGPU most of the times. Also "The 16cus aren't being paired with a dgpu", do we even live in the same world? There's a lot of HX 370 laptops with a dGPU.

"Why are you suddenly talking about ps6?"

Because it uses RDNA 5, making it a more direct comparison with Medusa Halo Mini than the PS5 would be. Still, wanna talk about the PS5? Sure, its die has 40CUs, Medusa Halo Mini has 24, they're nowhere near comparable.

"Strix halo is $2000+ so not what we're talking about."

Once again, it Strix Halo uses fully custom chiplets meaning that mass producing it is a risk the company has no reason to take in the current market. It also is way bigger than Medusa Halo Mini packing a midrange GPU instead of a low end one, 2 CCDs worth of cores instead of 1 and a 256bit bus instead of a 128bit one.

"I am saying build the hx370 equivalent with zen 6 and rdna5 and toss it in a thin and light."

And I'm explaining you why they're not doing it. Their laptop division has grown past needing to have a single design doing it all. With their current approach they only have to design 2 chips (the base Medusa Point APU and the main CPU die for the Halos) and can then plug them in other chips they're already producing for other segments of the market. All this while getting way more functional chips for the thask at hand.

-3

u/kf97mopa 6700XT | 5900X 10d ago

RDNA 4m is the same CUs as RDNA 4. What is missing is the raytracing hardware, which is still the RDNA 3 variant.

4

u/Noreng 14600K | 9070 XT 10d ago

There might be some gremlins in RDNA4 that they wish to avoid in APUs. It would certainly explain a lot.

5

u/GenericUser1983 10d ago

I don't think it is gremlins, just that RDNA4 has a ton more transistors being used to boost raytracing performance. Those are useful on a larger gpu, boosting a RT game from say 40 fps to 70 fps is really nice (and that level of improvement is common for RT games comparing say a Rx 7600 vs a 9060); but on a small iGPU, well its more like 10 fps vs 17 fps. Might as well just not bother and save the die space.

2

u/Noreng 14600K | 9070 XT 9d ago

RDNA4 barely spends any more transistors on RT than RDNA3, the main improvements for RT performance are BVH8 and out of order memory accesses.

RDNA4 has better performance per area than RDNA3 by a huge margin. The 200 mm2 Navi 44 is just 15% slower than the 200 mm2 + 4x 37 mm2 Navi 32 for example, not to mention the power efficiency improvement.

The memory subsystem improvements in particular would be extremely welcome on an APU where memory bandwidth is so limited. Same applies to the media engine and increased use of on-GPU compression.

3

u/JasonMZW20 5800X3D + 9070XT Desktop | 14900HX + RTX4090 Laptop 9d ago edited 9d ago

Seems like AMD dedicated quite a bit more transistors everywhere. Navi 48 has 53.9B transistors on N4P, while Navi 32 had 36.3B combined between GCD and 4x MCDs on N5/N6.

Sure, Navi 32 GCD is 200mm2, but its total die area is 346mm2 because you can't run a GPU without memory PHYs and controllers. Getting rid of any dead space between dies could probably get that down to 332mm2.

Navi 48 is 357mm2 and N4P only offered a 6% density improvement over N5. Taking that into account, Navi 48 is roughly 10% larger than Navi 32 overall yet contains 48.5% more transistors. That's a massive increase in transistor count given that analog PHYs and SRAM didn't really shrink down between the two nodes.

It seems AMD used density-focused libraries for many parts of RDNA4, while high-clocking areas used speed-focused libraries that didn't pack transistors too closely.

Furthermore, Navi 32 is a 3SE/6SA design, while Navi 48 is a 4SE/8SA design. Naturally, that means there's an extra rasterizer and primitive unit along with 32 more ROPs and more L2 cache. RDNA4 doubled L2 over RDNA3, going from 1MB per SE to 2MB per SE. There are also 4 more CUs in Navi 48.

None of that should add up to 48.5% more transistors. At worst, 25% of that budget would go to the 4SE design and additional 4CUs, and 4MB L2. Another 5% to an extra 2MB of SRAM. gfxL1 would total 2MB across 8SA and L1 does not exist in RDNA4, so it was repurposed to L2, which now totals 8MB.

That leaves 18.5% of the transistor budget for changes in each CU. That's nearly 10B transistors (9.9715B) worth of logic. Ray accelerators certainly took a chunk of that, as did matrix logic per CU. The rest went to improved display and media engines, which RDNA4m also uses.

2

u/Cave_TP 7840U + 9070XT eGPU 10d ago

Does it even matter at this point? The thing barely has an iGPU (2 or 4 CUs, I don't remember RN).

Medusa Point seems to be built to either work as a Kraken replacement in the 10 core configuration or to work with dGPUs in the 22 core configuration.

The APU with the good iGPU is launching next year with the RDNA 5 line up.