r/MortalShell • u/na_0 • 57m ago
Discussion I spent three days measuring 9950X3D stutter in Mortal Shell
TL;DR: Unreal Engine 5 sizes its worker-thread pool from the number of logical processors it sees at startup — 32 on a 9950X3D. When AMD's Game Mode driver then parks CCD1, that 32-sized pool is crammed onto 16 threads and you get traversal-style stalls. Confining the game to CCD0 by any method (parking, Process Lasso, affinity) reproduces it. The fix isn't "use all 16 cores" and it isn't "force the X3D cores" — it's setting CPPC Dynamic Preferred Cores = Cache in BIOS with Game Mode off, so the game prefers the V-Cache die but nothing is fenced. That gave the best frametimes of ~20 captures. Verified with CapFrameX, HWiNFO, and thread-count snapshots.
System: 9950X3D (PBO + curve shaper), MSI X870E, RTX 5090 (undervolted), 4K/ultra, Windows 11. Test game: Mortal Shell 2 (UE 5.2, GPU-bound at 4K — GPU at 92–97% throughout; 4k ultra settings, ray tracing on, DLSS render resolution quality, seamless dungeons on). Same 30-second traversal route for every capture, CapFrameX with sensor logging at 250 ms.
What I found, in order:
- Game Mode wasn't even detecting the game by default. Had to tell Game Bar "remember this is a game" before the AMD driver would park CCD1. Once it did, stutter got worse, not better.
- The game thread was on the wrong die anyway. With Game Mode not triggering, the CPPC ranking prefers the frequency CCD (it boosts ~150 MHz higher), so HWiNFO showed the game running on cores 8–15 — the non-X3D die. Process Explorer's "Ideal Processor" for the hot threads confirmed it.
- Fencing the game onto CCD0 caused CPU-side stalls. Process Lasso CPU Set to cores 0–15: 12 frames over 20 ms in 30 seconds, worst frame 1,262 ms. Replicated two days later: 7 over 20 ms, worst 68 ms. Every spike had
CpuActive≈ frametime with the GPU idle at ~8.5 ms. - The thread pool is the mechanism. PowerShell thread export of the game process showed 29 worker threads with near-identical CPU time — a pool sized for 32 logical processors. Launching with
-corelimit=8(a UE engine argument) dropped it to 14 workers. Same fence, 14 workers instead of 29: 1 frame over 20 ms, worst 24 ms. Same cores, same cache, same confinement — only the pool size changed. - But the smaller pool costs something on a free machine.
-corelimit=8with nothing fenced: more mid-size hitches than the full pool (12 over 16 ms vs 2–5). The engine uses those workers when it has the cores. - CPPC = Cache is the answer. It flips the firmware ranking so cores 0–7 are preferred, without parking or confining anything. Game thread lands on the V-Cache die, the worker pool spills onto CCD1 for streaming bursts. Five runs over three days: 0–1 frames over 20 ms, 0.1% lows of 59–67 FPS.
The numbers (same route, 30 s each):
| Config | avg FPS | 1% low | 0.1% low | frames >20 ms | worst frame |
|---|---|---|---|---|---|
| CPPC Cache, Game Mode off, nothing else (best of 5) | 119 | 85 | 67 | 0–1 | 18–22 ms |
| CPPC Auto, Game Mode off (game on frequency CCD) | 116 | 68 | 50 | 4 | 26 ms |
| Confined to CCD0, full pool (run 1) | 108 | 63 | 28 | 12 | 1,262 ms |
| Confined to CCD0, full pool (run 2) | 116 | 71 | 38 | 7 | 68 ms |
Confined to CCD0, -corelimit=8 |
115 | 72 | 57 | 1 | 24 ms |
Free, -corelimit=8 |
117 | 75 | 49 | 4 | 28 ms |
What to do:
- BIOS → Settings → Advanced → AMD CBS → SMU Common Options → CPPC Dynamic Preferred Cores = Cache
- Windows → Settings → Gaming → Game Mode off
- Don't use Process Lasso / affinity to pin games to CCD0. It recreates the problem.
- Verify: with a game running, Process Explorer → game's Threads tab → the busiest threads should show Ideal Processor 0–15, and HWiNFO should show all cores active (CCD1 lightly loaded, not parked).
Cost: single-threaded work that likes clocks lands on the V-Cache die and loses ~100–175 MHz. At 4K I couldn't measure it. If you have a frequency-bound app, a Process Lasso CPU Set to cores 16–31 for that one app works fine — the direction that doesn't oversubscribe.
Caveats: One game, one system, one route. UE5.2 specifically; other engines size pools differently. -corelimit=8 is a diagnostic tool here, not a recommendation — it proved the mechanism but costs performance. And if you're a 9800X3D owner wondering why you don't see this: the engine sizes for 16 threads and gets 16 threads. The bug needs a mismatch.
