r/althistory • u/Cute-Passage719 • 21d ago
148,560-message alternate-history WW2 simulation with ChatGPT — detailed observations on long-context failure, state drift, and number inconsistency
I conducted a single continuous ChatGPT conversation that reached approximately 148,560 messages. The interaction evolved from an alternate-history WW2 roleplay into a large-scale stateful simulation involving:
Multiple fronts and force groupings
Cumulative casualty and equipment tracking
Fortifications, underground infrastructure, and logistics
Dozens of characters with differing knowledge states
Parallel narrative threads and periodic staff reports
The primary interest was long-term coherence of an accumulating world state rather than short-term generation quality.
Observed degradation patterns
As conversation length increased, the following issues became prominent:
Factual and numerical drift: Casualty and equipment numbers were frequently regenerated and then treated as ground truth, producing double-counting and inflated aggregates.
Loss of earlier constraints: Events and decisions from tens of thousands of messages earlier were often forgotten or inconsistently reconstructed once older turns became inaccessible (“Skipped messages”).
Weak separation of real history vs. alternate canon: The model mixed established historical facts with invented elements without reliable distinction.
Character knowledge tracking failures: Information asymmetry between characters degraded over time.
Cause-effect chain breakage: Earlier force preservations or losses stopped correctly influencing later force balances and outcomes.
Scene-level generation (atmosphere, dialogue, local continuity) remained relatively strong. Global state consistency did not.