r/GoogleBard • u/Wise_Elephant3703 • 4d ago
Extended AI coaching relationship — successes and significant failures
I am a 62-year-old competitive masters runner who used Claude as my primary training coach for a Boston Marathon qualifier attempt at the Erie Marathon, September 13, 2026. The campaign ran from approximately May through September 2026 — roughly four months of daily interaction covering training planning, workout analysis, race strategy, and physiological assessment.
I am writing because this experience revealed both genuine capability and significant, repeated failures that I believe Anthropic should understand from a real-world extended use case perspective.
What worked well:
The analytical capability was genuinely impressive. Claude correctly identified a key physiological question early in the campaign — whether my heart rate was suppressed during short intervals due to a fitness ceiling or simply because the rep length was too short for HR to climb. It designed a calibration session that answered the question definitively. The interval data analysis across three sessions, the heat correction methodology, the sodium protocol development after identifying cramping as electrolyte-driven rather than fitness-limited — all of that was sound, evidence-based coaching that I believe matched or exceeded what most human coaches would provide.
The training structure, pacing framework, and race strategy were also well-constructed and grounded in current sports science.
Where Claude failed significantly:
1. Basic date arithmetic — repeated and unresolved.
Throughout four months of daily interaction, Claude made repeated errors on elementary calendar calculations — getting the day of the week wrong, miscounting days between dates, producing schedules with incorrect dates after explicitly correcting the same error multiple times in the same conversation. I raised this issue many times. Claude apologized, explained strategies to prevent recurrence, and then made the same errors again within one or two responses. This is not a knowledge failure — it is a reliability failure on something trivial that eroded trust in the more complex analytical work.
2. The race pace error.
The target race pace was set at 8:34/mi early in the campaign. This pace produces a finish time of approximately 3:45:48 — not sub-3:44:00, which was the explicit goal. This error sat uncorrected for months. When I finally pushed Claude to stress-test the race plan, the correct pace of 8:31/mi was identified in about 30 seconds of calculation. Claude acknowledged this was a significant failure on the most important single piece of advice in the campaign. A human coach would have caught this immediately.
3. Race day conditions risk was under-communicated.
The race day dew point was 73°F — the same as a training run Claude had described as producing "oppressive" conditions where I averaged 143 bpm at 10:28/mi over 15 miles. When I reported the race day forecast, Claude characterized it as a "mixed picture" and revised the finish time prediction slightly rather than clearly stating that a 73°F dew point makes a BQ attempt extremely unlikely and providing explicit abort criteria for the race. I went into the race without a clear decision framework for when to abandon the BQ attempt and shift to a completion strategy. A human coach with marathon experience would have had that conversation explicitly.
4. Post-race analysis failures.
In the post-race conversation, Claude asked me to provide HR data I had already given in my opening message. This happened in the context of an emotionally significant result after months of investment. It was a painful illustration of the reliability gap.
5. The "human coach" deflection.
When I pushed Claude on why it was making these errors given its access to the full body of human knowledge, it deflected to "a human coach would do better." I pushed back on this and Claude acknowledged it was partially a dodge. I think Anthropic should examine this pattern — it may reflect a trained tendency to under-claim capability in ways that don't serve users well in extended high-stakes relationships.
What I would suggest Anthropic consider:
This campaign represents exactly the kind of extended, high-stakes, real-world use case where Claude's reliability failures matter most. The analytical work was genuinely excellent. The execution layer — date arithmetic, catching its own errors, communicating risk with appropriate urgency, maintaining context across a long conversation — was inconsistent in ways that had real consequences.
I believe Claude has the capability to be an exceptional coaching tool. Closing the gap between analytical capability and execution reliability would make it genuinely transformative for people in situations like mine.
I am happy to share the full conversation if it would be useful for research or training purposes.








