Case study · Gymswell · August–September 2026
Three bugs from the beta, and what each one taught me.
Gymswell is a fitness app I have built since October 2025: React Native, an AI coach on OpenAI, native Swift for the Lock Screen. When it went to testers in August, three problems came back that were not what they looked like. This is how I found each one, with the numbers I measured on the way, taken straight from the commit log.
1 · The coach
The AI was answering. We hung up at 30 seconds and blamed the wifi.
What testers saw
What I measured
The cause
What I tried
What shipped
| Model | Time | Tokens | Speed |
|---|---|---|---|
| gpt-5.4 | 44.0 s | 4,096 | 93 tok/s |
| gpt-5.4-mini | 13.0 s | 2,356 | 181 tok/s |
| gpt-4.1 | 18.4 s | 1,542 | 84 tok/s |
Three runs on mini after the switch: 8.8 s, 16.4 s, 14.2 s. Vision was the risk worth checking, since meal photos ride the same model; both read a photo in about 1.7 s and described it equally well. The guardrails did not get weaker because the writer got smaller: the code fences model output with deterministic gates rather than trusting the prose.
Still open
2 · The paywall that wasn’t there
The build sold nothing, and the server was still charging ten a day.
What a tester said
What I got wrong first
What settled it
The second defect
What shipped
3 · The screen nobody could test
The workout screen had 10,700 lines and no test that could see it.
What testers saw
Why it got that far
What the audit found
And the timer
What shipped
What I take from it
- Measure before you believe the error message. Three of these looked like outages and none was.
- Read the user’s actual data. Two wrong theories cost more than opening one Firestore document.
- When you fix a bug, ask what else you broke, and write the answer down in the commit.
- Pin invariants in source with a ratchet, so the next change cannot quietly undo this one.



