luca rarau

Case study · Gymswell · August–September 2026

Three bugs from the beta, and what each one taught me.

Gymswell is a fitness app I have built since October 2025: React Native, an AI coach on OpenAI, native Swift for the Lock Screen. When it went to testers in August, three problems came back that were not what they looked like. This is how I found each one, with the numbers I measured on the way, taken straight from the commit log.

1 · The coach

The AI was answering. We hung up at 30 seconds and blamed the wifi.

What testers saw

Every AI feature started failing with “AI request timed out (30s). Please check your internet connection.” Nothing was broken. The requests were succeeding.

What I measured

Three consecutive analysis-sized requests against the live deployment took 45.7 s, 42.6 s and 40.2 s, every one returning HTTP 200 with a good answer, no throttling, 999 requests left on quota. A trivial prompt came back in 1.4 s, which is why it looked like an outage.

The cause

In July the provider had moved from a small model to a larger one that writes about 4,000 tokens for a workout analysis at 93 tokens a second. Four thousand tokens takes forty-three seconds no matter how good the connection is. The 30-second cut-off had never been enough for a full answer; it only started biting when the answers got longer. And the dialog sent people off to debug the one thing that was working.

What I tried

Capping the output tokens was the tempting fix and the wrong one: the model stops naturally at about 4,000, so a lower cap does not make it concise, it cuts the answer mid-sentence, and a test that only checks a response came back never sees it.

What shipped

First, the timeout moved to 90 s on the client and 120 s on the proxy, so a timeout fires in one place and it is the one that can explain itself. Then the coach moved to the smaller model. Measured on the same prompt:
ModelTimeTokensSpeed
gpt-5.444.0 s4,09693 tok/s
gpt-5.4-mini13.0 s2,356181 tok/s
gpt-4.118.4 s1,54284 tok/s

Three runs on mini after the switch: 8.8 s, 16.4 s, 14.2 s. Vision was the risk worth checking, since meal photos ride the same model; both read a photo in about 1.7 s and described it equally well. The guardrails did not get weaker because the writer got smaller: the code fences model output with deterministic gates rather than trusting the prose.

Still open

The wait itself. A long answer takes as long as it takes; the honest fix is streaming, so the user watches it arrive. That is a product decision and a bigger change than the bug deserved.

2 · The paywall that wasn’t there

The build sold nothing, and the server was still charging ten a day.

What a tester said

“even at the gym when i tried clicking the break pr and stuff it didnt work it said failed”. The modal said “Something went wrong. Please try again.”

What I got wrong first

Twice. First I blamed an API key that had been exposed in builds 155 to 169; it was exposed but never rotated, so those builds still worked. Then I blamed the trial and promo lanes missing from the server’s idea of a Pro user; that was a real defect, worth fixing, but the tester had neither, so it was never his problem.

What settled it

Reading his actual user document. Usage count for the day: exactly ten. Tier: free. No trial, no promo. The 1.0.0 build ships with monetisation switched off, so there are no products, no paywall, and the app believes nothing is metered. The AI proxy never got the memo and metered every caller at ten calls a day. It used to work because metering had been client-side, where the flag switched it off; moving it to the server reintroduced a limit the rest of the app assumes does not exist.

The second defect

It failed silently. The modal handled timeouts and network errors and dropped everything else into “Something went wrong”, including the server’s own limit refusal. No paywall, no mention of a cap, no way to tell a limit from a broken app.

What shipped

The server short-circuits metering while the monetisation flag is off, and a source ratchet pins the two flags together so they cannot drift apart again. The server’s Pro check now reads all three entitlement lanes the app reads, with the trial window derived from its start date rather than a boolean that goes stale. Both modals recognise a limit refusal by either route and go to the paywall, and deliberately do not match timeouts, because a weak gym connection must not start selling a subscription.

3 · The screen nobody could test

The workout screen had 10,700 lines and no test that could see it.

What testers saw

Build 38 died on its error boundary before you could log a set. A hook crash, in the one screen the whole app exists for.

Why it got that far

The screen had grown to 10,700 lines with nothing able to render it under test, so the crash had no test to fail. Fixing the hook was the small part. The real work was making the screen testable, so the next regression fails on my machine and not on a tester’s phone.

What the audit found

Asked whether the crash was the only thing I had broken, I went looking rather than answering. It was not. The set table reserved a note column whenever any row had a note, but only rows with a note rendered it, so in a mixed card two sets put their weight and reps in different places. Note-less rows now hold the column open, and a test pins that both rows carry the same number of cells.

And the timer

The same week’s tester reports included a rest timer that “works sometimes”. It was never miscounting. A paused rest displayed the configured target instead of the banked remainder, so 03:00 appeared to drop 98 seconds in a frame when you pressed Start. With no pause behind it, the target is the remaining time, which is why it only failed sometimes.

What shipped

Tests that render the workout screen, the note-column fix, the paused-timer display, the Live Activity pointer that had no writer, a units fix in the Swift widget, and a reps field that could not type a dash. Each with the defect that produced it recorded in the commit, so the next person does not have to rediscover it.

What I take from it

← Back to Gymswell on the portfolio