4 min read ·
The bug was in Oregon
Most flaky automation is not flaky. It runs somewhere you never chose, at an hour you never tested, and the fix is to check the map and the clock before you open the code. My browser was in Oregon.
I build a browser agent that creates customs jobs on a legacy ASP.NET portal with no API. For weeks, while I was still writing its SOPs, my test runs timed out at a different place each time: the login on one run, the document frame on the next. Each timeout looked like a bug in that step, so I treated it like one and tuned that step's timeout, then the next step's.
CI has a word for this, flaky, and a button for it, re-run, and a shrug that goes with both. It passed the second time. All three are ways to stop looking. "Flaky" sounds like a property of the test, as if some code were born nervous. Mostly it means nobody checked where the run happened, or when.
It is a lazy word, and for weeks it was mine.
Re-run gets one thing right. Sometimes the timeout is too tight. I doubled mine this week. The difference is that I knew where the browser was, and at which hour the portal slows down, before I touched the number.
Then I looked at the map. The cloud browsers ran on Browserbase in its default region, us-west-2, which is Oregon. The portal is hosted in India. From us-west-2 its pages took 23 to 33 seconds to load. From ap-southeast-1, Singapore, they loaded in about 2 seconds (my own measurement; I did not record how many pages).
So I had spent those weeks tuning timeouts against a vendor's default. A fill got 10 seconds and a wait got 30. A page that takes 23 to 33 seconds to arrive can beat a 30-second wait on one run and miss it on the next, and the timeout lands on whichever step draws the slow page. Nothing in that step's code says Oregon.
Setting the region was not enough. Stagehand's sessions.start took a region parameter and ignored it, and the Browserbase API showed the session still landing in us-west-2. Now the code goes around the SDK. It creates the session through Browserbase's REST API, with the region set, and connects Playwright to the connectUrl that comes back.
Then the fast region exposed a race.
In Singapore a sidebar click could now land while the accordion menu was still settling. In Oregon, the slow network had always given the menu time to finish. The fix was one flag, and my pull request for it says the rest: "These clicks never had force_click — they had always relied on incidental network latency to pace themselves."
That is load-bearing latency. Nobody designed it and nobody could see it. It was a folded napkin under a wobbly table. Speed a system up and you get the speed, plus every ordering assumption that the slowness had been paying for.
The map hid load-bearing latency. The clock is the other input, and it is harder to see, because it looks like luck. The portal slows down every afternoon. Around 2 to 3 pm IST a server-side CSV import can take up to about 10 minutes, so I raised the upload wait from 3 minutes to 12, on reasoning alone, because I could not make that load happen on demand. Every other timeout in the SOPs had a quiet-morning size: 15 seconds for a new window, 10 for a click, 30 for the loading overlay. A quiet-morning timeout is a coin toss at 2:30 pm.
One knob, SOP_TIMEOUT_SCALE, now stretches all of them, in place of the step-by-step bumps I made in June. It defaults to 2.0. I tested it the only way I could, by throttling the browser's network through CDP. Between 1 and 4:30 pm IST, a retry cools off longer, and longer again each time, because retrying into the afternoon slump only burns sessions.
The scale stretches the wait and never repeats the click, because "a window.open/postback control must not be re-clicked (double-submit risk)".
On a customs form, a retried click is not resilience. It is a second submit.
One number sat out of the knob's reach. A hand-rolled poll loop carried its own max_wait_time = 15, a bare literal that no config file would ever touch. It is 30 now. A timeout inside a loop is still a timeout. Grep for yours.
In June I argued that self-healing is a smell. Here a model staring at screenshots would have found nothing to heal. Every page arrived, half a minute late.
A stack trace has no field for a place or an hour.
