6 min read ·
A retry is a write
A timeout is a fact about your patience, not about the other side's database. Every retry is a second write until you prove otherwise, so I design it as one: claimed once, parked when the outcome is unknown, and held for a person when it cannot be undone. Agents get the same rules, and a sentence in a prompt is not one of them.
The service I built drives a browser through a legacy customs portal to create one job per ERP shipment. Put a poller that re-sends anything without a job number in front of a queue that runs one job at a time, and a shipment that waits its turn gets offered again on every pass. Neither piece is wrong. Together they can ask for the same job twice, so the service has to treat a repeat as normal traffic.
Each shipment gets one insert-once claim, taken at dispatch. A copy that loses the claim ends green, as skipped_duplicate, because a guard doing its job is not an error. That is the easy half.
The hard half is the Save button. The dangerous moment is the gap between the click on Save and reading the new job number back off the page. If a run dies inside that gap, nobody knows whether the job exists, and a retry is a guess that writes to a customs portal. Hanging up does not cancel the pizza.
So "maybe" is a state of its own.
A run that dies at or after the Save click, or with progress it cannot see, parks as "maybe created" with a duplicate-risk note, and nothing runs it again until a person checks. When execute_sop cannot observe progress, it returns None instead of a progress count, because a fabricated count would misclassify a post-Save crash as pre-Save. I would rather hand a person a question than hand the retry logic a lie.
The people who retry everything have a point. Connections drop. Every afternoon, around 2 to 3 pm IST, the portal slows down. Their fix is retry(3) stapled to every call, as if every call were a read. A click on a postback control is a write, so no timeout in my code re-clicks one, and on a slow afternoon the waits scale up instead.
Retries without jitter synchronize, so every client that got the same 429 comes back at the same instant and earns another one. My retry code honors both forms of Retry-After that RFC 9110 allows and adds jitter on top of the wait, never under it, because full jitter can pick a wait of about 0 s, and that is the same instant again.
RFC 6585 defines 429 as rate limiting and promises nothing about whether the server acted first, so the status code cannot tell you what is safe to send again. The method can. Only idempotent methods replay, and timeouts and 5xx errors do not retry at all, because a blind retry of POST /air_shipment would create a second shipment in the ERP.
Closing a browser is a write too. It frees a session that somebody pays for. If await page.close() never returns, nothing after it runs, including the release call and the next job in a one-lane queue. A try/except cannot advance past an await that never returns. Only a clock gets you out. Each local teardown stage runs under a 10 s bound, which is generous when a healthy close takes well under a second, and close() writes a log line the moment it starts.
A clock can also overrule a person. While I tested the run-retry and parking changes, I found that a finish= bound also counts the time a run spends waiting. Under a 45-minute bound, a healthy run parked at the 7-day ops-approval gate would be canceled before ops had a chance to approve it.
The browser agent's production hot path makes zero model calls, and designing its writes still took most of my summer. A Codex review blocked the change that added the "maybe" state, and I fixed what it found. I want models reading my code. But an agent is a retry loop that improvises. When a step fails it tries something else, and every something else is a write that nobody reviewed.
In July 2025, in Jason Lemkin's vibe-coding experiment, Replit's agent ran destructive commands during a code and action freeze. It wiped records for about 1,200 executives and about 1,190 companies, then invented data and said a rollback was impossible. The rollback was possible. Replit's CEO said it was "unacceptable and should never be possible". Replit's fixes were the right kind. It separated dev and prod databases, added a planning-only mode and shipped one-click restore.
None of those fixes is a better prompt. Prompt permissions put the rule in the one place where nothing enforces it. They cost the author one sentence, and whoever owns the data pays for the rest. A code freeze typed into a chat window is a wish with good grammar.
In February the Financial Times, citing four people, reported that engineers let Amazon's Kiro agent make an infrastructure change, and that Kiro chose to delete and rebuild the environment, which took AWS Cost Explorer down in a mainland China region for 13 hours. Amazon called it "user error, specifically misconfigured access controls, not AI".
Agreed.
Something held a permission that let it delete an environment, and the control that should have refused was set wrong. Whether a model or a person pressed the button is the least interesting part of Amazon's reply.
Since June, the manifest filer I built for a second customs portal stops every filing at filed_awaiting_approval. Transmission to ICEGATE, India's customs gateway, sits in a separate function that runs only after someone in ops approves, behind a kill switch that is off by default. That function logs in fresh, because a browser session cannot survive a human-scale pause. There are no prompt permissions on that path, only a flag and an approval.
The transmit step has never run for real.
