Skip to content

Back to Articles

5 min read ·

Self-healing is a smell

Two weeks ago I caught my agent's self-healing layer paying Gemini to search white screenshots for popups after every step. The three flaky clicks I was chasing in testing were one bug, in one click wrapper. An automation that needs a model to heal it has a bug that one engineer with a trace could fix.

The agent creates customs jobs on a legacy portal built on ASP.NET WebForms. The portal has no API, so the agent drives a real browser through written SOPs, mostly fixed Playwright steps. At the start of the month I was testing the job-creation SOP, and it kept timing out at a different click. Checklist in one log, Print in the next, Save in the third. It looked like three bugs in a flaky portal, and the obvious move was to make the agent tougher.

It was one bug.

The portal's popup controls do two things on a single click. They call window.open(), and they fire a postback. Playwright's page.click() timed out even when the popup had opened fine. The old code went looking for the new window only after the click returned. By then the click had timed out, and the window it had just opened sat there, orphaned. On top of that, a full-screen UpdatePanel loading overlay swallowed pointer events while the page worked. Between the orphaned window and the overlay, which click timed out changed from run to run.

The fix went into the one wrapper every click passes through. It expects the new page before it clicks, with context.expect_page(), waits for the overlay to clear, skips the navigation wait after a postback click, and runs a hook first to re-show the hover menus. One wrapper, not three patches. I ran it live on the portal against one job, and both checklist PDFs came out the other end.

Last December I wrote here that the model belongs at the edge, on the popup nobody wrote down, and I gave the agent a self-healing node to put it there. After a goto or a click, Gemini's vision model searched a screenshot for a blocking popup and tried to close it. It was built for exactly the problem I had this month, where something on the screen gets between the agent and its click. In the same piece I asked whether a vision call after every click was a fair price, and said to ask me in six months. It has been six months.

So what was it looking at in those test runs? It screenshotted "the frozen first page", which, in the words of my own PR, "produced blank/white screenshots". The healer never followed the new window either, the same blind spot as the code it was watching. After each step the agent paid for a vision call on a white rectangle and asked it to find the thing in the way. I assume it said no.

This is flake insurance. You pay a premium on every step so that nobody has to read the trace. It feels like diligence. It is laziness, billed per step. I can't tell you what it cost, because I never counted.

The pitch for flake insurance sounds grown-up. Portals change and selectors rot, so put a model in the loop and let it look. On a page nobody has mapped yet, or in a demo, that is a decent first guess. Everywhere else, read the trace instead.

A red step leaves a step name and a screenshot of the right window. A healed step leaves a green run. Even with every test run red, Checklist, Print and Save still looked like three flaky steps. I got lucky. My healer was blind, so all three stayed red long enough to line up. Give it working eyes and it closes a dialog here and presses Escape there, and I get fewer red runs, a ticket about a flaky Save button, and the same bug.

Then there are the quirks no model has seen, because they belong to one portal. The day after the click fix I found the next one on fresh test jobs, at the checklist step. A fresh job generates its own checklist when the popup opens, and the agent clicked Generate anyway. That extra click raised a confirm dialog the SOP had no step for. A vision model shown that popup sees a button called Generate. Of course it clicks. My code did. I took the click out, and a new job's checklist now generates on its own in about 50 s.

Some steps no screenshot can judge. The checklist PDF can run to 929 pages, and a picture of the screen can't tell you whether the bytes that came back are a PDF at all. The download step checks the bytes instead. It retries up to three times, and it keeps only a file that starts with %PDF- and weighs at least 1 KB.

I kept the healer for the page nobody has mapped. It has been off by default since June 3. Set self_healing: true and it comes back, and this time it screenshots the active window. Our SOPs are almost all fixed Playwright steps, and those handle their own popups. A healer is a prototype of an error handler, the way an agent is a prototype of a script.

What is yours looking at right now?