4 min read ·
Operations already wrote your spec
"Uses Playwright for deterministic actions (0 LLM calls)." That line comes from the first YAML SOPs I wrote for a customs portal, last month. The operations team's standard operating procedure, the one people follow every day, is the spec. Encode it step by step, and stop asking a model to guess it.
I joined Hexalog, a freight forwarder and customs broker in Gurugram, in September. My first Linear ticket, in October, was called "Browser Automation". The target is a legacy ASP.NET customs portal with no API. Agent demos avoid software like this. Its buttons post the whole form back to the server, and its popups turn up when they please.
The customs team has a procedure for the work I'm automating, and my job is to write it as YAML, one step at a time, so that a browser can follow it too. The procedure was written for a person, and a person fills its gaps without noticing. A browser has no such instinct, so each gap goes back to the people who do the job.
My own first version got this backwards. It handed Gemini a YAML SOP and a screenshot and let the model pick each next action, so the procedure sat right there in the prompt and every click was still a guess.
On November 13 each step got a method field that says who does it: Playwright, the model or a person. Fixed steps go straight to the tool and never reach the model.
The model keeps the part that no SOP can predict, such as the popup that appears on a bad day. On December 4 I added a node that runs after each goto or click, and Gemini's vision model looks for a blocking popup and gets three tries per page to close it. The next day every model decision started to record its reasoning and a confidence score.
Is a vision call after every click a fair price for the odd popup? I don't know yet. Ask me in six months.
Meanwhile AWS and GitHub spent 2025 reinventing the spec. In July, AWS launched Kiro in preview, an IDE built around "specs", its name for requirements and design files. In September, GitHub released Spec Kit, an open-source toolkit whose four phases are Specify, Plan, Tasks and Implement. Its announcement says coding agents "excel at pattern recognition but still need unambiguous instructions." GitHub calls specs "living, executable artifacts".
Welcome to operations.
Spec amnesia is building automation as if the company had never written down how the work gets done, and then applauding a tool vendor for suggesting that somebody should. A customs SOP was a living, executable artifact long before anyone put it in a repo. The executor is a person with a login.
The motive is vanity, and I had a mild case of it in October. An agent that figures it out demos beautifully. An agent that follows the SOP demos like a checklist, because it is one, and nobody claps for a checklist. So you throw away the flat-pack instructions and hand the Allen key to whoever sounds most confident.
The bill for spec amnesia arrives after the demo, in the pilot. In August the MIT NANDA report claimed that 95% of generative AI pilots show no measurable P&L impact. Its method is thin (150 interviews, a survey of 350 employees and 300 public deployments), and critics called the number a headline rather than a measurement. Fine. Throw out the number and keep the diagnosis, which blamed integration and "learning gaps" and not model quality.
Integration is the polite word for the week somebody finally has to say what the work is.
The model belongs at the edge, on the popup nobody wrote down, and it still has to earn that spot. The middle of the run belongs to the SOP.
In operations there nearly always is one, and it is the only spec in the building written by the people who pay for its mistakes.
