5 min read ·
An agent is a prototype of a script
Last fall my browser agent's first task was an Amazon search, and Gemini chose every click. By January 7 the first procedure for real customs work had 30 fixed steps and 1 model step, and I'm prouder of that 30 than of anything the model ever clicked. Let the model find the path, then ship the replay, because paying for reasoning on every run of a task you already understand is waste.
The best case for agents arrived this month. On February 5 Anthropic released Claude Opus 4.6 and described a C compiler that a team of 16 parallel agents wrote in about 2,000 Claude Code sessions over 2 weeks, for about $20,000 in API cost. It's about 100,000 lines of Rust, and it builds Linux 6.9 on x86, ARM and RISC-V. That's a real achievement.
But look at what the agents were given. A C compiler comes with a written standard, a huge test suite and GCC as an answer key. Then look at what shipped. Not the agents. A compiler, a program with no model anywhere inside it, and still the only kind of code generator I trust.
The $20,000 was a one-time fee.
I'm building an agent too, at Hexalog, a freight forwarder and customs broker in Gurugram. It exists to create customs jobs on a legacy ASP.NET portal that has no API, and that portal comes with none of the compiler's gifts. It has no spec and no test suite, and there's no second portal to diff against. The closest thing to a spec is what the operations team does every day, so my agent follows their standard operating procedures, or SOPs, which I write down as YAML. I made that case in December.
Having the SOP didn't stop me from paying a model to read it.
Most of us started last year with the model in charge, and so did I. My first commit was a LangGraph loop in which Gemini 2.0 Flash read the YAML and a screenshot and picked every next action for Stagehand to carry out. On October 14 I finished a ticket called "Cost optimization - intelligent model switching". The switching was a complexity analyzer that escalated a step to gemini-2.5-computer-use-preview-10-2025 whenever the step looked hard or kept failing. I never measured what it saved. I'm not sure it saved anything.
Every run paid a model to find a path that the SOP already described. That is the reasoning tax. Nobody hires a cartographer for the commute. Escalation raised the rate. When the cheap model got stuck on a step, my answer was a more expensive model thinking harder about the same step.
The way out was one field. Since November 13 every step in an SOP has carried a method. A playwright step is a fixed action, a stagehand step needs the model, and a hitl step goes to a person. A Playwright step skips the decision model and logs "Playwright step - skipping LLM (0 cost)".
The model stayed, demoted from driver to exception handler. The popup node from December still checks the page after each click, and since January a popup it cannot close goes to a human.
The reasoning tax is paid in tokens and in variance, and only the tokens show up on an invoice. The same shipment has to produce the same customs job on Tuesday and on Friday, field for field, and a model that picks the next click from a screenshot can pick a different click on a different day. Since December mine has written down its reasoning and a confidence score for every decision, so it can give a confident reason for either click. The script explains nothing. It does the same thing every time.
The standard objection is that scripts are brittle.
Good. Brittle is loud.
When the portal moves a button, a script dies at a named step with a stack trace, and the fix is one line in that step. An agent that improvises around the change might click the right thing or the thing next to it, and it can report success either way.
None of this is my discovery. In December 2024 Anthropic published "Building effective agents", which split LLM systems into workflows, where code sets the path, and agents, where the model directs its own steps. Its advice was plain: "When building applications with LLMs, we recommend finding the simplest solution possible, and only increasing complexity when needed." A company that sells tokens for a living told you to buy fewer. That post was more than 9 months old when I made my first commit. I built the agent first anyway, and I don't regret it.
A prototype is how you find the steps that need judgment. Once you know the path, a model on every step is indecision with an API key.
The place to pay for reasoning is once, when the code is written, not on every run. Inside the customs job, the model is on its way out. This week I started the move to Stagehand v3, and the plan for it makes the YAML mode the production path and the LLM mode the "fallback / legacy path".
An agent that does the same job every day is a prototype nobody finished. Finish yours.
