Skip to content

Back to Articles

6 min read ·

Stop paying a poet to fill a dropdown

Most of the AI in business software is a switch statement with a fuzzy input. This week I ran one over 11,450 emails with a model that cannot write a sentence, and the bill was about $1.88.

For years the standard way to pick the case has been to hire a poet. You write a prompt that begs for one word. The chat model sends back a friendly paragraph, or the word with a period after it. You write a regex to fish the label out, a fallback for the day the regex catches nothing, and a retry for the day the fallback guesses wrong. That is the prose round trip, and you pay per token in both directions to get back the enum you started with.

The fair reply turned up in the Hacker News thread for Jev's launch: "You can enforce structured output from an LLM too." True. That fixes the regex. It does not fix the confidence, which is one more field the model writes about itself, and this month I called that a mirror check. A confidence you can threshold is the one thing the prose round trip never had. Primeline, which ran a pre-registered test of Jev, wrote the rebuttal for me: "you can get the same answer out of any LLM... You cannot get the same number."

Jev is built for that number. TypeSafe launched it on September 15, 12 days ago, and calls it a System One model, after Kahneman's fast, intuitive mode of thinking. You hand it a state and a set of typed questions, each one a choice, a score or a yes-or-no question that TypeSafe calls a noul, and every answer comes back in one parallel pass with probabilities attached. Nothing else comes back. Jev cannot explain itself, and its own docs tell you to keep the arithmetic in code.

TypeSafe's CEO said it more plainly in the launch thread than any of the marketing: "choice maps to match statement, score maps to sorting, noul maps to if-statements." Input costs $0.042 per million tokens. Output is free.

Jev also makes you write the answers down before you ask. "You define the possible answers through primitives," the docs say. That one requirement is the homework the prose round trip let us skip. A chat model will take a question nobody finished thinking about and answer it anyway, at length.

Skeptics on Hacker News say BERT could do this years ago, and one points out that zero-shot classification with Facebook's BART "has been available for years." They are right, and it makes our habit look worse. The boring tool existed the whole time, and we reached for the chat model anyway because it was the one on stage.


A shared operations mailbox at the freight forwarder and customs broker where I work in Gurugram holds 19,043 emails from May 12 to September 23. Nobody had a list of what those emails are about, and I did not guess one. I ran bge-m3 embeddings and clustering locally first, at no API cost, and the clusters found the categories. Then Jev, through typesafe-sdk with jev-latest, read each email with its thread and gave it a primary team out of seven departments, plus sub-categories and a few signals, one of which is whether the email is waiting on us. It was a batch run over past mail, not live routing.

The latest version of my questions classified 11,450 emails for about $1.88, or 16 cents per 1,000, at about 15 emails a second. The whole mailbox would cost about $3.10. All of it is input, about 3,900 tokens per email, because every call carries the whole thread and the full rubric again and TypeSafe has no prefix caching. When output is free, the prompt is the invoice.

An earlier run, on an older version of the questions, stopped at 11,250 of 17,888 emails when my TypeSafe credits ran out and the API answered HTTP 402. It stopped the run and said why, which is more than most fallbacks do.


Jev's primary team matched the independent cluster labels on 72% of 8,825 emails. On which teams were involved at all, the match was 85%. Neither number is accuracy, since both only count how often two methods agree, and the only human check behind them is one label from a 200-email review sample.

TypeSafe's charts have the same hole, and TypeSafe says so. Its reference answers are "the average of GPT-6 Astra and Fable 5.1", which the company admits "biases answers towards OpenAI and Anthropic's models." Agreement with some other method is what most of us are calling an eval this year. The 0% hallucination figure is worse. The launch post explains it in two sentences.

Our number is not empirical. Schema matching is guaranteed, thus we can confidently add 0% into the plots.

By that standard my JSON parser has never hallucinated either. A valid label can still be the wrong label, and the CEO conceded in the thread that "it's also possible to be confidently wrong."

Then I found the flaw in my own question. Of the classified emails, 39.6% involve two or more departments. I asked for one primary team, a single answer, of a mailbox where two emails in five need more than one team. Jev answered the question I wrote.

The label was the bug.

Jev's calibration error in Primeline's test was 0.012 on yes-or-no questions and 0.086 on pick-one questions, about seven times worse, so mail like this should get a yes-or-no question for every department.

Primeline also showed that Jev's confidence on a choice is the top option's probability, rescaled for the number of options, so it tells you how sure the first guess was and nothing more. Everything under 0.6 went into a review pile, 10.9% of the classified mail, about one email in nine.


If you already have labels and categories that hold still, skip Jev. On Banking77, MindStudio found that a 22-million-parameter encoder with logistic regression scored 93.2% at 8 ms on a CPU, against 80.1% for Jev zero-shot. Train the small model and run it yourself. I can't yet. My categories came out of a clustering run this month, which is the situation a zero-shot decision model exists for.

Either way, own the exit. Jev is proprietary and hosted in the United States, and the company behind it, founded in 2024, was reportedly in talks at a valuation above $10B nine days after its $40M seed. Nobody knows what that company looks like in a year.

The interface is cheap to copy. Kev, an open family of Jev-like models under Apache 2.0, gets within a point of Jev out of distribution at its largest size. An independent reproduction on a DeBERTa encoder cost about 25 cents of compute to train. Keep your questions and thresholds in your own code, so the model behind them is one line of config.

Write the enum first.