20-PEOPLE-TO-3(7)2026-10-02

$ cat 20-people-to-3.md

20 people to 3: the part of an LLM pipeline that isn't the LLM

2026-10-02 · 3 min read · #llm #evals #human-in-the-loop #production

Twenty people. That's how many it took to keep one boring process running at a billion-dollar company.

The job was vehicle service history. Before a used car gets listed, someone has to know what's been done to it: services, repairs, the stuff that tells you whether a car was looked after or just driven. That data lived on a vendor's portal. No API. So every day, a team opened the portal, read the reports, and typed the fields into our systems by hand.

It cost about $10,000 a month. And it was the kind of work nobody wants, which means it's also the kind of work where mistakes creep in.

Now it takes three people. Here's how, and, more usefully, here's the one decision that made it work.

The pipeline is boring on purpose

Four steps:

  1. Fetch. A connector pulls the reports off the portal. No API means you work with whatever the portal gives you, so this part is mostly patience and retries.
  2. Parse. Turn the documents into text the model can work with.
  3. Extract. An LLM pulls out the handful of fields the business actually uses.
  4. Route. This is the interesting one.

Nothing in steps 1 to 3 would impress anyone. That's fine. A model can read a service report pretty well. That was never the hard part.

The hard part: deciding what the model doesn't get to decide

Here's the thing about LLM extraction. It's right most of the time, and when it's wrong, it's wrong confidently. A misread number looks exactly like a correct one.

So the question wasn't "can the model extract the fields?" It was "which reports can we trust without a human looking?"

Every extraction gets a confidence signal. Reports above the line go straight through. Reports below it land in a review queue, where a person checks them before anything touches a listing.

That queue is why it's 3 people and not 0. And honestly, 0 was never the goal. The goal was to stop paying humans to read the easy reports, so they could spend their attention on the hard ones.

Where the line goes

Picking that line is the actual engineering. Set it too low and wrong data slips into listings. Set it too high and you've rebuilt the 20-person team with extra steps.

There's no clever trick here. You take real reports, check them by hand, run the pipeline over the same set, and look at where the model is wrong. Not on average, field by field. Some fields the model basically never gets wrong. Others it fumbles in predictable ways. The line follows from that, and it's a business decision as much as a technical one: how much does a wrong field cost you, compared with one more report in the queue?

The queue also keeps the system honest after launch. The reviewers see exactly the reports the model was unsure about, so they're the first to notice when something drifts.

What I'd tell someone building one of these

  • Spend less time on the prompt than you think and more time on the routing.
  • Label real data before you pick a threshold. A hundred hand-checked reports will teach you more than any benchmark.
  • Design for the review queue from day one. The humans in the loop aren't a fallback. They're part of the system.
  • Measure what the old process cost, in money and in people. "20 to 3" is what got this taken seriously. "We added AI" wouldn't have.

The LLM was the small part of this project. The big part was deciding when not to trust it.

cd ..