Journal

Which model should we use? Build the evaluation set first

The question comes up in the first call, every time. The honest answer is that nobody knows yet, and that the way to find out costs less than a week.

Published 3 min readCardon Studio

The question behind the question

When someone asks which model we will use, they are usually asking two things. Will it be good enough, and will it be expensive. Both are answered by the same object, and it is not a benchmark from a vendor's page. It is a set of cases from your own work, with the answer your team would accept written next to each one.

Public benchmarks measure how a model does on someone else's problem. Your customers do not ask those questions. They ask about your contract terms, your product codes, your refund rules, in your language and with your typos. A model that tops a leaderboard can still fail at that, and a smaller one can pass it.

What an evaluation set is

A list of inputs and expected outputs. An input is whatever the system will receive: a customer message, a scanned invoice, a paragraph from a filing. The expected output is what a careful person on your team would produce, or the rule that decides whether an answer is acceptable. Sometimes it is an exact value. Often it is a short checklist: mentions the right clause, does not promise a date, cites the source document.

The set does not need to be large to be useful. A few dozen cases that cover the real spread of your work say more than a thousand that all look alike. What matters is that they include the awkward ones: the ambiguous request, the document with two versions, the question that should be refused.

Who writes it

Your team, with us in the room. The people who answer these questions today know which mistakes are embarrassing and which are expensive. We bring the structure and the tooling. They bring the judgment. Writing the set usually takes two or three working sessions during the assessment phase, and it is the most valuable document the project produces, because everything after it is measured against it.

It also settles arguments early. When two people disagree about what a good answer looks like, they find out while writing a case, not after launch.

Then the model is a measurement

With the set in place, choosing a model stops being an opinion. We run the candidates, hosted and open, against the same cases with the same retrieval and the same rules around them. Each one gets a pass rate, a cost per request and a latency. The smallest model that passes is the one we recommend, because it is cheaper to run, faster to answer and easier to replace.

Often the result surprises the team. A model a fraction of the size of the famous one passes, once the retrieval gives it the right paragraph and the checks catch the answers it should not give. The model is one part of the system, and rarely the part that decides whether it works.

Written down, so it can change

The choice goes into the assessment document with the pass rates next to it. Six months later, when a provider ships a new model or raises its prices, the same set runs again and the decision is revisited on evidence. Nothing has to be remembered, and nobody has to defend a preference.

This is also why the set stays yours. It is the one thing that lets you switch providers, switch studios or bring the work in house without starting over.