Journal
How will we know it still works after launch?
Launch is the moment the system meets the questions nobody wrote down. Knowing whether it still works is a routine, not a feeling.
What changes after launch
Three things move on their own. The model provider updates the model, sometimes without a version change, and answers shift in tone or length. Your data drifts: new products, new policies, documents that contradict older ones. And your users learn what the system can do and start asking things the evaluation set never imagined.
None of these announce themselves. The system keeps answering, confidently, and the first sign of trouble is usually a customer or a colleague noticing. A system in production needs something watching it that is not a person reading every answer.
Rules that check every answer
Before an answer is shown, it passes a set of checks written during the build. Some are simple: the answer cites a source document, it does not contain a price unless the price came from the catalogue, it does not promise a date. Some are a second model asked a narrow question about the first model's answer. An answer that fails is not shown. The user gets a fallback, and the failure is logged.
These checks are the reason the small model in the previous note is safe to use. They also produce the most useful number in the whole system: how often answers are being blocked, and for which reason. When that number moves, something upstream has changed.
A weekly sample, scored
Every week, a sample of real requests and their answers is drawn from the logs and scored against your evaluation set, using the same criteria your team wrote before launch. New cases from the sample that the set did not cover are added to it, so the set grows with real usage rather than with our guesses.
The score is a trend, not a verdict. One bad week after a provider update is a signal to look. Three weeks drifting down is a signal to act. Because the same set has been run since the prototype, the trend is comparable across the whole life of the system.
Alerts for the things that cannot wait
Some changes should not wait for the weekly sample. Cost per request rising, latency crossing the limit your product tolerates, the block rate jumping, the retrieval returning nothing for a class of questions. These raise alerts inside your cloud, to your channels, with a line in the runbook that says what to check first.
Alerts are tuned to be rare. An alert that fires every day is ignored by the second week, and then the one that matters is ignored too.
The monthly report
One page, in plain language, for the person who owns the system and does not read logs. The trend of the weekly scores. What was blocked and why. The incidents, what caused them and what changed as a result. The cost. What we recommend changing next month, and what we recommend leaving alone.
It is written so it can be forwarded to a director without a call to explain it. If it needs a call, it is not finished.
Ours or not
This routine does not depend on who built the system. A team with an assistant that works today and nobody watching it tomorrow can start here: an audit of what exists, an evaluation set written with the team if there is none, monitoring in their cloud, and the first report a month later.
More notes
-
Can it run inside our own cloud? Yes, and it usually should
The system and the data it reads are yours. The account they live in should be too. Here is what that looks like in practice, and what it costs you in effort.
Read more -
Which model should we use? Build the evaluation set first
The question comes up in the first call, every time. The honest answer is that nobody knows yet, and that the way to find out costs less than a week.
Read more