Journal

How will we know it still works after launch?

Launch is the moment the system meets the questions nobody wrote down. Knowing whether it still works is a routine, not a feeling.

Published 3 min readCardon Studio

What changes after launch

Three things move on their own. The model provider updates the model, sometimes without a version change, and answers shift in tone or length. Your data drifts: new products, new policies, documents that contradict older ones. And your users learn what the system can do and start asking things the evaluation set never imagined.

None of these announce themselves. The system keeps answering, confidently, and the first sign of trouble is usually a customer or a colleague noticing. A system in production needs something watching it that is not a person reading every answer.

Rules that check every answer

Before an answer is shown, it passes a set of checks written during the build. Some are simple: the answer cites a source document, it does not contain a price unless the price came from the catalogue, it does not promise a date. Some are a second model asked a narrow question about the first model's answer. An answer that fails is not shown. The user gets a fallback, and the failure is logged.

These checks are the reason the small model in the previous note is safe to use. They also produce the most useful number in the whole system: how often answers are being blocked, and for which reason. When that number moves, something upstream has changed.

A weekly sample, scored

Every week, a sample of real requests and their answers is drawn from the logs and scored against your evaluation set, using the same criteria your team wrote before launch. New cases from the sample that the set did not cover are added to it, so the set grows with real usage rather than with our guesses.

The score is a trend, not a verdict. One bad week after a provider update is a signal to look. Three weeks drifting down is a signal to act. Because the same set has been run since the prototype, the trend is comparable across the whole life of the system.

Alerts for the things that cannot wait

Some changes should not wait for the weekly sample. Cost per request rising, latency crossing the limit your product tolerates, the block rate jumping, the retrieval returning nothing for a class of questions. These raise alerts inside your cloud, to your channels, with a line in the runbook that says what to check first.

Alerts are tuned to be rare. An alert that fires every day is ignored by the second week, and then the one that matters is ignored too.

The monthly report

One page, in plain language, for the person who owns the system and does not read logs. The trend of the weekly scores. What was blocked and why. The incidents, what caused them and what changed as a result. The cost. What we recommend changing next month, and what we recommend leaving alone.

It is written so it can be forwarded to a director without a call to explain it. If it needs a call, it is not finished.

Ours or not

This routine does not depend on who built the system. A team with an assistant that works today and nobody watching it tomorrow can start here: an audit of what exists, an evaluation set written with the team if there is none, monitoring in their cloud, and the first report a month later.