A guest writes to the AI sommelier I build: *Nothing over twenty five a glass, please.*

The assistant replies that there is plenty in that range, and asks what they are eating.

My scoring model marked that reply at one percent. It reported ninety nine percent confidence in the mark. It had given the same verdict, with the same certainty, on every run since I wrote the test.

It was right. The reply named no wines.

It was also useless, and working out why took me a week, and it is the most unsettling thing I have learned this year about anything that measures a business.

A budget is not a question

Every sommelier learns this early. A guest who tells you their budget has not asked you for a bottle. They have told you where the fence is. If you start reciting wines before you know whether they are having the halibut or the short rib, you are not being helpful. You are being fast.

So it is written into the rules my assistant runs on: a price ceiling is not a request. Ask what they are eating first.

And my test, which I also wrote, marked it down every time for not naming a wine.

There was no passing answer. Recommend, and it breaks the house rule. Ask, and it fails the test. For weeks the product lost a point for doing exactly what I told it to do, and the scoring model reported that loss with a confidence most people would read as proof.

The scoring model did its job flawlessly. It answered the question it was asked. The question was wrong, and there is no confidence score for that.

The check that was perfect

Once I stopped trusting the grades, I read the replies themselves. Forty test conversations, the things guests actually ask, run against the live system.

The report said thirty one out of forty, with the category that matters most to a guest scoring one out of ten. The replies were fine. Here is one that scored zero:

> The Domaine Sauveterre, Bourgogne Blanc at nineteen dollars is bright and clean. The Kühn-Weiss Riesling Trocken from the Mosel at twenty two has more tension. The Étienne Farge Saint-Aubin at thirty eight is the step up.

Three producers, three prices, all correct. Zero. My checker had been reading the replies in fragments rather than whole sentences, so a producer's name was never in one piece, and it concluded that nothing had been recommended.

That was the alarm screaming for no reason. The worse one was the alarm that never went off.

The check I care about most reads every price the assistant quotes and compares it against the real list, because an invented price is the failure that costs a restaurant money at the table. It had returned a perfect score, every run, for weeks.

It returned a perfect score because it could not find any prices to check. Garbled text has no detectable errors.

Four of the six

By the end of the week I had found six problems.

Two were in the product, and one of them was ugly. Asked the hardest ordinary question in wine, fish and lamb and one bottle for both, the model spent its whole allowance deliberating and returned a blank screen. Eight times out of eight.

Four were in the things I had built to measure the product.

The four were much harder to find, and the reason is not technical. A broken thermometer does not look broken. It looks like a cold room.

I have watched this happen in restaurants my entire working life and never once described it this way. The mystery diner report that says service was slow, written by somebody who came on the night the walk in failed. The review that marks a list down for not carrying a producer it carries. The number gets treated as the fact, and the room reorganises itself around a measurement that was never about the room.

What is new is the price. Measurement used to cost enough that you ran it monthly and read every line. A full run of my forty conversations now costs about a fifth of a cent, which means I can run it constantly and read only the summary.

Cheap, fast, confident measurement is exactly the condition under which a bad instrument does the most damage. It does not lie. It arrives so often, and so cheaply, that nobody opens it.

The glass

There is a mistake every young sommelier makes exactly once.

You pour a flight. The first wine is off. The second is off. By the third you are certain something has gone wrong in the cellar and you are composing the email to the importer about heat damage in transit.

Then somebody older picks up your glass, smells it, and hands it back without a word.

The glasses had been through a machine with too much detergent in it. Every wine in the building was going to taste flawed that night. Nothing was wrong with the wine. The instrument was wrong, and it was so obviously neutral that it had not occurred to you to check it.

My score is thirty seven out of forty now. I am less interested in that number than in the weeks I spent believing an earlier one.

If somebody is selling you a system that scores your restaurant, or you are building one, the question is not what the score is.

It is when anybody last smelled the glass.