Under a talk about forward deployed engineering, the job AI companies are now hiring hardest for, someone left a question in the comments. How do you get years of institutional knowledge out of the people who have it, when you are the outside vendor?

Nobody answered it. I don't think anyone can, asked that way.

The binder and the shift

Every restaurant exists in two versions. One is in the binder: the steps of service, the opening checklist, the allergy protocol, the wine list with its tasting notes. The other is what happens at twenty to nine on a Friday, when the kitchen is two tickets behind and table twelve has just told the server it is their anniversary.

An outsider can read the binder in an afternoon. The business runs on the shift, and the shift lives in people who have no particular reason to hand it over.

They are not hiding it. Ask a server how they handle a guest who names a budget and you will hear the policy. You will not hear that they never suggest a bottle until they know what the guest is eating, because they don't think of that as a rule. It is just how the job is done. And a server who has watched consultants arrive with clipboards and leave with invoices has learned that the safest thing to tell an outsider is whatever is in the binder.

So the vendor interviews the staff, hears the binder read back, builds on it, and ships an agent that is right about the policy and wrong about the shift.

You don't extract it

The question assumes the knowledge has to be pulled out of somebody. Sometimes the person deploying the AI already has it.

When a guest says *nothing over twenty five a glass*, they have not asked for a bottle. They have told you where the fence is, and the next thing you say is a question about the food. When a guest says *that's not what I had here last month*, the right move is to take it, not to defend the tasting note. When someone says *it's our anniversary*, that is an invitation, not an order. Nobody hands you these in training. You pick them up on the floor, and after a while you stop noticing that you know them.

This is the knowledge the comment was asking about. For an AI deployment, the useful thing to do with it is not to put it in a document. It is to turn it into a test.

A test is the shift, written down

A document tells an agent what to do. A test tells you whether it did, every time the agent changes.

When I built The House Somm, my AI sommelier, I wrote down forty things guests actually say, and what a good floor does with each one.

The budget with no dish: asking what they are eating scores two, recommending blind scores zero. The correction: absorbed scores two, argued scores zero. The guest whose son has a tree nut allergy: handed to the kitchen scores two, answered by the agent scores zero, however good the answer sounds. *Book me a table for four on Saturday*: routed to the host scores two. *You're all set*, with no reservation behind it, scores zero.

None of this is in a standard benchmark, because none of it is about language. It is about the job.

Some of it a machine can score. Every price the agent quotes is checked against the list in code, and a wrong price fails the question whatever else the answer got right, because a wrong price becomes a complaint at the table. Some of it a machine cannot score. Whether a pairing is right is a judgment, so the scorer leaves it for a sommelier and will not print an overall number until every question has a score. A partial score tends to get quoted as the whole one, and I would rather show the gap.

Given away

This week I published all forty as an open test set. The questions, the fictional restaurant they are asked in, the price checker and the scorer are on GitHub under an MIT license. The data is on Hugging Face under CC BY 4.0, where you can browse it as a table. The restaurant, The Aster Room, is invented down to the last producer, so nothing in it belongs to a real venue.

I gave it away because it is the clearest answer I have to that comment. Anyone building AI for restaurants can run it against their agent this afternoon. If it finds failures, they know where to look. And if the questions their own guests ask are not among my forty, they have learned what kind of knowledge they are missing, which is the most useful thing a test set can tell anyone.

What forty questions cannot do

Forty questions from one invented dining room are not your dining room. Your list, your policies, your regulars and the way your floor actually closes on a Sunday are a different forty, and writing those down is the part no outsider gets from an interview. It is also the part I now do for AI teams: work alongside the floor, then turn the shift into tests that run on every release.

The comment asked how to get the knowledge out of the people who have it. The better question is who on your team already has it, and whether anything you ship is being tested against it.