Amplified Judgment
amplifiedjudgment.org October 2026

Tell, Don’t Ask

AI, human judgment, and who gets to open the gate

Contents · 7 sections
Chirag S ChamanOctober 2026
Begin reading ↓

Key points

If you only read one page, read this one.

  1. Intelligence is getting cheap. Judgment isn't.
  2. Much of workplace AI is built for Ask: someone at a screen with a question. Work runs on Tells: the report, the handover, the count. That mismatch helps explain why AI has reached the frontline least.
  3. The last mile of a decision runs on facts the business has to supply: what happened this morning, at this site, with this customer. Models bring breadth. People bring the judgment call, and own it.
  4. Every tough call an expert makes today can help a less experienced person make the right one tomorrow. Precedent doesn't remove the gate. It changes who's qualified to open it.
  5. When an expert's decisions can travel, the business stops waiting for the expert to walk in. That's how the businesses that were supposed to be hard to scale get bigger without getting worse: right call, any site, any shift.

Where this came from

When was the last time someone on a night shift typed a question into a chatbot with a pallet jack in the other hand?

They don't. They text the manager. They snap a photo. They count what's on the shelf and write the number on the clipboard by the door. That's how most of the world's work gets reported. They tell: in the nightly report, the handover, the question somebody carries to the boss at the end of a long day. Nearly everything written about AI pictures someone else. A person at a desk, asking.

I came to that gap by way of other people's writing. A handful of essays on AI and judgment, read more times than I'd admit to a stranger. They're credited at the end, and they deserve your time. This essay is partly my reply.

About a year ago, I set out to build something practical: an AI system that could run the day-to-day of my three businesses, so I could get back to everything else. Ten months and four attempts later, I'd had more arguments with Claude than I can count (we're still friends), and I finally understood the problem. A frontier model is a Ferrari. Point it at any road and it drives beautifully. Point it at the same road on Wednesday and it has forgotten where the potholes were on Monday.

Case study · My caféThe 1 am freezer argument

I wanted AI to handle inventory and ordering for the café. The larger goal was to have it run most of the work that didn't involve serving guests.

The LLMs knew how the most profitable operations did it. I had SOPs for taking inventory, agents finding suppliers and getting their pricing, even a plan for which days to stock to optimize schedules.

What the model didn't know well was how NYC traffic would affect a delivery. If my 4 pm delivery arrived at 6 pm, I had to pay overtime. We also had limited freezer space. I had to forecast sales to know how much room would open up. We'd found an elegant solution to that as well. Then another edge case turned up.

It's 1 am Friday morning, and I'm in a heated argument with a frontier model agent. The order hasn't gone out because we won't have freezer space to store it. If we don't send it in the next hour, we'll be out of certain products by Sunday.

It suggests getting a chest freezer, delivered before 4 pm. But I don't have the floor space. No problem, it says: there are standing freezers that can fit. It's calculated the square footage (only 19 sq ft), and I'm now learning about the latest in freezing technology.

I finally snap: "WTF do I have to become a warehouse, when I have a Costco & Aldi just ten blocks from me? Paying someone every day to go shop is cheaper than rent for 19 sq ft."

"You're right. I did not…" it responds.

The next day, Instacart became my just-in-time supplier for 30% of what we buy at the café. Two weeks later, the restaurant next door followed.

So I built the thing around the Ferrari: a system that knows what today needs and remembers where the potholes are. The model underneath it is called AJAM. You'll meet it in part V.

Fair warning: that gives me a stake in how this comes out. The last part makes five predictions with dates attached, and I'll come back and grade myself in public. For the record, I'm a computer scientist, investment banker, startup founder, bar owner, café owner and startup investor.

None of that is the story. The story is the people who tell.

Part I

The last mile of intelligence

Picture two ordinary afternoons.

Saturday, 2:15 p.m. In a building that runs some of the smartest models on earth, a battery alarm goes off in the middle of a maintenance window. The contract technician on site has the procedure open to page 14. It's clear enough. What he doesn't know is that last spring, when the same alarm went off during the same kind of work, the site did something different, for a reason that made sense at the time. It's in a shift log he has never seen. The senior facilities engineer who wrote that entry is at his kid's soccer game.

Friday, 4:40 p.m. At a grocery distribution center, an HVAC tech is looking at a dead freezer compressor and a building full of frozen inventory that would very much like a decision. His shop was bought by a larger platform in the spring. The old owner would have rented a reefer trailer on the spot, saved the stock, and sorted out the bill later. The new platform's price book says quote a replacement. The senior technician who'd know which way to go is on another job, and he retires in March.

Neither scene is exotic. Versions of them play out thousands of times a day, in plants and fleets and clinics and warehouses. In both, the knowledge needed to make a good call exists. It just isn't standing where the call is being made.

Everything leading up to those two moments is getting cheaper by the month: the alarm reading, the diagnosis, the manual, the price book. The last step isn't. Deciding what to do here, now, the way this business would decide it, depends on facts no model has been given. Last spring's shift log. The old owner's habit. The customer who pays late but sends ten referrals a year.

I call that stretch the last mile of intelligence: the distance between what a model knows and what this business needs done right now.

The idea · The last mile

Knowledge has to reach the person making the call.

The modelWhat the world knowsHere, todayThe facts this business suppliesThe judgment callA person with authorityThe last mile of intelligence
The modelWhat the world knows
↓ Here, todayThe facts this business supplies
↓ The judgment callA person with the authority to decide
General knowledge, local context and decision rights meet at the point of work.

At the end of it there's a gate, where someone with the authority and the know-how has to make the call: the decision nobody else is cleared to make. Follow page 14 or not. Rent the trailer or quote the replacement. In most businesses that gate opens for the owner and a handful of senior people. Everyone else waits, or guesses. This essay is about who else could be trusted to open it, and how they earn that trust.

Part II

AI the frontline will actually use

If the last mile of intelligence is where judgment calls are made, you'd expect AI to be crowding toward it. It isn't. The tools went where the screens were. The work stayed in the truck, on the line, at the clipboard by the door. The people doing it make dozens of decisions a day with nobody to ask, and they're the ones AI has reached least.

Ask is for desks

The AI interface most people know works the same way. You sit at a screen, you type a question, a model answers. Call it Ask. It's a wonderful interface if your job happens at a desk, and it has spread through desk jobs at remarkable speed.

It has not spread to everyone else. In Gallup's surveys of US workers, AI use in jobs that can be done remotely went from 28% in mid-2023 to 66% at the end of 2025. In jobs that can't, it went from 15% to 32% (Gallup, 2026). The gap didn't close. It more than doubled, from 13 points to 34.1 In healthcare support, manufacturing production and retail sales, somewhere between 70% and more than 80% of workers say they never use AI in their jobs at all (Gallup, 2026).

The evidence · Adoption

The gap is getting wider.

The gap between job types

13 → 34percentage points
AI use at work, including occasional use. Remote-capable jobs: 28% in Q2 2023, 66% in Q4 2025. Non-remote-capable jobs: 15% to 32%. These categories do not map exactly to desk and frontline work. Source ↗

Businesses show the same split. In the Census Bureau's May 2026 analysis, about one in five US companies used AI. Among firms with 250 or more employees it was 37%; among firms with fewer than 20, use hadn't meaningfully moved over the preceding six months (US Census Bureau, 2026). That matters, because most of the world's work isn't done at desks. One widely cited estimate puts deskless workers at about 80% of the global workforce (Emergence Capital, 2020).

It's easy to read that gap as reluctance to adopt AI. I think the interface has a lot to answer for. Ask makes the work come to the AI: find a screen, find the time, find the right words. Frontline teams and small businesses are being asked to bend their day around a tool built for someone else's. Fewer than half of US employees say their employer has built AI into how they work at all (Gallup, 2026). A chatbot waits for the user to come to it. AI that's going to work on the frontline has to do the opposite, and meet people where the work already happens.

And the frontline talks all day. Much of what it says isn't a request for an answer. It's a Tell.

Every business runs on Tells

Every chatbot runs on Tells. You just never see them. Before you type a word into ChatGPT or Claude, the model has already been told a great deal. Who it is. How to talk. What it should and shouldn't do, how to format a list, when to say no. Pages of instructions sit in front of every conversation, written by the people who built it. The labs worked out early that good behavior from a model starts with telling it, in detail, how things work around here. Every Ask rides on a stack of Tells.

Businesses run the same way, at every level. A CEO doesn't steer the company by typing questions into a box. She steers it on what she's told: the Monday numbers, the board deck, the chief of staff's "you need to hear this before your 10 am." A plant manager runs on the shift log. A technician runs on the dispatcher's note. Tells are how work talks, from the night shift to the boardroom. And almost none of them make it into a model's training data. They're local, they're about today, and by next week they're buried in last week's report. Instructions tell a model how to behave. Reports tell it what happened. Both matter, but a report isn't permission to act.

Once a conversation starts, though, a chatbot does something odd. Built to answer, it hears nearly everything as a question. Describe a situation and you get a solution, often to a problem you didn't have.

We've all been on one side of this, or both. Picture my wife and me at the end of a long workday. She says, "The train was packed again this morning." I get into ChatAI husband mode and start fixing: leave earlier, take the bus, should we look at apartments closer to work? She wasn't asking for a plan. She was telling me about her morning. I needed to listen first: was this a bad morning, or a pattern we needed to do something about? A lot of arguments with your favorite ChatAI could be avoided if it just listened first.

Every Ask rides on a stack of Tells.

Work is the same. Say the operator at an anaerobic digestion plant writes, "Digester 2 ran a degree cool this morning." That isn't necessarily a problem. It's an observation. A degree cool on three Monday mornings in a row, each one after a weekend feedstock delivery, is a pattern, and that one might matter. You can't tell which you're looking at from a single message. Jump on the first one and you've invented a problem. Ignore all of them and you miss the real one. Capture the Tells over time, and the pattern can show up or not.

That's the job: listen first, answer second. A good manager does it without thinking. An AI built only to answer has to learn the habit.

A Tell is anything a business already records about what happened: the form someone fills out at close, the spreadsheet that tracks the freezer temperature, the weekly update to the owner, the photo of a gauge sent to the plant manager. Most valuable of all are the questions people bring to the boss, because each one marks a spot where judgment was needed and wasn't there.

Tells are not chatter. The team group chat, with its memes and shift swaps and "who took my charger," is mostly noise. The signal is in the routines: the same report, the same count, the same handover, filed day after day, so patterns can show up and exceptions can stand out.

Side by side, the difference looks like this.

Ask Tell
Who starts it Someone seeking an answer Someone doing their job
What it takes Open the app, find the words The report, count or photo they already send
What the AI learns What the user thinks to ask about What happens, day after day
Who it serves Whoever brings it a question Everyone reporting the work, from the night shift to the CEO

If you remember one line from this essay, make it this one. Ask is how you use AI. Tell is how AI learns your business.

It also suggests why workplace AI tools can struggle to become a habit. Asking a busy person to fit another tool into the day is a hard bet. Reading the habits they already have is a better one. Tells are already happening, on WhatsApp, on paper, in email. Start listening there.

Without Tells, an AI has no context about the business, no patterns to learn from, and no way to know when a person is actually needed. With them, it can do something more useful than answer questions. It can notice.

Part III

Judgment is the scarce part

Once an AI can hear what's happening, the next question is what to do about it. The people building agents and the people funding them keep returning to the same answer, which almost never happens. Intelligence isn't short. What's short is the person who knows which of ten good answers this business would pick.

Execution is cheap

Olivia Moore at a16z ran a neat experiment: she handed a fresh social media account to an autonomous agent and told it to grow as fast as possible. The agent nailed the mechanics, from posting cadence to copy to charts. It fell flat on the one thing that mattered, which was having something worth saying (Moore, 2026). Her summary: AI can run a playbook. It can't write one.

Her colleague Yoko Li makes the point through agent loops, the generate-check-revise cycles that power today's agents. "Done," she argues, is a judgment made by the system around the work: the test, the reviewer, the budget (Li, 2026). Take the people out and someone still has to decide what the owner would have signed off on.

Matan-Paul Shetrit, who builds agent products for a living, put it most bluntly. AI doesn't reduce work, he wrote. It refactors it. People move from doing tasks to authorizing them and absorbing the consequences. Fewer decisions land on their desks, and each one weighs more (Shetrit, 2025). Anyone who has run a business knows that trade. The owner who stops doing the books doesn't stop worrying about cash. She just worries about it with less to go on.

Even the most bullish voices concede the gap. Sequoia's partners called long-horizon agents "functionally AGI" and declared 2026 their year, while admitting those same agents can "charge confidently down exactly the wrong path" (Grady and Huang, 2026). I've hired that person. Most founders have. And Sequoia's Julien Bek drew the cleanest line of all: intelligence work follows rules, however complicated; judgment work runs on experience and instinct (Bek, 2026).

They agree on the bottleneck. What happens to it next is less settled. Bek expects judgment to become intelligence over time, as models learn from enough examples. Shetrit worries it will pile up with a few people, hidden inside software where nobody can see it. Part IV takes both seriously.

I'm not betting against the models. They're getting better quickly, and I use them every day. The argument is about what they can't see, not what they can't think. A brilliant new hire on their first day can't see last spring's shift log either.

Same tool, opposite results

Handing everyone the same smart tool doesn't guarantee that everyone gains.

Rembrand Koning at Harvard Business School and his colleagues ran an experiment with 640 small-business owners in Kenya. One group got an AI business advisor, built on GPT-4 and delivered over WhatsApp; the other got standard written business guides. There was no statistically detectable average effect. Underneath that average, it pulled the group apart. Owners who were already doing well saw profits and revenue rise by roughly 10 to 15%. Owners who were struggling did about 8% worse (HBS, 2025).

The researchers' account suggests why. Strong owners picked advice that fit their business, down to which breed of chicken to raise, and used it well. Struggling owners picked moves like cutting prices or buying ads, which cost money and often backfired. Same advisor. Judgment looks like part of the difference: telling good advice for this business from good advice in general.

Now look at a study that went the other way. Erik Brynjolfsson, Danielle Li and Lindsey Raymond followed 5,179 customer support agents as an AI assistant rolled out. Productivity, measured as issues resolved per hour, rose about 14% on average, and about 34% for novice and lower-skilled agents, with little change for the veterans (Brynjolfsson, Li and Raymond, 2023). The authors found signs that the assistant was spreading the practices of the company's best performers. The newer agents weren't left to sort business advice on their own. They had help with the specific problems that business saw.2

The evidence · Two studies

Two studies. Two very different patterns.

Kenya · 640 business owners
AI business advice

Change in profits and revenue

Kenya: stronger owners gained 10–15%; struggling owners declined about 8%Stronger owners+10–15%Struggling owners−8%−10%0+10%

No statistically detectable average effect.

HBS, 2025 ↗
Support · 5,179 agents
AI at the point of work

Change in issues resolved per hour

All agents≈ +14%
Novice & lower-skilled agents≈ +34%

Little change for experienced agents.

NBER working paper, 2023 ↗
The studies involve different jobs, interventions and outcome measures. They suggest a question about how expertise travels; they do not isolate the effect of precedent. Kenya’s positive estimate is shown as a range. The support study uses a separate 0–40% bar scale.

Put the two studies side by side and a useful question emerges. What would happen if more AI systems passed along a business's best judgment, instead of leaving each user to sort the advice alone?

Every business that tries to grow runs into this. Think of a café chain going from three locations to thirty. Location three is run by the founder's best manager. Location thirty is run by someone hired last month. Give location thirty a chatbot and the new manager still has to choose which answer fits. Give that manager the calls the founder and the best manager have already made, in situations like the one in front of them, and there's less to figure out from scratch.

That second thing has a name, and it's old. It's called precedent. Part V comes back to it.

The frontier never closes

Models learn the world as it was. Businesses run on the world as it is at 4:15 this afternoon.

There are two kinds of knowledge the business has to supply. The first is temporal: it just happened. The truck is late, the reading moved, the customer phoned back. The second is local: it's true here and nowhere else. This freezer is older than the others. This client hates surprises. This valve sticks in the cold. A model can use those facts. It won't discover them in its general training.

To see both at full volume, look at a stadium tour. A tour at the scale of the Eras Tour, 149 shows on five continents over 21 months (Pollstar, via NBC), rebuilds a small city every few nights: new venue, new local crew, new promoter, new weather, new rules.

Picture a night in Lisbon. It's 4:15 p.m. The promoter has just moved doors up 90 minutes. The local crew is six people short. The production manager has to decide which parts of the build to drop without breaking the show sixty thousand people paid for. None of that was in anybody's training data last week. Most of it wasn't true this morning. And the people who know how the last forty nights went are spread across three buses and a hotel.

Julien Bek's line that "today's judgement will become tomorrow's intelligence" is right about the direction. Calls that get made often enough, with enough recorded outcomes, do turn into rules a machine can follow. Everything that follows depends on it. For an operator, though, the next frontier matters as much as the one just crossed. The world keeps printing new situations. Every new site, season, customer and regulation brings a fresh frontier with it. The last mile isn't a gap that closes. It moves, every day, with the business.

And even where it could close, there's a reason to keep a person at the gate that has nothing to do with what models can or can't do. Owners want to own their decisions. Customers, staff and regulators want a name attached when something goes wrong. Plenty of calls a model could make are still calls a business should make, until it has good reason to hand them over. Keeping people in the loop shouldn't be an apology for a weak agent. It should be a deliberate choice about who decides what, and when that changes.

Part IV

Three futures

So judgment is scarce, and people hold it. That raises an awkward question for anyone selling AI: if the machines keep getting stronger, why keep paying people to stand at the gate? The first answer is about the person already on shift. The second is about who ends up holding judgment once AI is everywhere.

A pair of hands

The case for replacement is easy to see: cheaper decisions, more consistency, less dependence on the one expert who always gets the call. Capture the judgment once and let the machine use it all day.

But the person on shift brings more than a decision. They hold local facts, they know what happened this morning, and they come with hands. They can open the panel, phone the customer, walk the line. No robot required.

The agent still needs what the production manager in Lisbon saw ten minutes ago, and the means to act on it. When is supplying both cheaper than helping the person already there make a better call?

People have limits too, and the usual one is overload. Hand someone every fact at once and the call gets worse, not better. That's where the model earns its keep: sorting the facts by what the business has decided before, and putting only what matters in front of the person who has to decide. For the person already on shift, that pairing is hard to beat.

Automated, concentrated, amplified

If judgment is scarce, the big question is what happens to it as AI spreads. I see three futures. They're less predictions than choices, and businesses are making them right now, mostly without noticing. A business can combine all three. The choice is who gets to decide, and who can see and change the rules.

Automated. This is Bek's path. Judgment is captured in data, turned into rules, and absorbed into models, until the machine makes the call. For a lot of routine decisions, that's fine, even good. The risk is what happens to the business's own reasoning along the way. If the call disappears inside a vendor's model, beyond the business's ability to read it, question it or change it, it has rented its judgment back from someone else.

Concentrated. This is Shetrit's warning, and it's the one that worries me. As software takes over the doing, judgment doesn't spread out. It funnels up to fewer people, and it hides, buried in routing rules, confidence thresholds and prompt settings nobody on the floor ever sees. People keep their tasks but quietly lose their decision rights. Shetrit calls it judgment capture (Shetrit, 2025). The risk is a business where a few people decide everything, out of sight, and everyone else stops deciding at all, which is the last thing a growing business needs.

Amplified. This is the argument here, and it's built as an answer to that risk. Judgment is captured from the people who own it, written down in plain words anyone can read, and sent out to the front line, with clear rules about who can change it. The rules about who gets to decide are visible. Every call has a reason and a named person responsible for the decision or the rule behind it. And the people on the front line get more decisions over time, not fewer.

Bek is right about the direction judgment travels. The open question is where it ends up. I think it belongs inside the business, where its owners can read it, not inside someone else's model.

Part V

Amplified judgment

Everything so far points at one design, and it isn't a smarter model. It's a way of getting each call to the person who should make it, and making sure the business keeps what it learned when they do. Think of it less as a brain and more as a good dispatcher with a long memory.

A router, not an oracle

The model is called the Amplified Judgment Action Model, or AJAM. The ideas don't depend on any product. Anyone could build it.

The core idea is simple to state. The agent's job isn't to know everything. It's to run the business's playbook all day, handle what's already been decided, and get the real judgment calls to the right person before it's too late. A router, not an oracle.

It works as a loop with five steps.

The model · Five steps

A router, with a human at the gate.

The AJAM decision loop Five steps: Notice branches to Act for a covered call, or through Route, a human decision and Set precedent for a judgment call. The outcome is checked and delegation approved before the precedent informs the next call. NoticeRead the Tells ActRun approved rules RouteFacts to the right person DecideA person owns the call Set precedentCall, reason and owner Already covered Needs judgment Check the outcome · Approve who can use it next The next call, after approval
1 · NoticeRead the Tells
Already covered
2 · ActRun approved rules
Needs judgment
3 · RouteFacts to the right person
4 · DecideA person makes the call and owns it
5 · Set precedentRecord the call, reason and owner.
Check the outcome; approve who can use it next.
↩ Approved precedent informs the next call
The five steps, in detail
  1. Notice. The agent reads the Tells: the reports, counts, handovers, photos and questions the business already produces.
  2. Act. Anything the playbook already covers, it handles, the same way every time.
  3. Route. Anything that needs judgment goes to the person who should make the call, with the facts they need and nothing they don't.
  4. Decide. A person makes the call and owns it. This is the step the whole model is built to protect.
  5. Set precedent. The call, the reason behind it and who made it are recorded. Its outcome is checked, and an approver decides who can use it next time.
The AJAM loop keeps a person responsible for the judgment call. A recorded decision becomes usable precedent after outcome review and approval.

Go back to the freezer on Friday afternoon. The tech sends a photo of the compressor and a two-line note. The agent knows what's on his truck, what the service contract says, which reefer trailers are available nearby, and, because the old owner's calls were captured in the first month, how this shop handled the same problem before it was acquired. It doesn't make the final call. It routes the question to the owner with all of that attached. The owner adds something nobody wrote down: this customer pays late but sends ten referrals a year. He decides in a minute. Rent the trailer at cost. That decision, and why, goes into the record. Once the outcome has been checked and the owner approves the delegation, the lead tech can make that kind of call himself.

AJAMDecision logReview pending

Freezer 3 · Compressor failure

Friday · Grocery distribution center
2 min16:40–16:42
  1. “Compressor’s dead on freezer 3. Box is climbing.” Photo attached.

    TechTell
  2. No compressor on the truck. Service contract and nearby reefer trailers pulled up.

    AgentNotice
  3. Customer told a decision is coming within the hour, under the shop's communication policy.

    AgentAct
  4. Routed to the owner with the photo, contract, trailer options and how the old shop handled it.

    AgentRoute
  5. Rent the trailer at cost. Quote the compressor Monday.

    OwnerDecide
  6. Saved for review: this customer, freezer down on Friday afternoon, trailer first. Outcome and delegation approval to follow.

    PlaybookSet precedent
Decision log · Friday, grocery distribution center · Illustration

One owner's judgment call, made in a minute, can become a call the lead tech is trusted to make.

Usman Rabbani at Brighton Park describes a parallel idea for enterprise software: the durable value isn't in the model, but in the harness around it, the code that gathers context, enforces permissions, escalates to people and keeps the record. His rule of thumb is that a model should retrieve facts rather than recall them (Rabbani, 2026). AJAM applies the same rule to judgment. The model doesn't remember how this business makes calls. It looks them up, every time, in a record its owners can read.

Two things keep this honest. The first is that autonomy is earned one kind of decision at a time, never granted wholesale. The second is that the gate visibly moves. If the model is working, you should see fewer escalations to the owner each month, more calls made by junior staff, and a record of how often people overrode a precedent and why. Those changes have to come with outcomes that hold up. Fewer calls to the owner mean little if people have just stopped asking. People stay in the loop on purpose: to keep control of the calls that matter, and to hand them down, deliberately, as they move on to bigger ones.

Precedent: how judgment scales

Courts worked this out centuries ago. Nobody puts the chief justice on parking tickets. Instead, judges write down what they decided and why, and later judges facing a similar case follow it unless there's a good reason not to. That's how common law spread consistent judgment across thousands of courtrooms without cloning the best judges.

Most business decisions aren't right or wrong in the abstract. They reflect which values a business applies. Two excellent tour managers will handle the same exception differently, and both can be right. One protects the crew's rest; the other protects the fans' night out. A model can't know which values you hold. A run of your decisions shows them, plainly, in a way no mission statement ever does.

Precedent doesn't remove the gate. It changes who's qualified to open it.

Every precedent in AJAM starts with four things: the situation, the decision, the value that was weighed, and who made the call. What actually happened gets attached as the result comes in. With enough of them, the agent has something better than a blank page: cases the business has already decided. The ambition is to make its job narrower, matching the case in front of it to that record rather than reasoning from scratch. Researchers will recognize case-based reasoning here, and a close cousin of what Gary Klein called recognition-primed decision making: experienced people mostly decide by recognizing a situation as like one they've seen before (Klein, 1998).

The Kenya study raises a fair question. If AI advice hurt struggling owners in that experiment, why expect precedent to help less experienced people? The reason to try is that precedent gives them this business's calls, on this business's situations. Its usefulness depends on a new case genuinely matching an old one. Anything new, or a match the system can't establish, still goes up to someone who can judge it.

Who signs? Healthcare separates permitted tasks from clinical judgment. An aide needs the authority, competence and supervision to do a delegated task, and remains responsible for it. The nurse retains accountability for overall care and nursing judgment (NCSBN and ANA, 2019). A business can borrow that discipline: widen the work someone handles, and name the responsibility that goes with it. Precedent alone doesn't grant permission.

Precedent can also go wrong, and a system built on it has to plan for that. The wrong precedent can get matched to a case that only looks similar. A good call can go stale when conditions change. A bad call can get recorded and repeated. The fixes are unglamorous: check precedents against what actually happened, retire the ones that stop working, and keep track of how often people override them. If nobody ever overrides a precedent, you don't have judgment. You have a rubber stamp.

Precedent doesn't remove the gate. It changes who's qualified to open it. Experts record the hard calls, and the results let them decide which ones to hand to people who weren't in the room.

Memory with gates

Where something is remembered should depend on what it costs the business to forget it.

That sounds obvious, but memory that's everything in one big pile, or whatever happened to fit in the conversation, won't do the job. AJAM organizes memory in six layers, from the most general to the most specific, each with a gate suited to what it holds.

The memory · Six layers

Different knowledge. Different gates.

1World

What the model already knows

Open to everyone
2Profession

How good practitioners work

Open within the industry
3Institution

This business’s approved playbook

Read as needed; approvers change it
4Temporal

Today’s shifts, seasonal quirks and open issues

Scoped by role; expires
5Outcomes

Situation, decision, values, owner and result

Gated by role and team
6Tacit

Knowledge still held by a person

Reached by asking that person
Tacit knowledgeAsk the person who knows
Outcomes recordSave the call and check the result
Institutional ruleAn approver promotes the precedent
What each layer holds, in detail
  1. World. What the model already knows. Open to everyone.
  2. Profession. How good practitioners in this field work: what a well-run digester or service call looks like anywhere. Open within the industry.
  3. Institution. This business's playbook, the rules it has decided on. The people who need it can read it; only approvers can change it.
  4. Temporal. What's true right now: today's shifts, this season's quirks, open issues. Scoped by role, and it expires.
  5. Outcomes. The precedents: situations, decisions, values, who decided and what happened. Gated by role and team.
  6. Tacit. What's still only in people's heads. Never stored directly. It's reached by routing a question to the person who knows.
Learning crosses layers through recorded outcomes and human approval. Access follows the purpose of each layer.

Over time, judgment hardens. Take Gary, the process engineer at that digestion plant. He knows the feed valve on Digester 2 sticks when it's cold out. For years that fact lived only in his head, which meant the night operator found out the hard way every winter. The first time the agent routes a cold-morning question to Gary, his answer goes into the outcomes record. The second and third time the same situation comes up and his answer holds, an approver can make it a playbook rule for winter. Gary is still the expert. He just doesn't get the same phone call every January.

Impact helps decide where something sits and how far it reaches. An approved limit on digester temperature gets written into the playbook and enforced. "The night operator prefers texts to phone calls" stays in temporal context and gets used when it's relevant. Reach matters as much as depth: a precedent that proves itself at one site can be promoted to the master playbook for every site, but only when a person approves it, because what works at one site may flop at another.

Seema Amble at a16z makes a distinction that fits here: memory isn't learning. Learning comes from corrections and outcomes, and it comes in two kinds, learning the profession and learning the institution (Amble, 2026). Those are layers two and three, kept apart on purpose.

This also changes what a playbook is. In most businesses today, the playbook lives in people's heads and in a folder in the cloud that gets updated once or twice a year, usually after something goes wrong. Under AJAM, the playbook is run every day and updated every time a precedent earns its place.

Part VI

What it's worth

A good idea that fits nowhere is a hobby. Amplified judgment fits some businesses far better than others, and many of them are attracting buyers looking to roll them up. If that's a coincidence, it's an expensive one.

Where it pays

This isn't for every business. It's for businesses where the know-how is in the building but rarely at the wrench when it's needed.

Six questions decide the fit.

  1. Is the work done away from a desk, across shifts, sites or trucks?
  2. Are routines already written down somewhere, even on paper or WhatsApp?
  3. Are there a few experts and many doers?
  4. Do the same kinds of calls keep coming back, never quite identical?
  5. Is a wrong or slow call expensive?
  6. Does knowledge have to move: to new sites, new hires, or the next generation?

The giveaway is when the manual exists and the work still goes wrong: same manuals, very different results. In Uptime's 2025 survey, nearly 40% of data center operators reported a major outage caused by human error over three years. Of those incidents, 85% involved staff not following procedures, or procedures that were flawed (Uptime Institute, 2025). In its 2026 report, most respondents said their last significant outage cost more than $100,000; one in five said more than $1 million (Uptime Institute, 2026). In Aquant's field-service benchmark, the best-performing companies fix the problem on the first visit 86% of the time and the worst manage 53%; a miss adds an average of two more visits and about 14 days (Aquant, via Industrial Machinery Digest). How much of that gap is judgment that didn't travel?3

The evidence · Field service

One visit. Two very different results.

Same benchmark. Different results.

33percentage-point gap
First-time fix rates in Aquant’s 2025 benchmark: 86% for the best-performing group, 53% for the worst. This comparison does not isolate the cause of the gap. Source ↗

By those six questions, the strongest fits look like this.

  • Water and wastewater. A strong fit across all six. About 49,500 community water systems serve the US, and 81% of them serve 3,300 people or fewer (Congressional Research Service, 2026); in 2020, EPA estimated roughly a third of the workforce could retire within ten years (EPA, 2020). Operators already keep daily logs for their permits. Picture a plant the morning after a storm, ammonia creeping up, and the chief operator who's seen this twice before on vacation. A wrong call is a permit violation and a notice in the local paper.
  • Field service. HVAC, elevators, commercial refrigeration. Senior techs are retiring, everyone else lives in a truck, and buyers are circling: PitchBook counted 55 private equity HVAC deals across North America and Europe in 2024, up 72% on the year before (PitchBook, 2025).
  • Food and beverage processing. The US had 42,708 food and beverage plants in 2022 (USDA ERS), with sanitation and changeover routines that depend on a few quality leads. A pallet with last month's lot code is a small decision with a large downside.
  • Data center operations. Strong evidence that failure can live in routines. The fit is best where a few experienced engineers support many sites or shifts. The shelf is crowded with monitoring software, but the alarms are covered and the call isn't.
  • Home care and senior living. The 2024 Activated Insights report put caregiver turnover in home care near 79% (Activated Insights, via Home Health Care News), and Ziegler's 2026 survey, mainly of not-for-profit providers, put senior-living turnover near 35% (Ziegler, 2026). The agent keeps what the last aide learned about a client. A nurse still makes every clinical call.
  • Veterinary groups. Consolidators bought clinics from dozens of owners with dozens of habits. Start with the operating routines; clinical judgment stays with the vets.

Then there are the ones nobody thinks of. Funeral homes, where regional buyers are rolling up family firms whose know-how sits with one director. Golf course maintenance, where a June 2025 report described one private equity-backed operator running 70 courses after more than 18 acquisitions in three years (Front Office Sports, 2025). Quarries, dairy farms, airport ground crews.

It doesn't fit every decision. Aviation maintenance still requires approved data and the proper sign-off. Chairside dentistry usually puts the expert at the point of work (the front desk is fair game). Where existing systems cover the call, or it's cheap to get wrong, there's less reason to add this. Clinical and other regulated calls keep their gates. Wherever a wrong answer could cost a life, authority stays with qualified people and their safety systems, full stop. The fit is a class of decisions, not a whole industry's license to delegate.

The playbook premium

Most small businesses sell for what's on the books. The playbook in the owner's head walks out the door at closing.

The price gap is wide. Main Street businesses that sold in 2025 went for a median of $350,000 and an average of 2.61 times their cash flow (BizBuySell, 2026). Private equity deals in the first half of 2025 averaged 7.2 times EBITDA, and 10 times for deals between $100 million and $250 million (GF Data, via ACG). Those are different earnings measures and very different businesses, so don't read them as an exact spread or a measured playbook premium. Buyers do pay for businesses they can grow. In HVAC, service-heavy companies with recurring revenue, strong customer retention and regional density often fetch 10 to 12 times EBITDA or more, while undifferentiated contractors trade lower "unless bundled into a larger platform" (PitchBook, 2025).

And everyone's buying. Add-on acquisitions, where a platform company buys smaller ones and bolts them on, made up 72.9% of US private equity buyout deal count in 2025 (PitchBook, via Cherry Bekaert, 2026). The sellers are lining up too: 51% of US employer-business owners were 55 or older in the Census Bureau's 2018 data (US Census Bureau).

What surprised me is where the money came from in Bain's study of 44 buy-and-build deals. It found that the ones relying on buying small and selling big returned 1.4 times the money invested, while the ones that actually grew the businesses and improved margins returned 2.2 times (Bain, 2024). The same 2024 report describes Caliber Collision, then with more than 1,700 shops, sending integration and training teams into the shops it buys. That team is a playbook on legs.

The evidence · Buy-and-build

The return from running it better.

44 deals. Two sources of return.

1.4 → 2.2times invested capital
Bain’s study of 44 buy-and-build deals: 1.4× invested capital for deals relying on multiple arbitrage; 2.2× for those focused on growth and margin improvement. These are investment returns, not acquisition valuations. Source ↗

Roll-ups also break, and how they break is instructive. Some veterinary consolidators are under real strain: lenders have marked loans below face value, and Octus links the pressure partly to staffing and retention problems (Octus, 2026). ATI Physical Therapy faced litigation over its disclosures about therapist attrition; its 2024 annual report records a $26.5 million aggregate settlement (ATI, 2024 annual report). Buying a business doesn't make its people stay. What happens to the playbook if the experts it depends on walk out? Even franchises, the original written-down playbook, survive their early years better than independents by only five or six points, and the edge fades (University of Michigan, 2018).

AJAM takes a different route. It builds the playbook while it runs it. There's no six-month binder project with consultants. The routines the business already files, and the calls its experts already make, become the playbook. My bet is that a playbook written by the people who know the work is one they're more likely to trust, and stay to run. Most playbooks get written, then run. This one gets written by running.

That's the difference I'm after: a roll-up that works like a team. An elite sports team has a coaching staff and analysts who help the players run the right play at the right moment, while the players still make the call on the field. A Fortune 500 company can afford that kind of staff. A ten-person company never could. Now it can have its playbook run around the clock.

Speed matters here too, but only one kind. Call it decision velocity: how fast a signal on the ground turns into the right action, and how many good calls a business makes in a day. Fast is worthless if every site applies a different rule, so a fast decision counts only when it passes the test from the first page: same call, any shift. Or any store, or any truck. A precedent that proves itself at one site gets promoted, by a person, to every site, and consistency becomes something you can actually see.

Most playbooks get written, then run. This one gets written by running.

Which leaves a question I can't answer for you yet. The figures above don't tell us what a working, written-down playbook adds to the price of a business. So consider it yourself. If a buyer could see every call your business made, who made it and why, what would that be worth?

Part VII

Where to start, and what I'm betting

Diagnosis is cheap. Here's the part you can act on next week with a notepad, whether you run a business, build AI or fund it. Then five bets with dates on them, so you can check later whether I was right or just confident.

Where to start

If you run the business. Write down the ten judgment calls only you make. Not the tasks, the decisions: the refund you'll approve, the order you'll bend, the customer you'll make an exception for, and why. That list is your first playbook, and every item on it is a precedent waiting to happen. Then look at what your team already files every day. The nightly report, the handover, the count. Those are your Tells. You don't need a new app to start listening to them.

If you build AI products. Stop chasing the agent that does everything. Build the one that knows when to ask, asks the right person, and remembers the answer. Build it on what people already send, not on a new habit you hope they'll form. Then track the three things from part V, and check that the outcomes hold up as the gate moves.

If you invest. In industries everyone says are hard to scale, the companies worth backing are the ones that turn their experts' decisions into precedent. Look at how a company captures what its experts know, how fast that knowledge reaches every site, and whether decision velocity holds up as it spreads. Ask how many Tells each worker sends in a week, in month one versus month six, and how many arrive without anyone asking for them. And when you look at a roll-up, ask who wrote the playbook. If the answer is a consultant, ask how many of the experts are still there.

Bets, with dates

An essay that can't be wrong isn't worth much, so here are five bets. Each one has a date and a way to check it. I'll report back on all of them, in public, including the ones I lose.

1

More than 50% of US workers in non-remote-capable jobs use AI at least occasionally at work, up from 32% at the end of 2025

Check by
How we'll knowGallup's workforce survey
2

More than 25% of US firms with fewer than 50 employees report using AI in at least one business function

Check by
How we'll knowCensus Business Trends and Outlook Survey: share of firms with 1–49 employees reporting AI use in any business function in the previous two weeks
3

At least three agent companies selling into field operations, each with more than $50 million in revenue in its latest completed fiscal year, disclose a delegation rate: within a defined class of calls previously handled by experts, the share now handled by frontline staff and the share handled by agents, reported separately

Check by
How we'll knowEarnings releases, filings, investor letters
4

At least one large field-service, care or utility operator publicly lets an agent make a defined class of routine calls without per-case approval, under a named accountable owner, because its own precedent showed it was safe

Check by
How we'll knowPublic announcement or filing
5

An acquirer cites a target's decision record, meaning who decided what and why, as part of the reason for the price it paid

Check by
How we'll knowDeal announcement, earnings call or investor letter

Bet 3 is a wager that delegation rate becomes the number boards ask about, the way they ask about churn today. Bet 4 is the gate moving, not vanishing. The first two bets track adoption; the others ask whether businesses change how they delegate and what they value.

Saturday, again

It's 2:15 in the afternoon, and the battery alarm goes off again. This time the technician knows what the site decided last spring, and why. At the grocery warehouse, the HVAC tech already knows the call his new owners would make, because they wrote it down, checked the outcome and authorized him to use it. Nobody had to disturb anyone at a soccer game. Right call, any afternoon, any truck.

Credits

Reading these five writers pushed me to build this, sometimes because I agreed with their ideas and sometimes because I didn't. They're listed in the order I'd want to argue about those ideas with them over dinner. Others sharpened it along the way, and they're in the sources.

  • Matan-Paul Shetrit, on judgment as the final frontier and on where it concentrates (1, 2)
  • Julien Bek, on intelligence work versus judgment work (Sequoia)
  • Olivia Moore, on agents that can run a playbook but not write one (a16z)
  • Rembrand Koning and colleagues, for the evidence that the same AI advisor can help some operators and hurt others (HBS)
  • Usman Rabbani, on an architecture for enterprise software built around context, permissions and the human role (Brighton Park)

Sources

Source notes

  1. Gallup, Q2 2023 and Q4 2025. Use includes occasional use. Remote capability is not an exact dividing line between desk and frontline jobs. The occupational non-use figures come from the separate October 2026 article. ↩
  2. Kenya figures follow the HBS September 2025 summary; its account of advice selection is exploratory. Support-agent figures follow the April 2023 working paper, also preserved by Stanford SIEPR. Different people, tasks and interventions mean these studies do not isolate precedent as the cause of their different results. ↩
  3. Aquant's 2025 benchmark compares groups within its dataset. It does not separate judgment from differences in staffing, parts or other operations. ↩