top of page

Your AI Business Case's Blind Spot

Sep 10
11 min read

Car side mirror with the blind spot warning indicator illuminated, road visible behind.

I have worked on hundreds of business cases. Every one of them was written inside a department or a division, and written for the benefit of that department or division. That is not a criticism. It is what a business case is for.


The only one I ever saw that mentioned disaster recovery was a business case about disaster recovery.


That was fine for thirty years. It is not fine now, and the reason has nothing to do with the quality of the cases or the people writing them.


The business case and the continuity plan never meet

Two owners, one assumption, no shared document

A business unit builds a case because it wants to be better at the thing it does. Faster claims processing, cleaner forecasts, fewer touches per order. The case compares the world with the project against the world without it, over three or five years, and both of those worlds are assumed to keep operating.


Continuity belongs to somebody else. IT owns it, writes the plan, tests it on its own cycle, and reports on it to a different committee.


Between those two functions sits an assumption that has never needed to be written down, because it has never been wrong.


IT could assume the business had the people.


The continuity plans I have read all have a section saying the work will be done manually and reconciled when service returns. Underneath that section sits an assumption: somewhere in the business are people who know how to do the work without the system. Not named, not budgeted, not tested. Present.


That assumption survived mainframes, client-server, ERP, the web, mobile and the cloud. Every one of those waves changed how the work was done and left behind the people who knew how to do it. The process moved. The knowledge stayed in the building.


The first wave that can consume the fallback

Earlier automation moved the process and left the practitioners

Earlier waves removed people. Mainframes removed clerks, ERP removed accounting staff, self-service removed transaction processors. The distinction is not headcount.


What those waves left behind was enough human capability to reconstruct the process or run it another way. The people who understood the work were still in the building, and parts of the job were still done by hand, so the understanding stayed current without anybody maintaining it.


An AI deployment is funded on the basis that it absorbs volume. That is the case. It is usually a good case, and the savings usually land.


What lands with the savings is the headcount change that made the numbers work. Some of those people leave the company. Others are redeployed, which sounds gentler and often is, though it depends entirely on how completely and for how long they leave the work.


The relevant question is not whether they are still on the payroll. It is whether anyone is still doing the thing by hand, and how recently.


The continuity plan does not change. It goes unedited, because the room is not thinking about it and the person who owns it was not invited.


The gap is not a missing number. It is a seam between two owners.


And a second seam opens after the deployment goes live. The business has reduced its human fallback and increased its dependence on a provider it does not control. Two dependencies moving in opposite directions, neither of them recorded in the document that authorized the move.


The case is scoped to a division and scored on that division's outcome. The cost lands in a function that was not in the room, against a plan written before the deployment existed. There is no meeting where those two things are netted against each other, and no form where the trade would be recorded if somebody wanted to record it.


Nobody skipped a step. No step exists.


Stated as a CFO would state it: the case captures the labor avoided. It does not capture the replacement cost of the capability being consumed, because that capability has never appeared on any ledger.


What the deployment does to the people who stay

What thirty years of research says about reliable systems

The obvious answer at this point is that the people are still there, just supervising instead of doing. A human stays in the loop, the process gets reviewed, and the capability is preserved.


Two people wrote well about that this summer and it is worth reading both (available on LinkedIn).

  • Iain Brown argued in July that a person appearing on a process map is not the same as a person who can change what happens next, and that oversight has to be designed around whether intervention is actually possible.

  • Mo Saker wrote in August about being the human in the loop himself, with an approval queue of forty-seven items at 11:38 at night, a median review time of eleven seconds, and one bad approval that cost about ninety-six thousand dollars.


Neither of them cites the research, and the research is thirty years old.


Reliability degrades the reviewer. Parasuraman, Molloy and Singh ran a multitask flight simulation in 1993. Participants monitored an automated system while handling other jobs. For one group the automation was consistently reliable. For the other it varied. After about twenty minutes, the group watching the reliable automation was substantially worse at catching its failures.


The reliability itself was what degraded the monitor. And the effect disappears when monitoring is the only job, which matters because enterprise reviewers rarely have monitoring as their only job.


Rarity degrades detection. Wolfe, Horowitz and Kenner published in Nature in 2005. Observers searched simulated baggage x-rays, and the only thing that changed between conditions was how often a target was actually there. At fifty percent prevalence they missed seven percent. At one percent they missed thirty. Changing nothing but frequency produced a fourfold increase in errors.


The mechanism shows up in the reaction times. At one percent prevalence, observers were answering "nothing here" faster than they were answering "found it." They were abandoning the search in less time than it took, on average, to find the thing.


Think of it as Where's Waldo, except Waldo is only on one page in a hundred. You would find him on page one. By page forty you would be flipping. By page ninety-nine you have stopped looking in any way that counts, and when he finally shows up you turn straight past him.


Nobody decides to stop looking. The absence of a decision is the whole problem.


Wolfe tested the obvious fix, which is to salt the queue with additional targets so the reviewer finds things more often. It did not protect the rare target, whose miss rate went up rather than down.


He is also honest about the limit, and it belongs here. He writes that because the experiments are burdensome, it is not clear whether the effects occur in the field. Twenty years later that is still the honest position.


Reliability and rarity meet in the same place

Two conditions, and a successful deployment creates both

The two research traditions rarely cite each other, and they describe the same operational danger from opposite ends. High automation reliability and low error prevalence create the same condition for the person watching: the thing they are looking for becomes rare. Ninety-nine percent accuracy, from the reviewer's side of the desk, is an environment where an error shows up once in a hundred.


So two conditions govern whether a human in the loop is a control or a formality.


Errors frequent

Errors rare

Reviewer does nothing else

Works. Frequency keeps them calibrated, attention is undivided

Prevalence pulls detection down, though single-task monitors held up regardless of reliability

Reviewer has a queue and a day job

Degraded and still functioning. Finding things regularly keeps them in the task

Both effects compound. The quitting threshold falls and attention goes to the work that visibly needs it


This is a way of locating the problem rather than a model of it. Bottom right is where many enterprise review roles end up, and it is where a successful deployment can put you.


The deployment moves you right, because the errors the reviewer sees become rarer as the model improves. The same deployment moves you down, because the headcount that justified it is gone and the people left absorb other work. Both movements are consequences of the project delivering what it promised.


You do not pick that cell. You arrive in it, and the arrival looks like a win in every report written about it.


Reviewing is not doing

Recognition and construction are different skills

The loop is assumed to do a second thing, and it matters more to the seam than the first does.


Keeping a person in the loop is supposed to keep the capability alive. If the model goes away, the reviewer can do the work.


Checking a proposed answer is recognition. Building the answer from source material is construction. From outside they look adjacent, and they train different things.


A senior property accountant, twenty years in, can look at a rent roll and tell you within about four seconds that something is wrong before she can tell you what. The number does not sit right against a building she knows, in a market she knows, at a time of year she knows, and the reasons surface afterwards.


That did not come from approving rent rolls. It came from building them, being wrong, and being corrected, over and over.


Hire her replacement into a review role and they never acquire it. Keep her in one for three years and it fades, because the thing that maintained it was the doing.


So the continuity assumption is not just weakened by the people who left. It is weakened by the people who stayed.


Two clocks start the day you approve, and the second one has a number

The fallback capacity ratio, in three numbers you already own

One of them is the payback period. Everybody watches it. It appears in the case, in the post-implementation review, in the slide that says the project delivered.


The other one measures how long until the fallback is gone. Nobody computes it, and it can run faster than people assume, because skill can decay faster than headcount disappears.


You can read the second clock in an afternoon. Three numbers, and you already own all of them.

  • Daily Volume. Units per day through the process, whatever the unit is.

  • Manual Throughput. Units one trained person completed per day before the deployment. Somebody has this, because it was in the case that funded the project.

  • Capable Headcount. People who have actually done the work by hand. Not the org chart for that function. Names, and when each of them last did it.


Multiply the second by the third to get manual daily capacity. Divide that by daily volume. Call the result the fallback capacity ratio.


At 1.0 the fallback keeps pace. Above it, the work gets absorbed. Below it, the fallback cannot keep pace with incoming volume, so backlog grows for every day the model is unavailable. Materially below 1.0, the continuity plan describes a state the organization cannot reach.


Then run it again with an honest view of how good those people are now, three years into reviewing rather than doing. That second answer is the one to take into the next approval.


The ratio is not a forecast and it does not need to be precise. It needs to be spoken out loud in the room where the case is approved, because right now no number occupies that seat at all.


Then the next case arrives

This is where the seam stops being a one-time problem.


The next deployment gets proposed against the fallback as it stands now, not as it stood before the last one. The manual alternative has quietly become worse, which makes automating further look less risky than it did, because the thing you would fall back to has degraded since anyone last looked at it.


Each case is approved on the evidence available in the room. Each one lowers the bar for the next. And the direction is not symmetric: capability is rebuilt by doing the work repeatedly and being corrected, which takes years and requires somebody who still knows how to correct you. That population shrinks every cycle.


The exposure does not stay the same size. It compounds, and every turn of it was somebody's good quarter.


The second dependency, the one you acquired

When a provider disappears for reasons that are not technical

Everything above is about the fallback you spent. The other half of the seam is the dependency you took on in exchange for it.


The scenario worth planning against is not a model producing a wrong answer. That is what the reviewer is for, degraded or not.


It is the model being unavailable while the work still has to happen.


On June 9, 2026 Anthropic released its two most capable models. Three days later the US Department of Commerce issued an export control directive requiring the company to restrict access by any foreign national, anywhere. Anthropic said it had no reliable way to verify nationality in real time, so it suspended access to both models for all users. Access was restored on July 1st, after the controls were lifted.


Eighteen days without the capability, and it began three days after customers were invited to build on it.


Nothing broke. No outage, no bug, nothing to detect. A regulator acted, a provider complied, and the capability was gone.


The standard advice for this is to abstract behind a gateway and keep a second model warm. That is sound and it does not help here. No abstraction routes around a jurisdiction, and a second vendor inside the same one fails on the same morning.


Continuity planning assumes failures are largely independent. Your systems and your competitor's do not go down on the same morning for the same reason. Shared providers and shared jurisdictions break that assumption, and the people who price risk noticed first.


Kevin Kalinich of Aon told the Financial Times the industry could absorb one company's agent misfiring but not an upstream failure producing a thousand losses at once.


The market moved on that reasoning. In January 2026 three new generative AI exclusions became available for commercial general liability. ISO form CG 40 47 removes bodily injury, property damage and personal and advertising injury arising out of or attributable to generative AI. Two companion forms extend it. They are optional endorsements rather than automatic terms, which makes their existence the signal: the industry built the language before most buyers asked about it.


Coverage counsel put the underwriting logic more plainly than an essay can. A bad employee email reaches one customer. A bad chatbot answer may reach thousands.


Transferring this risk is getting harder, not easier, at exactly the point more organizations are taking it on.


Who owns this

Five questions for the room where the case is approved

I do not think there is a general answer, and I have seen enough organizations to be wary of anyone offering one.


Where this sits depends on how the company is built, who holds the operating model, and whether continuity reports somewhere with the standing to ask a division a hard question.


What I am confident about is the shape of the problem.


The case is written by the division that benefits. The exposure lands on a function that was not in the room. The assumption connecting them was never written down, because until now it was never wrong.


So the questions are ownership questions before they are technical ones.

  • Who was in the room when the case was approved, and who owns the continuity plan for that process? If those are different people who did not speak, that is the seam.

  • What does our continuity plan assume about the people, and who last checked whether it is still true? The plan says the work will be done manually. It does not say by whom.

  • Which processes still have someone who can do them by hand, and when did they last do it? Names and dates. This is the only question on the list with a factual answer, and most organizations cannot produce it.

  • What is our fallback capacity ratio? Volume, per-person throughput, honest headcount. If it is below 1.0, say so out loud before the next case is written.

  • Which processes are we willing to run with no fallback at all? This is a legitimate answer. Plenty of businesses depend on capabilities that would simply stop, and pretending otherwise is its own kind of dishonesty. What is not legitimate is arriving there without having chosen it.


The business case was never designed to hold any of that. It was built to compare two versions of a division's operating performance, and it does that job well.


It just cannot see the thing it is spending.


There are four seats in the name. The fifth one is yours.

Comments


bottom of page