An AI Pilot Needs an Operating Thesis
An AI pilot is not an operating plan.
It is a container. It tells us that a company intends to try some technology for a limited period. It does not tell us which work will change, who will do less of it, where judgment will remain, or how anyone will know the trial helped.
That missing statement is the operating thesis.
A useful operating thesis sounds like this: routine order-status calls can be handled from approved facts and documented automatically, while exceptions arrive in front of an experienced CSR with the context required to act.
That can be tested. It names the work, the boundary, the human role, and the expected change. "We are piloting an AI phone agent" names a technology and leaves the operation blank.
Name The Behavior That Changes
An operating thesis describes how work should behave differently after the pilot.
It should answer a few plain questions:
- What recurring work enters the pilot?
- What happens today when that work arrives?
- Which step should become faster, more consistent, or easier to inspect?
- What still requires a person?
- Where does an uncertain or failed case go?
- Which measure would show that the change was useful?
These questions force the proposal out of demonstration language. A model can answer a phone, draft an email, classify a request, or summarize a case. Those are capabilities. The operating thesis explains what happens around the capability and why the result matters.
For a support operation, the thesis might be that routine calls can follow one approved playbook across every shift while experienced representatives focus on exceptions. For a legal workflow, it might be that new matters arrive with required documents checked and missing information marked before an attorney reviews them. For a weekly reporting process, it might be that source changes are gathered and compared before the manager spends attention on the result.
Each thesis describes changed work. None depends on vague promises about what AI can do.
Start With The Failure Mode
The best pilot boundary usually begins with a recurring failure mode.
Customers wait because the same routine question sits in a call queue. Notes vary by representative, so the next person has to reconstruct the interaction. A manager keeps answering the same procedural question. An outage turns a normal process into improvised work. A handoff loses the reason the case was escalated. A report takes hours because someone must collect the same five sources every week.
These are observable conditions. They can be counted, timed, sampled, and compared. They also reveal what the system must preserve. If the current process depends on an experienced person recognizing a dangerous exception, the pilot needs an escalation path. If the source material changes often, the pilot needs an owner for that material. If a wrong action would affect money, rights, or a customer promise, the pilot needs review before the action occurs.
This is why most AI problems are operational problems. The model may be capable of the visible step while ownership, source quality, permissions, or handoffs remain unresolved. A pilot that ignores those conditions can produce an impressive result without improving the operation.
Give The Pilot A Narrow Lane
A useful pilot is smaller than the aspiration behind it.
"Handle customer service" is not a lane. "Answer routine order-status calls after verification, document the result, and escalate every exception" is a lane. The second version has an entry condition, approved facts, an ordinary path, stop conditions, and a destination for work the system should not finish.
That narrowness is not timidity. It makes the operating thesis measurable. The team can inspect routine completion, verification failures, escalation reasons, documentation quality, repeat contacts, tool failures, and supervisor time. It can learn whether the system reduced repeated work or merely moved it somewhere less visible.
We used that boundary in the AI CSR build. The AI handles one bounded inbound workflow from approved facts. Experienced CSRs supervise, receive difficult handoffs, review sampled interactions, and improve the playbook. The operating change is not "a voice model answers the phone." It is "routine work becomes more consistent while experienced attention moves to exceptions, recovery, and improvement."
That is a better pilot because failure can also teach us something. A high escalation rate may reveal that the lane is poorly defined. Repeated missing information may expose an intake problem. Supervisor corrections may show where the playbook is incomplete. A workflow trial should make those findings visible instead of treating every human intervention as a defect.
Measure The Operation, Not The Novelty
Novelty produces bad pilot metrics. Number of conversations, model response speed, or a polished demonstration can matter during implementation, but they do not establish that the operation improved.
The measures should follow the operating thesis. If the claim is better continuity, measure whether eligible work was completed during the hours or conditions that previously caused delay. If the claim is consistency, sample whether required checks and documentation appeared each time. If the claim is better use of experienced staff, measure supervisor attention per completed case and the quality of escalated context. If the claim is faster recovery, measure how quickly failed or uncertain cases reach a responsible person.
The pilot also needs a stopping rule. If routine completion remains low, exceptions overwhelm the reviewers, customers repeat themselves, or the team cannot maintain the approved source material, the operating thesis may be wrong. The responsible result may be a narrower lane, a workflow change before automation, a conventional software purchase, or no build.
That is why a Workflow Assessment comes before the build. The assessment names the bottleneck, maps the ordinary and exceptional paths, and decides which change is worth testing. The pilot then has a job beyond proving that a model can perform.
Remove AI From The Sentence
There is a simple test for an AI pilot proposal.
Remove the words "AI pilot" and describe the operating change by itself.
Will routine requests receive a consistent answer from approved facts? Will a handoff reach the right person with enough context? Will a weekly review begin with the source material already gathered? Will the operation continue through a common disruption? Will experienced people spend more time on judgment and less on reconstruction?
If that change is not clear or valuable, adding AI will not rescue the proposal. If the change is specific, measurable, and worth making, then the team can decide whether AI is the right component.
The pilot is the test. The operating thesis is the reason to run it.
