Journal · AI and Enterprise Control
A Pilot Should End in a Decision
Many AI pilots show promise and still go nowhere. The problem often starts with how the pilot was designed.
The demonstration worked. What happens next?
An AI pilot produces promising results. The demonstration is convincing. The team reports time saved, users express interest, and the business sponsor sees potential.
Then the questions begin.
Will the results hold with production data? Who handles exceptions? What happens when the AI gets a decision wrong? Who owns the process after the project team leaves? Does the productivity gain translate into a measurable business outcome?
These are not implementation details to resolve after approval. They are part of the investment decision.
A composite example: the pilot that answered the wrong question
The following is an illustrative composite, not a report of a specific client engagement.
A team pilots an AI assistant to help employees handle service requests. Responses are generated quickly, and employees spend less time drafting routine answers.
The team measures response-generation time and user feedback. The results look promising, so the pilot is declared successful.
Now consider the same pilot with a decision plan agreed before the first test.
The team tracks three measures:
- Quality: How often does a response meet the agreed standard without correction? A result below the agreed minimum triggers redesign or a stop decision.
- End-to-end economics: What is the cost per successfully completed case, including human review, correction and exception handling? A result above the agreed cost threshold blocks scale-up.
- Operational reliability: How often does the assistant fail, escalate or require intervention under representative conditions? Performance outside agreed tolerances triggers further testing or redesign.
The business owner also defines the target outcome, risk boundaries, decision authority and date for reviewing the evidence.
The pilot now has a purpose beyond demonstration. Its evidence will inform a decision.
The results might still be disappointing. If the experiment identifies a constraint early enough to prevent a poor investment, it has served its purpose.
The failure is declaring success before defining what success needs to prove.
Feasibility is not investment readiness
A pilot needs to distinguish between two questions.
Feasibility: Can the technology perform the task under defined conditions?
Investment readiness: Does the capability produce a measurable business outcome, within acceptable risk and operating cost, in the environment where it will be used?
The first question matters. It is rarely sufficient on its own.
McKinsey's August 2026 State of AI survey illustrates the gap. Respondents report 44% enterprise-wide AI scaling and 37% reporting some AI-related EBIT impact. About one in five report constraints from AI operating costs. Nearly three-quarters of high performers report fundamentally redesigning workflows, compared with one-quarter of other respondents.
These are self-reported survey findings. They do not establish why individual pilots stall. They do show why adoption, individual productivity and enterprise financial impact should not be treated as interchangeable outcomes.
The workflow findings matter too. Inserting a tool into an existing process and changing how work gets done are different interventions.
Design the pilot to answer the investment question
Five practices make the evidence more useful.
1. Test representative conditions
Use data, users and workflows that reflect the intended operating environment. Include difficult cases, exceptions and realistic volumes where feasible.
A controlled demonstration establishes what works under controlled conditions. It does not establish how the capability behaves in production.
2. Measure the complete task
Measure more than generation speed or task completion by the AI.
Include human review, corrections, overrides, exception handling, integration, support and monitoring. Where relevant, assess cost per successfully completed case, not simply the cost of producing an AI response.
3. Name the business owner from day one
The business owner is accountable for the outcome, not merely for approving the pilot.
This person owns the target measure, the workflow change and the scale-up recommendation. Technology and delivery teams contribute evidence, but they do not substitute for business accountability.
4. Test the operating model
Establish who reviews outputs, who handles exceptions, who can override the AI and who intervenes when performance deteriorates.
Identify the required data access, system integration, security controls, training and ongoing monitoring.
NIST's AI Risk Management Framework offers a useful reference through its four functions: Govern, Map, Measure and Manage. These functions address risk across the AI lifecycle, including after deployment.
5. Trace value from plan to evidence
Use four distinct labels:
- Planned value: the outcome and benefit expected when the investment is approved.
- Actual value: the result measured after implementation.
- Claimed value: the benefit reported by a team or sponsor but not yet fully substantiated.
- Validated value: the benefit supported by evidence and an agreed validation method.
For example, if an AI assistant saves employee time, establish what happens to the released capacity. Does the organisation process more cases, reduce overtime, improve service or redeploy people to higher-value work?
Attribution needs a method. Where practical, compare results with a suitable control group. Otherwise, use a before-and-after baseline and document other changes that might explain the result.
Without a credible connection between the intervention and the outcome, time saved remains a productivity indicator, not proof of business value.
Write the decision before running the pilot
Agree the decision criteria before the experiment begins. Otherwise, the team risks redefining success after seeing the results.
MIT CISR's 2018 research briefing on test-and-learn innovation describes how Deutsche Telekom addressed a portfolio problem: it had become too easy to start experiments and too difficult to stop them. The company introduced staged funding, with tangible milestones defined in advance and reviewed before further investment.
This is an analogous practice, not AI-specific evidence. The briefing concerns digital innovation portfolios and includes one company's experience. Its relevance is the principle: commit to the evidence required for the next investment decision before committing the next tranche of funding.
The one-page AI pilot charter
Agree this charter before the pilot starts. It is the decision mechanism, not an additional layer of paperwork.
Decision sought
What must be agreed
The investment decision the pilot will support.
Business outcome
What must be agreed
Planned value, baseline, target and measurement method.
Scope and conditions
What must be agreed
Intended users, process, data, volume and exclusions.
Success criteria
What must be agreed
Quality, reliability, economics and risk thresholds.
Accountability
What must be agreed
Named business owner, decision authority and control owners.
Decision rules
What must be agreed
The scale, redesign, stop and defer rules below.
Dates
What must be agreed
Pilot end date, evidence review, decision date and retest date, where needed.
Next funding step
What must be agreed
What approval releases and the conditions for further investment.
The success criteria define readiness for the specific use case. They should be agreed in advance, proportionate to risk and supported by evidence.
Scale
Evidence and required action
Business outcome meets the agreed target; reliability, risk, ownership and operating economics meet requirements. Release the next funding stage subject to production controls.
Who decides
Named business owner and designated investment authority.
By when
At the agreed decision gate.
Redesign
Evidence and required action
A specific gap prevents the capability from meeting criteria, and a credible intervention exists. Name the change, retest and required evidence.
Who decides
Business owner with delivery and risk leads.
By when
By the charter's review or retest date.
Stop
Evidence and required action
The business case fails the agreed threshold, a material risk remains unacceptable, or no credible path to the required outcome exists. Close the experiment and record the learning.
Who decides
Designated investment authority, informed by business and risk owners.
By when
At the decision gate, or earlier if a stop condition is triggered.
Defer
Evidence and required action
A prerequisite remains unresolved, such as data access, process ownership or a necessary control. Assign an owner and a dated review.
Who decides
Business owner and authority responsible for the prerequisite.
By when
On the specific review date recorded in the charter.
Make departures from the charter visible. Record every decision that departs from the agreed criteria, including the reason and the person authorising the departure. For consequential use cases, require an independent reviewer who was not responsible for championing the pilot or its scale-up.
This safeguard matters because the business owner may have a strong interest in the pilot succeeding. Pre-commitment has little force if the same people can change the criteria without scrutiny.
A defer decision without an owner and review date is not a controlled decision. It is an open-ended extension.
Readiness depends on the use case
A low-risk internal drafting assistant and an AI capability initiating consequential financial or operational actions do not need identical controls.
For a low-risk assistant, sample-based human review, access restrictions and usage monitoring might be proportionate. For a system initiating consequential actions, requirements might include human approval before action, detailed audit logging, independent testing and defined emergency intervention.
The appropriate controls depend on the impact of an error, the reversibility of the action, the exposure of affected people and the operating context.
When this advice does not apply
Not every pilot should aim to scale.
Some experiments exist to answer an early question: is the idea technically plausible, is a new approach worth exploring, or what unknowns need investigation?
These experiments should be judged by the quality of the learning, not by immediate production readiness.
MIT CISR's test-and-learn research supports defining hypotheses, measuring results and using evidence to decide whether an experiment should continue or be discarded. Its findings concern innovation more broadly, rather than AI pilots specifically.
A learning experiment should state what the team expects to discover and what it will do with the result. A production-oriented pilot should test the conditions required for an investment decision.
Confusing the two creates problems in both directions: promising experiments are judged too early by scale-up criteria, while demonstrations are approved for investment without sufficient evidence.
What evidence would make us scale, redesign, stop or defer?
Ask the question before approving the pilot.
Then agree who owns the decision, which evidence will determine it and when the decision will be made.
If the team cannot answer, the organisation has not yet defined what the pilot is supposed to decide.
Sources
1. McKinsey & Company (2026). The state of AI in 2026: On the road to ROI. Published 25 August 2026. Figures cited are from the August 2026 edition; accessed 11 October 2026. https://www.mckinsey.com/capabilities/quantumblack/our-insights/the-state-of-ai
2. National Institute of Standards and Technology (2023). Artificial Intelligence Risk Management Framework (AI RMF 1.0). NIST AI 100-1. Accessed 11 October 2026. https://doi.org/10.6028/NIST.AI.100-1
3. Fonstad, N. O. and Ross, J. W. (2018). Learning How to Test and Learn. MIT Sloan Center for Information Systems Research, Research Briefing XVIII-2, 15 February 2018. Accessed 11 October 2026. https://cisr.mit.edu/publication/2018_0201_TestAndLearn_FonstadRoss
Related
Read next.
Journal · AI and Enterprise Control
Why Most Enterprise AI Pilots Never Become Operating Capabilities
Most AI pilots do not fail because the model stops working. They fail because the organisation has not built the operating capability around the model.
Journal · AI and Enterprise Control
Before an AI Agent Touches a Real Process
Before an AI agent touches a real process, define the boundary.
Journal · AI and Enterprise Control
When AI Starts Acting, Who Owns the Outcome?
AI changes the conversation when it moves from answering questions to taking action.