From pilots to production, deliberately
By now most leadership teams have an AI pilot portfolio and a board asking why none of it has reached production. The gap is rarely technical. Pilots stall because nobody chose the process by its economics, measured the baseline it was meant to beat, checked that the data would survive contact with production, or decided who owns the output when it is wrong. This playbook sequences those choices. It pairs with the Before Implementing AI guide, which sets out the questions to settle before spend is committed; the playbook assumes you are proceeding and walks the implementation end to end, from candidate selection to scaling. One framing carries through every step: you are not adopting a technology, you are redesigning a process that will now include a component with a measurable error rate. Organisations that hold that framing make unglamorous, compounding progress. Organisations that chase capability demonstrations accumulate pilots.
Step 1: Choose the process by its economics
Select the first production candidate with an explicit filter, not by which demonstration impressed the executive team. Score every candidate process on four measurable properties: volume (how many times a day the task runs, because the economics need repetition), cost per execution today (measured, including rework), error tolerance (what one wrong output costs, and how visibly it fails), and data availability (whether the inputs exist digitally, consistently, now). The strongest first candidates are high-volume, moderate-cost, error-tolerant and data-rich: document triage, first-draft responses, classification, extraction from standard forms. The weakest are the ones most often proposed: rare, high-stakes judgements with thin data, where the system cannot learn and the errors cannot be afforded. Rank the portfolio, pick one process, and write down why the others lost. That record restrains the next wave of executive enthusiasm and shortens the next selection round. Expect the filter to be unpopular, because the processes it favours are rarely the ones executives volunteer: visible flagship projects score badly on exactly the properties that make first implementations succeed.
Step 2: Baseline the process you intend to beat
AI cases are routinely argued against an imagined perfect process instead of the real one, which distorts the decision in both directions. Spend two to four weeks measuring the incumbent process before any vendor conversation: actual unit cost including rework and supervision, actual cycle time including queues, and, most usefully, the current human error rate, sampled honestly, because almost nobody knows theirs and it is frequently worse than assumed. The baseline changes the conversation twice over. It gives the business case a denominator finance will accept, and it converts the accuracy debate from "is the model ever wrong" to "is the model wrong more often than the process we run today", which is the only version of the question that leads anywhere. Publish the baseline before vendor selection so nobody can renegotiate it afterwards, and keep the measurement method, because you will re-run it on the system's output in Step 6.
Step 3: Test the data before you believe an accuracy claim
Every accuracy claim, vendor or internal, is conditional on data the claim never describes. Before selecting tooling, run the Data Readiness Assessment against the specific process from Step 1, not the organisation in general: do the inputs exist digitally for the full population or only the easy cases, are they consistent across regions and systems, is there labelled history to evaluate against, who owns corrections, and are you contractually and legally permitted to use the data this way, including customer terms and residency constraints. Then build the evaluation set that will govern the whole programme: a representative sample of real cases, including the messy ones, with correct answers agreed by the people who do the work today. This set is the instrument that every vendor claim, pilot result and production check gets measured against. Building it is unglamorous, and it typically reveals more about the organisation's data than any strategy deck has, which is exactly why it cannot be skipped.
Step 4: Design the oversight model before the pilot
Decide where humans sit in the new process before piloting, because oversight design determines both the risk profile and the economics, and retrofitting it after go-live is expensive in both.
| Oversight design |
Where it fits |
Effect on the case |
| Human reviews every output |
Low error tolerance; early confidence building |
Safest start; savings limited to drafting time |
| Human samples a percentage |
Stable accuracy; moderate error cost |
Savings scale; the sampling design becomes the control |
| Human handles exceptions only |
High volume; well-understood failure modes |
Full economics; depends entirely on exception detection |
| System acts autonomously |
Reversible actions with automatic checks |
Rarely the right first step, whatever the vendor says |
Step 5: Run the pilot as a production rehearsal
The purpose of the pilot is not to prove the technology works; vendors have already proven that on friendly data. The purpose is to measure what production will be like. So run it on live, unfiltered inputs including the messy cases, with the Step 4 oversight model operating as designed, staffed by the people who will own the process rather than by the vendor's engineers, and for long enough to see volume peaks and input drift, which usually means a full business cycle rather than a demonstration fortnight. Measure against the Step 2 baseline and the Step 3 evaluation set: accuracy on your cases, exception rates, time per human review, and failure modes by type, because the overall accuracy number matters less than which cases fail and how detectably they fail. Agree the production thresholds in writing before the pilot starts, so the go decision is a comparison against pre-agreed criteria, not a negotiation with your own sunk costs. Resist the temptation to tidy inputs mid-pilot to help the system along; production will not extend that courtesy.
Step 6: Build the cost case at production volume
Pilot economics flatter every programme: volumes are small, the vendor is attentive and the hidden work is absorbed as enthusiasm. Before committing, rebuild the case at production volume with the full cost stack: inference or licence costs at real throughput, integration build and maintenance, the monitoring from Step 7, the human oversight time from Step 4 priced at loaded cost, periodic re-evaluation as inputs drift, and the exception-handling capacity that never appears in vendor proposals. Run the comparison through the AI vs Human Decision Calculator using your measured baseline rather than industry claims, then stress the result: if accuracy on your data is ten points worse than the pilot suggested, if volumes halve, if the sampling rate has to double after an incident, does the case survive? A case that only works at pilot attention levels and optimistic accuracy is not a production case, and finding that out now costs a spreadsheet rather than a programme.
Step 7: Put accountability in writing before go-live
Production readiness is an organisational state, not a technical one. Before go-live, have four things written and signed. A named business owner for the system's outputs, distinct from the IT owner of its infrastructure. The error thresholds that trigger review, pause or rollback, with the person who can invoke each named. The monitoring that runs continuously: accuracy sampling against the evaluation set, drift on inputs, exception-rate trends. And the customer and regulator script for when a consequential output was machine-made, because drafting it during the incident is too late. The AI Analysis vs Human Judgement comparison is useful here for drawing the boundary formally: which classes of decision the system may make, and which always route to a person. Then rehearse one failure before customers find one. Inject known-bad cases, and watch whether the monitoring catches them and whether the rollback actually rolls back. Organisations that handle AI incidents well are the ones that had the ownership argument before go-live, on paper, with names attached.
Step 8: Scale by process, not by enthusiasm
Once the first system stabilises, pressure arrives to do this everywhere at once. Scale deliberately instead: return to the Step 1 ranking, take the next process, and repeat the sequence, which now runs faster because the evaluation-set discipline, the oversight patterns and the monitoring stack already exist. At the portfolio level, re-run the AI Readiness Assessment periodically and keep a live register of every model in production with its owner, thresholds and last evaluation date, reviewed quarterly at executive level, because a set of systems that are each individually fine can jointly exceed the organisation's capacity to supervise them. Each addition also brings a vendor dependency and a data-sharing agreement, and that accumulation should be adopted knowingly rather than discovered at audit. The discipline of one process at a time is what turns the first success into a capability rather than an anecdote.
Where independent operators change the outcome
AI decisions are unusually hard to challenge internally. The enthusiasts overweight the demonstrations, the sceptics overweight the failures, and very few people in either camp have operated an automated process through its second year, when the vendor's attention has moved on and the input data has drifted. Independent challenge from operators who have owned automation outcomes pays at the moments this playbook marks: at Step 1, when the candidate list reflects internal politics more than economics; at the Step 5 exit, when pilot results are being read generously; and at Step 6, when the production cost case is asking to be believed. Selected senior operators from the Global Board review the case confidentially and report on what they would not sign, before the commitment hardens into contracts and infrastructure. It is a small, fast input into a class of decision that is consequential, technical enough to resist internal challenge, and cheaper to pressure-test than to relive.