A pilot that only ever gets demonstrated is not a pilot. It is a performance. The pattern is easy to recognise: a lighthouse team builds something impressive, the sponsor presents it to the board, and communications turns it into a case study. Six months later, the tool still serves one workflow, with the original builders quietly correcting its mistakes behind the scenes. Take an invoice-classification assistant trained for one finance team. It handles twelve carefully selected document formats during the demonstration, but nobody has tested supplier invoices from another business unit. The presentation shows ninety-five per cent accuracy. The operating reality is one team, one dataset, and no independent users. A real pilot leaves the stage. It meets unfamiliar inputs, impatient users, competing processes, and Tuesday-afternoon failures. If it never touches a second workflow, it has not proved expansion. It has proved that a prepared demonstration can succeed.
Innovation Theatre: How To Tell A Pilot From A Performance
A polished demo is not evidence of capability. The real test is whether ordinary people can use the system on ordinary work without the team that built it standing nearby.
The tell
Why theatre is the default
Innovation theatre is not usually caused by lazy people. It is produced by rational incentives. A demo photographs well, gives the sponsor a clean success story, and confines risk to a controlled room. Distribution does the opposite. It exposes weak documentation, awkward permissions, uneven skills, and teams that do not want their workflow redesigned. Consider a customer-service copilot presented at an executive meeting. The sponsor gets credit when it drafts a convincing reply in eight seconds. Nobody gets comparable credit for the following ten weeks spent integrating identity controls, rewriting escalation rules, and training three regional teams. Those jobs are slower, harder to explain, and politically boring. So the organisation funds the visible moment and underfunds the operating system behind it. The result is predictable: pilots accumulate while capability remains flat. The system rewards the show because the show produces immediate evidence of activity, even when it produces no durable change in work.
Four diagnostic questions
Four questions cut through the presentation. Who used it this week who did not build it? What task became measurably shorter? Would it survive the original champion leaving? Can a sceptic reproduce the result unaided? Score each answer zero or one. No partial credit for planned onboarding, estimated savings, or documentation that nobody has tested. Imagine a legal research assistant used by its two developers and one friendly lawyer. It produces a strong contract summary during review, but elapsed drafting time has never been measured. Its prompt lives in the champion’s personal account, and a sceptical associate cannot reproduce the output without help. The score is one out of four: somebody outside the build team used it. That is not failure. It is an honest statement of maturity. The danger begins when one point is reported as four. A pilot earns credibility through independent use, measured work, operational continuity, and reproducibility. Everything else is supporting material.
The floor test
Adoption is a change in workflow, not a count of licences, training seats, or meeting attendance. The floor test asks whether routine work now happens differently when nobody senior is watching. Consider subscription detection in a finance product. Weak logic flags any merchant appearing twice, producing false positives for irregular purchases. Useful logic checks that charges recur at plausible intervals and that amounts remain consistent within a defined tolerance. It is then shipped into the transaction pipeline, covered by regression tests, monitored for precision, and used by support when customers challenge a result. That is adoption because the operating workflow changed. Now compare it with a polished subscription dashboard launched to two hundred licensed users. It received strong feedback during training, but telemetry shows that fewer than five people opened it twice. The dashboard exists. The capability does not. Measure repeated use, cycle time, error rates, and decisions changed. Attendance proves exposure. Only changed work proves adoption.
What to do instead
Stop funding isolated brilliance and start spreading named expertise. Give teams one shared method for selecting tasks, testing outputs, measuring time saved, and deciding when human review remains mandatory. Give them one shared language for confidence, failure, ownership, and escalation. A procurement team, for example, might choose supplier comparison as Friday’s real task. Three buyers use the same template to extract payment terms from live proposals, record corrections, and publish the verified result into the existing approval workflow. By Friday afternoon, twelve comparisons are complete, the error pattern is visible, and another buyer can repeat the method without the facilitator. That matters more than a set-piece demonstration of fifty features. Instrument the spread: first independent use, second team adopted, task completion time, correction rate, and percentage of runs completed without builder support. The unit of progress is not applause. It is useful work completed repeatedly by people outside the original team.
The two-year view
The gap will not come from who ran the most pilots. It will come from who distributed dependable practice across the widest part of the organisation. Over two years, one company might produce eight executive demonstrations, each built by a central AI team. Another might equip forty operational managers to redesign one recurring task each quarter, using common controls and shared measurement. The first company finishes with eight case studies and a queue for specialist support. The second finishes with 320 tested workflow changes, a larger pool of competent reviewers, and managers who know where AI fails. That difference compounds. Each distributed success creates users who can identify the next task, challenge weak evidence, and teach somebody else. Performance creates spectators. Practice creates capacity. Organisations that stop performing capability will pull ahead because they can change ordinary work without waiting for a lighthouse team. Raising the floor is slow, repetitive, and unglamorous. It is also the whole game.
Want to go deeper?
If this article raised questions about your own AI strategy, we're happy to talk it through. No pitch. No pressure.
Start a Conversation →This article provides general information and opinion. It does not constitute legal, financial, or technical advice. Always consult qualified professionals for decisions specific to your organisation.