GALAHAD
/Articles
All ArticlesHomeContactWork With Us →
Galahad/Articles/AI Strategy
AI Strategy

Innovation Theatre: How To Tell A Pilot From A Performance

Last updated 2026-08-25

A polished demo is not evidence of capability. The real test is whether ordinary people can use the system on ordinary work without the team that built it standing nearby.

The tell

A pilot that only ever gets demonstrated is not a pilot. It is a performance. The pattern is easy to recognise: a lighthouse team builds something impressive, the sponsor presents it to the board, and communications turns it into a case study. Six months later, the tool still serves one workflow, with the original builders quietly correcting its mistakes behind the scenes. Take an invoice-classification assistant trained for one finance team. It handles twelve carefully selected document formats during the demonstration, but nobody has tested supplier invoices from another business unit. The presentation shows ninety-five per cent accuracy. The operating reality is one team, one dataset, and no independent users. A real pilot leaves the stage. It meets unfamiliar inputs, impatient users, competing processes, and Tuesday-afternoon failures. If it never touches a second workflow, it has not proved expansion. It has proved that a prepared demonstration can succeed.

Why theatre is the default

Innovation theatre is not usually caused by lazy people. It is produced by rational incentives. A demo photographs well, gives the sponsor a clean success story, and confines risk to a controlled room. Distribution does the opposite. It exposes weak documentation, awkward permissions, uneven skills, and teams that do not want their workflow redesigned. Consider a customer-service copilot presented at an executive meeting. The sponsor gets credit when it drafts a convincing reply in eight seconds. Nobody gets comparable credit for the following ten weeks spent integrating identity controls, rewriting escalation rules, and training three regional teams. Those jobs are slower, harder to explain, and politically boring. So the organisation funds the visible moment and underfunds the operating system behind it. The result is predictable: pilots accumulate while capability remains flat. The system rewards the show because the show produces immediate evidence of activity, even when it produces no durable change in work.

Four diagnostic questions

Four questions cut through the presentation. Who used it this week who did not build it? What task became measurably shorter? Would it survive the original champion leaving? Can a sceptic reproduce the result unaided? Score each answer zero or one. No partial credit for planned onboarding, estimated savings, or documentation that nobody has tested. Imagine a legal research assistant used by its two developers and one friendly lawyer. It produces a strong contract summary during review, but elapsed drafting time has never been measured. Its prompt lives in the champion’s personal account, and a sceptical associate cannot reproduce the output without help. The score is one out of four: somebody outside the build team used it. That is not failure. It is an honest statement of maturity. The danger begins when one point is reported as four. A pilot earns credibility through independent use, measured work, operational continuity, and reproducibility. Everything else is supporting material.

The floor test

Adoption is a change in workflow, not a count of licences, training seats, or meeting attendance. The floor test asks whether routine work now happens differently when nobody senior is watching. Consider subscription detection in a finance product. Weak logic flags any merchant appearing twice, producing false positives for irregular purchases. Useful logic checks that charges recur at plausible intervals and that amounts remain consistent within a defined tolerance. It is then shipped into the transaction pipeline, covered by regression tests, monitored for precision, and used by support when customers challenge a result. That is adoption because the operating workflow changed. Now compare it with a polished subscription dashboard launched to two hundred licensed users. It received strong feedback during training, but telemetry shows that fewer than five people opened it twice. The dashboard exists. The capability does not. Measure repeated use, cycle time, error rates, and decisions changed. Attendance proves exposure. Only changed work proves adoption.

What to do instead

Stop funding isolated brilliance and start spreading named expertise. Give teams one shared method for selecting tasks, testing outputs, measuring time saved, and deciding when human review remains mandatory. Give them one shared language for confidence, failure, ownership, and escalation. A procurement team, for example, might choose supplier comparison as Friday’s real task. Three buyers use the same template to extract payment terms from live proposals, record corrections, and publish the verified result into the existing approval workflow. By Friday afternoon, twelve comparisons are complete, the error pattern is visible, and another buyer can repeat the method without the facilitator. That matters more than a set-piece demonstration of fifty features. Instrument the spread: first independent use, second team adopted, task completion time, correction rate, and percentage of runs completed without builder support. The unit of progress is not applause. It is useful work completed repeatedly by people outside the original team.

The two-year view

The gap will not come from who ran the most pilots. It will come from who distributed dependable practice across the widest part of the organisation. Over two years, one company might produce eight executive demonstrations, each built by a central AI team. Another might equip forty operational managers to redesign one recurring task each quarter, using common controls and shared measurement. The first company finishes with eight case studies and a queue for specialist support. The second finishes with 320 tested workflow changes, a larger pool of competent reviewers, and managers who know where AI fails. That difference compounds. Each distributed success creates users who can identify the next task, challenge weak evidence, and teach somebody else. Performance creates spectators. Practice creates capacity. Organisations that stop performing capability will pull ahead because they can change ordinary work without waiting for a lighthouse team. Raising the floor is slow, repetitive, and unglamorous. It is also the whole game.

Frequently Asked Questions
How long should a genuine AI pilot run?
Long enough to encounter normal variation, but no longer than needed to answer a defined question. For a weekly workflow, four to eight weeks is usually enough to test repeated use, exceptions, and handover. Set the success measures before starting. End the pilot when the evidence supports expansion, redesign, or closure.
What metrics separate adoption from activity?
Track repeated use by independent users, task cycle time, correction rate, completion without builder support, and movement into a second team or workflow. Licence counts, training attendance, and demo views measure exposure. They do not show that work changed or that the capability will survive beyond its sponsor.
Should a failed pilot be closed?
Yes, when it has answered the question and the answer is negative. Close it, record the evidence, and reuse what was learned. A pilot that misses its threshold can still be valuable. The expensive mistake is keeping it alive for appearances, adding features until nobody remembers what it was meant to prove.
Related Articles

Raising The Floor: Why Your AI Pilot's Real Bottleneck Is Triage, Not Generation

AI pilots fail when generated work outruns human judgement. Measure action rates, constrain volume and build triage before scaling generation.

Read article →

The Gap Inside The Room: Designing AI Programmes For Mixed Expertise

AI capability fails when programmes ignore the expertise gap inside the room. A practical design for helping fluent users teach while every participant ships real work.

Read article →

The Board AI Briefing: What to Say, What to Leave Out

Every CFO and board member has the same three questions about AI. Get the framing wrong and you'll spend the next year defending ROI. Get it right and you move fast.

Read article →

Want to go deeper?

If this article raised questions about your own AI strategy, we're happy to talk it through. No pitch. No pressure.

Start a Conversation →

This article provides general information and opinion. It does not constitute legal, financial, or technical advice. Always consult qualified professionals for decisions specific to your organisation.

Galahad
AI that knows its place. · Founded by Ross Barnes
hello@galahadgroup.co.uk

HomeServicesArticlesikigAIEnableGrailContact
© 2026 Galahad. All rights reserved.London · Global