GALAHAD
/Articles
All ArticlesHomeContactWork With Us →
Galahad/Articles/AI Strategy
AI Strategy

Raising The Floor: Why Your AI Pilot's Real Bottleneck Is Triage, Not Generation

Last updated 2026-08-19

Your AI pilot does not need another model upgrade. It needs a credible way to decide what deserves attention, who must act and what gets killed when nobody does.

The demo lie

A successful demo proves that a model can produce an answer. It does not prove that an organisation can absorb one. That distinction disappears in most pilot reporting. The team shows ten plausible outputs, leaders nod, and generation becomes the headline. Then the tool enters a real workflow. A service agent receives a suggested response but cannot see which policy sources support it. Sending the reply takes thirty seconds; checking it takes six minutes. The agent returns to the old template because the queue is growing. The model worked. The organisation did not act. Adoption stalls in the gap between output and accountable action: verification, routing, approval, escalation and ownership. If those steps are absent, better generation only creates more material for people to ignore. The demo tested whether AI could speak. The pilot must test whether work moves.

A worked failure

Fifteen proposals in a fortnight sounds productive until the status column is examined. Zero approved. Zero rejected. All fifteen expired. That result is easy to misread as weak model performance, but nobody made a judgement about the outputs. The system produced proposals, placed them into an approval queue and waited. Reviewers had no protected time, no ranking signal and no consequence for silence. By day fourteen, every proposal had crossed its expiry threshold untouched. The correct diagnosis is not that the model generated bad ideas. It is that the operating system around the model could not evaluate them. Rejection would have produced useful evidence about quality. Approval would have produced action. Expiry produced neither. An unread proposal has no measurable value, however sophisticated its reasoning. When every item dies in the queue, fix triage before touching the prompt.

The saturation point

Every approval process has a carrying capacity. Cross it and the queue stops being a decision mechanism. It becomes storage. Consider thirty-seven pending approvals spread across six concurrent projects. Each item takes twelve minutes to inspect properly, including checking evidence and recording a decision. That is more than seven hours of review before meetings, interruptions or follow-up questions. No single reviewer owns the whole queue, so everyone selects the familiar items and leaves the ambiguous ones behind. Age becomes the accidental prioritisation rule. Strong proposals expire beside weak ones because the system cannot distinguish urgency, value or risk. This is the saturation point: supply has outrun judgement. Adding another agent or increasing generation frequency makes the backlog worse. Capacity must be explicit. If reviewers can assess ten proposals each week, the system cannot responsibly produce fifty. Generation needs a throttle tied to actual review bandwidth.

Precision over volume

The useful metric is the percentage of generated output that leads to human action. Everything else is supporting evidence. A pilot that produces 200 recommendations and prompts eight actions has a four per cent action rate. Another produces twelve recommendations and prompts nine actions: seventy-five per cent. The first will look larger on a dashboard. The second is changing work. Action should be defined before launch: approved, rejected with a reason, implemented, escalated or converted into a tracked task. Views and impressions do not count. Neither does an output copied into a document and forgotten. Break the rate down by workflow, risk level and reviewer so the bottleneck becomes visible. Low action may indicate poor relevance, missing evidence, unclear ownership or excessive volume. Output counts conceal all four. Measure what people do after generation. If the answer is nothing, the system is producing inventory, not value.

Fixes that hold

Durable fixes reduce demand on judgement before adding capacity. Raise the proposal bar: require evidence, expected value, named ownership and a clear decision request. Narrow scope so one queue handles one class of decision. Instrument approval, rejection and expiry rates by week. Then delete expired items rather than preserving a flattering backlog. Suppose a system generates forty change proposals monthly, reviewers act on eight and twenty expire unread. Cap generation at ten, rank by expected impact and require a reviewer assignment before submission. After a month, seven are approved, two rejected and one expires. Total output falls by seventy-five per cent while useful decisions barely move. That is an improvement. At portfolio level, the same discipline means deferring or cancelling projects whose review burden exceeds their likely value. A pilot without capacity limits is not ambitious. It is unmanaged demand dressed as innovation.

Raising the floor, not the ceiling

The next spectacular demo will not fix a broken operating model. Raising the floor means making ordinary use reliable for the many people who must review, route and act. Take a claims team where one specialist can build an impressive agent but forty handlers cannot tell why it recommended escalation. The ceiling is high: the specialist can demonstrate complex reasoning. The floor is low: handlers lack evidence links, decision rules and a fast correction path. Fixing those basics may look less exciting than launching another agent, but it turns capability into routine work. Anti-innovation-theatre means funding the unglamorous layer: queue ownership, review time, thresholds, audit trails, training and stop conditions. It also means refusing pilots that cannot name the decision they will change. The standard is simple. Generate less, make judgement easier and measure action. An AI capability exists only when the organisation can use it without depending on a handful of exceptional people.

Frequently Asked Questions
What should an AI pilot measure instead of output volume?
Measure the action rate: the percentage of generated outputs that receive a defined human response. Track approvals, reasoned rejections, implementations, escalations and expiries separately. Segment the results by workflow and reviewer. A high output count with a low action rate indicates accumulated inventory, not adoption.
How do you know when an approval queue is saturated?
Saturation starts when incoming proposals consistently exceed the number reviewers can assess properly. Warning signs include rising queue age, unread expiries, reviewers selecting only familiar cases and decisions arriving after their operational value has passed. Set a weekly review capacity and throttle generation against it.
Should expired AI proposals remain in the backlog?
No. Remove them from the active queue and record their metadata for analysis. An expired proposal should not compete with current work or inflate the appearance of demand. Repeated unread expiry is evidence that scope, ownership or generation volume is wrong. Fix that system before producing more.
Related Articles

Innovation Theatre: How To Tell A Pilot From A Performance

A practical test for separating genuine AI pilots from polished performances, measuring workflow change, and building capability that survives beyond the demo.

Read article →

The Gap Inside The Room: Designing AI Programmes For Mixed Expertise

AI capability fails when programmes ignore the expertise gap inside the room. A practical design for helping fluent users teach while every participant ships real work.

Read article →

The Board AI Briefing: What to Say, What to Leave Out

Every CFO and board member has the same three questions about AI. Get the framing wrong and you'll spend the next year defending ROI. Get it right and you move fast.

Read article →

Want to go deeper?

If this article raised questions about your own AI strategy, we're happy to talk it through. No pitch. No pressure.

Start a Conversation →

This article provides general information and opinion. It does not constitute legal, financial, or technical advice. Always consult qualified professionals for decisions specific to your organisation.

Galahad
AI that knows its place. · Founded by Ross Barnes
hello@galahadgroup.co.uk

HomeServicesArticlesikigAIEnableGrailContact
© 2026 Galahad. All rights reserved.London · Global