Skip to content

Why an AI pilot does not become production

A pilot does not fail at the moment something breaks. It fails while it works — sitting beside the process it was built for, rather than inside it.

AploraThe Aplora team4 min read
A board of sticky notes sorted into backlog, this week and in-progress columns

A failed AI project rarely has a dramatic ending. Usually everything looks fine: the solution ships, the demo lands, the team says thank you. Then an ordinary working week arrives and it turns out the old way is faster — not objectively, but for one specific person on one specific Tuesday. A month later two people use it. Three months later, nobody.

1. The process has no owner

The most common and most underestimated reason. An owner is not the project sponsor and not whoever signed the budget. It is the person accountable for the process outcome who also has the authority to change how people work inside it.

Without that person, rollout hits an approval at every step. The solution is technically ready, but using it requires changing a standard operating procedure, and nobody is empowered to change it. The project settles into “works but unused” — the most expensive state possible, because the money is spent and the effect is not.

A check before the start: ask who will answer the question “why is this metric where it is” six months from now. If two names come back, or the answer is a department, there is no owner.

2. There is nothing to compare against

A pilot without a recorded baseline can neither be confirmed nor refuted. The conversation about its future becomes an exchange of impressions — and in that conversation the winner is not whoever is right, but whoever is more confident.

A separate trap is redefining the metric mid-flight. If you started by counting “share of calls reviewed” and end up counting “share of calls the system produced any output for”, that is a different quantity and the comparison is meaningless.

3. No stop criterion was defined

A pilot without a stop criterion does not end — it gets extended. Each extension looks reasonable: two more weeks, one more dataset, one more prompt iteration. The problem is that the decision to continue is made by the person already invested in it, and at each step walking away costs more than it did at the last one.

A stop criterion has to be a number and a date, written down before the start. Not “if it doesn't work out” but “if by this date the quantity has not reached this value at this volume, the work ends”.

A weak initiative stopped in time is the budget for the next one. Stretched out, it becomes two failures instead of one.

4. Adoption was treated as secondary

A tool that monitors a person without giving them anything does not stick. This is not a motivation question and training does not fix it: people rationally choose the way of working where their own result is higher and their risk is lower.

  • The solution must be useful to the person using it, not only to the person watching

  • Individual scores are better hidden at first: until the criteria are calibrated they create resistance and devalue the system

  • In the first weeks, work through contested cases with people rather than sending them a verdict

  • An adoption metric belongs in the project from day one — otherwise it surfaces when it is too late to fix

5. “Done” meant the wrong thing

If done is defined as “the code works”, the project ends exactly there. A working definition looks different, and it is worth agreeing on it up front — then nobody is surprised by the amount of work in the final stage.

“Done” in the project“Done” in the process
The solution works on test dataThe solution sits inside the daily route of work
The demo happenedPeople use it without reminders
The code was handed overAn owner is named on the client side
Documentation existsQuality and cost monitoring is running
The ticket is closedThe effect is measured over a comparable period

What to check before the start

  1. 1

    Who owns the process, and can they change the standard procedure

  2. 2

    Which quantities are fixed as the baseline, and from which source

  3. 3

    What value by what date counts as confirming the hypothesis

  4. 4

    At what value the work stops

  5. 5

    What the person who will use the solution gets out of it

  6. 6

    Who clears exceptions, and how many hours a week that takes

None of these questions is about technology, and all six are answered before development begins. That is precisely the part of the work that is easiest to skip and most expensive to have skipped.

Methodology notes, by e-mail

The same material we publish here: how to model the economics of an initiative, where rollouts break, and what to verify before work starts. Once a month at most, no market news and no sales e-mail.

By submitting you agree to the privacy policy — privacy

More reading

  • A pricing worksheet with rows for materials, labour and expenses, a hand and pen over it

    Modelling the economics of an AI initiative before you start it

    A baseline cannot be reconstructed after the fact. What to measure before work starts, three ways to model the effect, and the cost line that gets forgotten most often.

    Aplora5 min read
    Read
  • A rep in a headset writing notes in a notebook during a call

    What a sales book needs before it can be automated

    The difference between a description and a criterion, why “build rapport” cannot be checked, and what a wording looks like when the system and the sales lead reach the same conclusion.

    Aplora3 min read
    Read

Let's apply this to your own workflow

Thirty to forty-five minutes on one specific task: clarify the problem, judge the fit, and name the next justified step.