← All writing
Enclave · Writing

Why Your AI Pilot Never Made It to Production

The demo worked. Six months later nothing is live. The reasons are boringly consistent, entirely predictable, and almost never about the model.

The pattern is so consistent it has become a genre. A team builds something in a fortnight. It demos beautifully. Everyone in the room can see the value. Budget appears. And then, six or nine months later, there is still nothing in production, and nobody can quite explain where the time went.

We can explain where the time went. It went into the seven things that were never on the plan, because the plan was written by people who had only ever seen the demo.

The demo and the system are different products

A demo has one user, one file, one happy path, and a friendly operator who knows which questions work. A production system has every user, every file, every path, and a stranger typing something nobody anticipated on a Tuesday afternoon.

Those are not the same product with different amounts of polish. They have different requirements, different failure modes, and different costs. The demo is a proof that the capability exists. It is not a proof that the system is close, and treating it as one is the single most expensive mistake in this field.

Seven reasons, in rough order of frequency

1. The data was not ready. It lives in six systems. Two have no usable API. The one that actually matters is a shared drive of scanned PDFs with a naming convention nobody has understood since 2014. Every AI timeline that skipped this discovery is now a data-engineering timeline wearing a disguise — and data engineering is slower, less glamorous, and much harder to get funded a second time.

2. Nobody defined correct. Without an evaluation set, "it seems better" is your only quality signal. That is fine for a demo and fatal for a system. You cannot tune what you cannot measure, you cannot defend a change you cannot justify, and every prompt edit becomes a coin flip that nobody is allowed to lose. Teams in this state ship nothing, because shipping requires someone to say the quality is good enough, and no one can.

3. Ninety-two percent turned out not to be a product. A model that is right nine times in ten makes a spectacular demo and an expensive support queue. Essentially all of the product work lives in the tenth case: detecting that the system is unsure, containing the damage, and routing the case to a human before it reaches a customer. That work is invisible in a demo and is most of the build.

4. Permissions were the actual project. The prototype indexed everything, because indexing everything is the fastest way to a good demo. Then someone asked whether a contractor could retrieve the compensation review, and the honest answer was yes. Retrofitting per-user authorization into a working prototype is usually a rewrite, not a patch, because the authorization model has to sit at query time and the prototype has no concept of who is asking.

5. The bill arrived. Tokens, times real volume, times retries, times all the context quietly re-sent on every call. Almost no pilot runs that arithmetic, because at pilot volume the number is a rounding error. At production volume it is a line item, and the conversation stops being about capability and starts being about whether this was ever worth doing.

6. The workflow never changed. Bolting a model onto a broken process automates the broken process, faster and at higher cost. A large share of the value in any credible AI project is a process redesign wearing an AI costume — and that part is organisational work, not engineering. It needs someone who can have the conversation about why step four exists at all.

7. It only worked for the person who built it. Adoption is not a launch email and a wiki page. If the new way is not visibly faster than the old way on the first attempt, on a normal day, under normal pressure, people go back to the spreadsheet within a week. Quietly, and without telling you.

Notice what is not on the list

The model. Not once.

That is the uncomfortable part. Almost every failed AI project failed at integration, evaluation, permissions, cost, or adoption — and almost every one of those failures was knowable in advance. None of them required a better model to avoid. Several of them get worse when you upgrade the model, because a more capable system reaches further into your data and fails in more interesting ways.

What actually changes the outcome

Two things, and the order matters.

Build the evaluation set before the system. A few hundred real cases with the right answer attached, drawn from work your team has already done. It feels like a delay. It produces nothing you can demo. It is also the only thing that converts "does this seem better?" into a number, and everything downstream — model choice, prompt changes, the decision to ship — depends on having that number. Teams that skip this step do not save two weeks; they lose the ability to make any decision with confidence for the rest of the project.

Then ship the narrowest possible slice, end to end, into production. Real data, real permissions, real users, behind a feature flag. Not a notebook, not a sandbox, not a broader prototype. A narrow slice in production surfaces the expensive problems in the first fortnight, when they are still cheap to respond to. A broad prototype surfaces them in month five, when the budget is committed and the roadmap is public.

Everything else is downstream of those two decisions.

The uncomfortable question worth asking early

Some of what gets scoped as an AI project is a reporting problem, a data-quality problem, or a process problem that a model will make faster and no better. It is worth asking, out loud, in week one, whether that is what you have. The answer costs nothing then. In month four it costs a year.


If you have a pilot that stalled somewhere in this list, that gap is the whole practice here. Book a thirty-minute scoping call — you describe the workflow, and you will leave knowing whether it is viable and what it would genuinely take.