All posts

Most AI pilots do not fail. They just never ship.

08 Jul 2026AIDelivery3 min read

An engineer with a good model and a weekend can build something that makes a leadership team lean forward. Nothing about that experience prepares anyone for the eleven months it then takes to put the same capability in front of customers. And because the demo was so convincing, the eleven months reads as failure rather than as the actual cost.

The demo is not a lie. It is just a different artefact from a system, and the distance between them is where almost every AI initiative quietly stops.

What the demo skips

A demo answers a question you chose, from data you prepared, for a user who is you. A production system answers questions you did not anticipate, from data that is wrong in ways nobody documented, for a user who will screenshot the bad answer.

Between the two sit the parts nobody demos: access control, so the system cannot tell one department what another department earns. Evaluation, so you can tell an improvement from a regression. Cost modelling, so the unit economics survive contact with real usage. Failure semantics, so a partial answer is labelled as one. Audit trails, so a regulated customer can adopt it at all.

None of that is glamorous. All of it is the actual work.

The part that stops projects dead

On Matrix the blocking constraint was not technical. Everybody agreed that pointing a language model at enterprise data would produce enormous value. Everybody also agreed that a CISO would refuse it, because it meant enterprise records leaving the estate to be processed by someone else's model.

That refusal is where most enterprise AI pilots die: not at the demo, and not at the business case, but at the security review that was never in the plan. So on Matrix it went first. Before any feature was built, the boundary was drawn: what may cross the perimeter, what may not, and what that rules out. The architecture follows that line, which is why it has to be established before there is an architecture to change.

The result is a closed loop. Querying, planning, document generation and agent execution all happen with the customer's data inside the customer's infrastructure. That constraint cost real capability, and it is the only reason the thing is adoptable.

Measure before you generate

The second rule we hold to is that the evaluation harness comes before the features it measures.

A system that answers confidently and wrongly is worse than no system, because people act on it. On Matrix the harness was built first and stayed a human responsibility throughout. On ContentMorph we assembled a set of real content with human-judged good and bad outputs before writing the generator, which is the only reason we could tell tuning from guessing when the output started improving.

This is the least popular advice we give, because measurement feels like overhead when the demo already works. It is also the difference between a system that gets better every week and one that gets changed every week.

We do not run pilots

If a capability is worth building, we scope it as a production system from the first conversation, with access control, evaluation and cost modelling in the scope, because those are the parts that determine whether it ships.

The organisational problem is harder than the technical one. Pilots are funded as experiments and production systems are funded as programmes, and nobody wants to be the person who says the afternoon demo needs a quarter and an owner. The honest version of that conversation is shorter than the eleven months.

Create 10x more value with technology.

Your roadmap shouldn't be waiting.

30-minute call  ·  No deck  ·  Engagement lead on the call