savas.one Back to writing

Enterprise AI

Why do enterprise AI projects stall at the pilot stage?

3 min read

An AI pilot can produce impressive results in a few weeks. Move the same work into an environment with real users, real data and real accountability, and the picture changes. A pilot looks at technology; a product must look at the whole system.

A small glowing pilot cube on a copper base and a running production system, joined by a lit bridge that stops halfway
01

Pilots often answer the wrong question

Pilots usually ask whether the model can perform a task. Production asks different questions: who owns the output, how errors are detected, when users should not trust the system, and what happens when source data changes?

Technical capability is necessary, but insufficient for a product decision. A good pilot reduces uncertainty in the next decision rather than merely showcasing capability.

02

What happened to us on a contracts project

When we first integrated AI models into our systems, the pilot went really well. We gave it the documents of a few contracts, the model read them through RAG, reviewed them and came back with meaningful interpretations. It pulled the right information out of the contracts. A job that took 2-3 days in the pilot got one week in UAT; we thought it would be done and we would go live right after.

Once the work grew in UAT, problems showed up. Some contracts were Word files, some were signed PDFs. Some were long-running deals continued through renewals, and the information they depended on was not in the shared folders. Because the model has to give some answer, it started guessing wherever information was missing. When we began checking results one by one, it became clear this was not that easy: garbage in, garbage out.

We extended UAT and decided to clean the data first, using the models for that as well. We collected contract details and files from the shared folders, checked whether they met our criteria and cleaned them up. We revisited the RAG architecture and chunk settings, and built one process for PDFs and a separate one for Word, Excel and similar documents. It took a long time, but in the end we had very well prepared data. Going live was honestly the easiest part.

03

Ownership cannot be added later

The model produces an answer, but it does not own the business outcome. When the boundaries between process, data, engineering and operations remain vague, progress slows as the system approaches production.

A product definition therefore needs approval, feedback, monitoring, escalation and human handoff—not only features.

I spent years on the systems and application side. That is where I learned that everything that goes live will one day get someone called at midnight. AI is no different: a source document changes and the answers break. Who notices, who fixes it, who tells the users? Without an answer, the pilot ends but the product never starts.

04

Measurement is broader than model quality

A few successful examples may sell a demo. A product needs continuous measurement across answer quality, latency, cost, attribution, user behaviour and operational failure.

When success is not defined up front, the pilot may be declared successful without anyone knowing what should scale.

Think of a simple example: an internal support assistant. In the pilot you ask "are the answers correct?" In production the real questions are different: how many questions were handed to a person, what did users do when they did not like an answer, what did it cost this month, how many seconds did an answer take? If you do not collect these numbers from day one, all you have at the end of the pilot is an impression.

05

A pilot should end with a decision

Every pilot should lead to one of three decisions: scale, change or stop. “Let us experiment a little longer” often signals an unclear product question rather than a technical gap.

Enterprise AI moves beyond the pilot through a clear problem, visible ownership, measurable success and an operable product model—not only a better model.

The lesson I took from the contracts project: before a pilot starts, write down which problem you are solving, which number will measure success and who makes the call when the pilot ends. And one more question: does the production data look like the pilot data? Ours did not, and we learned it in UAT.