Skip to main content
BiltIQ AI logoBiltIQ AI logo
Back to Blog
Enterprise AI

Why AI pilots fail to reach production

The pilot worked and production stalled. Five gates decide whether an enterprise AI project ships — and four of them have nothing to do with the model.

BiltIQ AI
6 min read

There is a pattern common enough to be predictable. A team builds a proof of concept in a
few weeks. It answers questions about a folder of documents impressively well. Leadership
sees it, likes it, and asks for production. Then the project stops — not with a rejection,
but with a series of meetings that never resolve.

The reason is that a pilot and a production system are asked different questions. A pilot
is asked can this work? Production is asked can this be operated, evidenced and
defended?
Five gates stand between them, and only one is about the model.

Gate one: permissions

This is the gate that stops the most projects, and it is the least anticipated.

A pilot runs against a curated folder. Someone chose the documents, so everything in the
index is something everyone in the pilot is allowed to read. Production means indexing the
real estate — a file share with twenty years of accumulated structure, a document
management system with inherited permissions, a mail archive nobody has audited.

The hard part of ingestion is not parsing. It is permissions. An index that does not
inherit the source estate's access model will surface documents to people who could not
open them in the original system. Retrieval does not respect a folder ACL unless you build
it to.

What makes this dangerous rather than merely difficult is that the failure is invisible
in testing.
Pilot users are usually senior and have broad access, so nothing looks wrong.
The defect surfaces months later when a junior analyst receives a passage from a document
they had no right to see — and by then the system is in front of the whole organisation.

Any serious AI transformation treats permission inheritance as a
first-class requirement in the architecture, not a hardening task scheduled for later.

Gate two: citation

A pilot is judged on whether the answer sounds right. Production is judged on whether the
answer can be checked.

These are different systems. A model asked to summarise will produce a fluent paragraph
whether or not the underlying passages support it. Without a citation path — answer to
passage, passage to page, page to document — a reviewer has no way to verify anything, and
the burden of correctness quietly transfers from the system to whoever reads it.

That transfer is what compliance functions object to, usually without articulating it in
those terms. What they are saying is: if this is wrong, who finds out, and when?

There is a related failure worth naming, because it is subtle. Aggregation questions are
not retrieval questions.
"How many of our contracts have an uncapped indemnity" requires
every contract to be examined, not ten relevant passages to be found. A retrieval system
will happily answer from ten passages and produce a confident number that is simply wrong.
That is the specific failure mode to test for before production, and it is not a model
quality issue — it is a question-routing issue.

Gate three: the operating model

A pilot has an owner who cares. Production needs an operating model.

Who retrains when the corpus drifts? Who reviews flagged outputs? What happens when a
document is retracted — does the index forget it? What is the escalation path when the
system gives a wrong answer to a customer-facing team?

These questions are unglamorous and they are usually the reason a project sits for a
quarter. Nobody rejects it. It just has no owner, no budget line and no runbook, so it
stays a pilot indefinitely. The organisations that get through this gate decide the
operating model before the build, not after.

Gate four: the evidence obligation

This gate is new and it is dated.

The DPDP Act's substantive obligations commence on 13 May 2027. Rule 6 of the DPDP
Rules 2025 requires access control, access logging and monitoring, and retention of access
logs for at least one year. Breach reporting runs on a 72-hour clock with no
materiality threshold
— every personal-data breach is reportable.

Read that as an engineering requirement rather than a legal one and its significance
changes. It is an evidence obligation. You cannot discharge it with a vendor's assurance;
you discharge it with a record. A system that cannot produce a per-decision log covering
who asked, what context was retrieved and what was returned is not going to clear this
gate, regardless of how good its answers are.

For regulated sectors the constraint arrives earlier. RBI's payment-data localisation
circular requires end-to-end payment transaction data to be stored only in India, and the
Master Direction on Outsourcing of IT Services means a cloud AI API used on regulated
data is an IT outsourcing arrangement
— with audit rights and supervisory access
attached.

One correction, because it is repeated constantly and it damages credibility with the one
person who checks: DPDP does not impose a general data residency requirement. Section
16 is a restriction list, not a prohibition. Hard localisation for financial data comes
from the RBI.

Gate five: the economics, honestly assessed

The last gate is arithmetic, and it is the one where a vendor should be willing to lose the
deal.

Below roughly 90,000 queries a month, no on-premise deployment pays back inside three
years — including the cheapest one available.
That calculation ignores electricity and
staff time, so the real floor is higher. A team of twelve running around 400 queries each
per month spends roughly ₹1,200 a month on frontier API access, or about ₹43,000 over three
years, against ₹8.1 lakh of hardware. That organisation should not buy hardware.

Above the floor the picture inverts, and volume decides how fast. A mid-enterprise
reference fleet crosses over at 18 months at one million queries a month, 12 months at 1.5
million, and 9 months at two million. What moves the crossover more than hardware price
does is context length
— long contexts, premium model tiers and agentic workloads that
chain several calls per user action all pull it earlier, because they raise token
throughput without raising capital cost.

The failure here is not choosing wrong. It is choosing without knowing the volume, then
discovering the payback assumption was fictional a year in.

Getting through

None of these five gates is about model quality, and four of them can be designed for at
the start rather than discovered at the end. Permission inheritance, a citation path,
an operating model, a per-decision audit record, and a sizing exercise honest enough to
say no.

Projects that build those first are unglamorous for a month and then ship. Projects that
build a chatbot first demo beautifully and stall at exactly the point this article
describes.


Next: The five-stage transformation roadmap
· AI transformation ROI, and when not to buy

👨‍💻

BiltIQ AI

Expert team at BiltIQ AI providing cutting-edge AI solutions.

Contact our team →
Share this article:

Book an Architecture Consultation

30 minutes. No sales pitch. We assess your current stack, identify where agentic AI creates measurable value, and give you a concrete deployment path — with timelines and costs.

Your Data. Your Premises. Your AI.