Skip to main content
BiltIQ AI logoBiltIQ AI logo
Back to Blog
Enterprise AI

AI transformation ROI, and when not to buy

How to calculate the return on an on-premise AI deployment, which variable actually moves the crossover, and the volume below which nothing pays back.

BiltIQ AI
6 min read

Return-on-investment models for enterprise AI have a credibility problem: they are usually
built by the party selling the outcome. This one is built the other way round, starting
with the conditions under which the purchase is wrong — because those conditions are real,
they are knowable in advance, and a model that cannot produce a "no" is not a model.

Start with the floor

Below roughly 90,000 queries a month, no on-premise deployment pays back inside three
years — including the cheapest one we build.
That calculation ignores electricity and
staff time entirely, which biases it in favour of owning. The real floor is therefore
higher than 90,000.

Computed against an entry-tier deployment at ₹8.1 lakh of capital:

Queries / month Cloud spend 3-year cloud cost Break-even
5,000 ₹1,232/mo ₹0.44 L never — 55 years
25,000 ₹6,160/mo ₹2.22 L never — 11 years
50,000 ₹12,320/mo ₹4.44 L never — 5.5 years
~90,000 ₹22,500/mo ₹8.1 L 36 months — the floor
137,000 ₹33,750/mo ₹12.2 L 24 months
274,000 ₹67,500/mo ₹24.3 L 12 months
500,000 ₹1.23 L/mo ₹44 L 7 months

The worked example worth internalising: a team of twelve running about 400 queries each per
month spends roughly ₹1,200 a month on frontier API access — about ₹43,000 over three
years
, against ₹8.1 lakh of hardware. They should use a cloud API. Anyone selling them
infrastructure is selling them a mistake.

Then the crossover, with the volume attached

Above the floor, the question becomes how fast. There is no single crossover figure, and
any vendor quoting one without a volume is quoting a scenario without telling you which.

Crossover is throughput divided by capital. For a mid-enterprise reference fleet at ₹41.7
lakh:

Sustained volume Crossover
1 M queries/month (700-token contexts) 18 months
1.5 M queries/month 12 months
2 M queries/month 9 months

An entry deployment crosses at around four months, because the hardware costs a tenth as
much.

You will see the band "8–14 months" quoted widely. It is real, but it is a consequence of
volume
, not a property of on-premise AI. It holds at entry tier from about 0.4 M
queries/month upward, and for the reference fleet once sustained volume exceeds roughly
1.3 M queries/month. It does not hold for the reference fleet at 1 M queries/month on
700-token queries — that is 18 months — and it does not hold at any volume below about
0.5 M queries/month at reference scale.

The rule this implies is simple and worth applying to any vendor's numbers, including ours:
a crossover figure without a volume attached is not a figure.

The variable most models get wrong

Hardware price is the number everyone negotiates. It is not the number that decides the
outcome.

Context length moves the crossover more than hardware price does. Take the reference
fleet at one million queries a month and vary only the average context:

  • 300 tokens — never pays back inside 36 months
  • 700 tokens — 18 months
  • 1,500 tokens — 8 months
  • 4,000 tokens — 3 months

Same hardware, same query count, same capital. The only variable is how much context each
query carries, and it swings the answer from "never" to "one quarter".

This matters because it is not a tuning parameter you choose — it is a property of the
workload. Retrieval-augmented and agentic workloads run long contexts by construction:
retrieval stuffs passages into the prompt, and an agent chaining several model calls per
user action multiplies token throughput without touching your capital cost. An organisation
doing serious RAG and agent work is structurally on the favourable side of this curve; an
organisation doing short-form classification is not.

Which leads to the practical instruction for building your own model: measure average
context length before you measure anything else.
Most business cases we see estimate
query volume carefully and assume context length, which is exactly backwards.

Running cost, measured rather than modelled

Capital is only half the comparison, and the operating side is usually estimated badly in
both directions.

Draw Annual electricity
Reference fleet, typical ~300–500 W ~₹35,000
Reference fleet, worst case 24/7 ~1.7 kW ~₹1.5 L
One 8×GPU cloud-class server 10.2 kW ~₹8.9 L plus cooling

A 240 W node is cooled by ordinary office air conditioning. No server room, no chiller, no
water — which also removes a facilities project that is often larger than the hardware
purchase.

On a per-token basis the comparison lands at roughly ₹15–20 per million tokens for
owned inference, against $2–15 per million on commercial APIs depending on tier. The spread
is wide because the API side is wide; the point is not the precise ratio but that the owned
cost is fixed and the rented cost scales with adoption.

What belongs in the model that usually is not

Staff time. Someone operates this. If your model shows payback in nine months and
assumes zero operational headcount, it is wrong.

The cost of not knowing. DPDP Rule 6 requires access logs retained for at least a year,
with breach reporting on a 72-hour clock and no materiality threshold. Building that record
later, retrofitted onto a system that was not designed to produce it, is more expensive
than building it in. That is a real line item, even if it is hard to price precisely.

Residual value. Hardware has some at year four. Rented capacity has none, ever.

What you should not put in: productivity gains you cannot measure. A model that leans
on "40% faster document review" without a baseline measurement will not survive contact
with a finance function, and including it damages the credibility of the lines that are
solid.

The honest summary

Owning infrastructure pays when three things are true together: sustained volume above the
floor, context lengths that are long by workload construction rather than by hope, and data
whose accumulated context is strategic enough that renting it back indefinitely is the
wrong trade.

When those are true, the case is strong and does not need overstatement. When they are not,
the correct answer is a cloud API — and knowing which situation you are in is the entire
purpose of the first stage of any AI transformation worth running.


Next: Build, buy or rent enterprise AI
· What DPDP 2027 means for your architecture

👨‍💻

BiltIQ AI

Expert team at BiltIQ AI providing cutting-edge AI solutions.

Contact our team →
Share this article:

Book an Architecture Consultation

30 minutes. No sales pitch. We assess your current stack, identify where agentic AI creates measurable value, and give you a concrete deployment path — with timelines and costs.

Your Data. Your Premises. Your AI.