AI transformation that ends inside your building.
Most AI programmes stall in the same place: a pilot everyone liked, and a production system nobody will approve. The gap was never the model. It is everything around the model — retrieval over your own data, agents that act under budget, a privacy rail, and a record an auditor accepts.
We build that layer and install it on your infrastructure. You own what comes out.
30 minutes. We size your real workload and tell you plainly if owning is the wrong call.
Four failures, and none of them are the model’s fault.
The pilot that cannot be promoted
It worked on a curated folder. Production means reading the real estate — permissions, retention and all — and nobody will sign that off.
Answers nobody can check
Fluent and uncitable. A confident answer with no source is worse than no answer, because someone will act on it.
Intelligence you rent
Every prompt, document and correction improves a system you do not own. Stop paying and the accumulated context leaves with the contract.
Shadow AI
Staff are already pasting internal material into consumer tools. The choice was never "AI or no AI" — it was governed or ungoverned.
None of these are solved by a better model. They are solved by architecture — the part most AI programmes never buy.
A model is a component. The system around it is the product.
India’s AI conversation is model-first — foundational models, benchmarks, parameter counts. That work matters, and it is national work worth backing.
But no organisation can deploy raw weights. A hospital needs retrieval that inherits its access model. A bank needs an action to be budgeted, scoped and logged. A ministry needs an answer to carry a citation. All of that lives in the layer around the model — and that layer is what survives an inspection.
BiltIQ builds the layer. The models inside it are open-weight today and domain-tuned over time. They are components in our system, not our pitch.
Everyone else sells you a model. We install the plant.
Checkable, not claimed.
Winner — NHA PM-JAY AI Challenge 2026
Problem Statement 2, IISc Bangalore finale, 8 May 2026. Five-agent claims analysis returning in 25–40 seconds; 40 fraudulent claims identified in 82 minutes.
Finalist — HIMSS 2026 Emerge, Las Vegas
Clinical AI work.
ISO/IEC 27001:2022 Certified
Cert 305026060958IS, valid to 8 June 2029. ISO 9001:2015 — 305026080584Q.
4 patents filed
Complete specifications, Patent Office Kolkata. Numbers listed in the stack section below.
NVIDIA Inception Partner
Since 2025. DPIIT Recognised Startup — DIPP239966. Udyam MSME UDYAM-JH-06-0038439.
Open source you can audit
biltiq-privacy (MIT, PyPI) · ManthanQuant · upstream vLLM patches.
The team
10 engineers, 80+ combined years. Forward-deployed engineering — we build inside your environment.
Papers and deep dives
The thesis paper — Intelligence Over Your Data — plus the deployment anatomy, the multi-agent architecture, and the sector papers, open in full. Start with the thesis paper or browse the resources.
Five stages. Each one ends in something you own.
01 Establish the floor
Size the real workload — query volume, context length, concurrency. Decide honestly whether owning hardware is right for you at all.
You own at the end: A sizing model and a go / no-go you can defend to finance.
Engagement: AI Factory Advisory
02 Ground the estate
Connect the document estate. Parse it, chunk it layout-aware, and — the hard part — inherit the source system's permissions into the index.
You own at the end: A retrieval index over your corpus, access-faithful.
Engagement: Structured Data for AI · ATC Manthan
03 Answer, with citations
Grounded question-answering in front of the people who need it. Every answer carries its source.
You own at the end: Cited answers over your knowledge, on your hardware.
Engagement: Document & Conversational AI
04 Act, under budget
From answers to actions. Agents execute scoped, budgeted, logged steps — with circuit breakers, not open-ended autonomy.
You own at the end: Workflows that execute, with an audit trail per action.
Engagement: Agentic AI · ATC Flow
05 Scale and own
Extend across functions on the same core. Connect internal systems over MCP. Models, index and memory stay in-house.
You own at the end: A platform, not a subscription.
Engagement: AI Orchestration · ATC Connect
Deployment runs in days, not quarters — because the stack is built once and configured per organisation, not rebuilt each time.
Four layers, one request path, no bypass.
Inference
Open-weight models served by vLLM on your own GPUs. A model registry maps every endpoint by role and pins versions; routing sends each task to the right-sized model rather than the largest one.
You get: Predictable cost; no vendor swapping the model under you.
Retrieval — Manthan
30+ document parsers, layout-aware chunking. Hybrid retrieval — BM25F lexical plus dense vector, fused by reciprocal rank fusion — re-scored by a multimodal reranker. Five-layer memory: working, session, project, knowledge graph, immutable audit trail.
You get: Answers from your documents, with page-level citations.
Agents — Agent OS
Agents defined in YAML running ReAct loops under explicit step and token budgets with circuit breakers. 30+ built-in tools plus MCP tool access. Agent-to-agent delegation signed with Ed25519. Per-project scope isolation.
You get: Actions that are bounded, attributable and reversible.
Privacy rail
Presidio plus NER detecting Indian identifiers (Aadhaar, PAN, mobile) alongside international PII. HMAC pseudonymisation keeps records linkable without exposing originals. SHA-256 tamper-evident hash chain over every decision.
You get: Your security team reads the code instead of trusting a description.
Compliance mode on_prem_required: fail-closed, no cloud egress path, enforced in code. A misconfiguration cannot silently transmit data, because the failure mode is refusal.
Why you can check this yourself
- biltiq-privacy — the privacy rail is published as an MIT-licensed library on PyPI. Read the recognisers rather than accept a description of them.
- ManthanQuant — 3-bit Lloyd-Max KV-cache compression: 5.12× at 0.983 cosine similarity.
- vLLM patches upstream — including PR #38826 backport. Our GB10 serving work is in public repositories.
- Four patents filed (complete specifications, Patent Office Kolkata): trajectory-aware KV-cache compression 202631090661 · cryptographically verified fail-closed sovereignty boundary 202631094208 · deterministic effectivity-aware answer validation with provenance-bound tool invocation 202631094273 · tenant-isolated model resolution with policy-preserving egress-gated fallback 202631094336.
One core. Ten applications. All on your infrastructure.
| Function | Product | On-premise, it… |
|---|---|---|
| Knowledge & legal | ATC Manthan | Ask contracts, records and drawings real questions; get answers cited to page level. |
| Learning & workforce | ATC Quest LMS | Regenerates courses when a regulation or SOP changes, SME-approved. 22 Indian languages. |
| Higher education | ATC Campus | Learning, research and admin AI grounded in institutional data, run on campus. |
| Content & comms | ATC CMS | Draft, manage and publish with agents on your infrastructure. |
| Marketing | ATC Social | Schedule, respond and analyse with your brand voice kept in-house. |
| Contact centre | ATC Voice | Agents that answer, qualify and route calls in Indian languages on your hardware. |
| Support & internal help | AI Chatbots | Grounded chat over your documents — cited answers, no cloud dependency. |
| Operations | ATC Flow | Agents executing multi-step processes under governance and human checkpoints. |
| Platform engineering | DevOps Tools | Provision GPUs, serve models, monitor the stack without outsourcing operations. |
| Integration | ATC Connect | Connects agents to internal systems over MCP, with signed auditable hand-offs. |
These are ten applications of one core, not ten products stitched together. That is why the second deployment in an organisation costs less than the first.
Sector detail: Healthcare · BFSI · Government & PSUs · Education · Manufacturing · SMBs — or see the AI Factory platform itself.
We will tell you not to buy this.
Below roughly 90,000 queries a month, no on-premise deployment pays back inside three years — including the cheapest one we build. That figure ignores electricity and staff time, so the real floor is higher.
A team of twelve running about 400 queries each per month spends roughly ₹1,200 a month on frontier API access — about ₹43,000 over three years, against ₹8.1 lakh of hardware. They should not buy hardware, and we say so.
Where it does pay back
| Deployment | Sustained volume | Crossover |
|---|---|---|
| Entry tier | from ~0.4 M queries/month | 8–14 months |
| Mid-enterprise reference fleet (₹41.7 L capex) | 1 M queries/month, 700-token contexts | 18 months |
| Same fleet | 1.5 M queries/month | 12 months |
| Same fleet | 2 M queries/month | 9 months |
What moves the crossover more than hardware price does — context length. Longer contexts, premium model tiers, and agentic workloads that chain several model calls per user action all pull crossover earlier, because they raise token throughput without raising your capex. RAG and agentic work run long contexts by construction. That is the actual commercial argument.
The running-cost comparison, measured on our own cluster — not modelled
| Draw | Annual electricity | |
|---|---|---|
| Our reference fleet, typical | ~300–500 W | ~₹35,000 |
| Our fleet, worst case 24/7 | ~1.7 kW | ~₹1.5 L |
| One 8×H100 cloud-class server | 10.2 kW | ~₹8.9 L plus cooling |
A 240 W node is cooled by office air conditioning. No server room, no chiller, no water.
We run your numbers, including the case where the answer is no.
You are choosing the architecture you will be audited on.
The DPDP Act’s substantive obligations commence on 13 May 2027. An organisation procuring AI infrastructure now is choosing the architecture it will be audited on, before the audit exists. That is a scheduling fact, not a scare.
DPDP Rule 6 requires reasonable security safeguards — encryption and masking of personal data, access control, access logging and monitoring, backups, and retention of access logs for at least one year. Breach reporting runs on a 72-hour clock with no materiality threshold: every personal-data breach is reportable. How the duties map to architecture.
Three corrections we publish against our own interest
- DPDP does not require your data to stay in India. §16 is a restriction list, not a general prohibition. Hard localisation for Indian financial data comes from the RBI, not from DPDP. Anyone claiming otherwise is wrong, and your DPO will know it.
- DISHA imposes no obligation on anyone. It has never been enacted.
- HIPAA does not yet mandate encryption or MFA. Those are proposals in the January 2025 NPRM, not current law.
The argument that survives is the stronger one. DPDP Rule 6 is an evidenceobligation. You cannot discharge it with a vendor’s promise; you discharge it with a record. That is the case for architecture over assurance, and it needs no overstatement.
Frequently Asked Questions
What is AI transformation, in practical terms?
Moving from pilots to systems your organisation can operate and evidence: retrieval over your own data, agents taking bounded actions, a privacy rail, and an audit record — installed on infrastructure you control.
How long does deployment take?
Days, not quarters. The stack is built once and configured per organisation rather than rebuilt each time. Sizing and data-estate work happen first and depend on the state of your document estate.
Do we have to replace our existing systems?
No. ATC Connect exposes your internal systems and data sources to agents over MCP, so the platform reaches into what you already run.
Is this ready for DPDP obligations?
It is built for them by architecture: fail-closed on-premise mode with no egress path, access logging, and a tamper-evident audit chain over every decision. DPDP's substantive obligations commence 13 May 2027, and Rule 6 requires access logs retained for at least a year — an evidence obligation you meet with a record, not a promise.
When is on-premise AI the wrong choice?
Below roughly 90,000 queries a month, no on-premise deployment pays back inside three years — including the cheapest one we build. At that volume you should use a cloud API, and we will tell you so.
Can we still use frontier models like GPT or Claude?
Yes, by policy rather than by default. Commercial APIs are policy-gated and default-off. Sovereign by default, frontier by choice — the routing decision is yours and it is logged.
Go deeper on any stage.
- What AI transformation actually means
- Why AI pilots fail to reach production
- The five-stage transformation roadmap
- Build, buy or rent enterprise AI
- Transformation ROI, and when not to buy
- What DPDP 2027 means for your architecture
- Shadow AI: what your staff are already doing
- From answers to actions: enterprise AI agents
- AI transformation in regulated sectors
- How to staff an AI transformation
Start with the sizing, not the software.
A 30-minute architecture consultation. We size your real workload, tell you where the crossover sits at your volume, and tell you plainly if owning is the wrong call.
