Skip to main content
BiltIQ AI logoBiltIQ AI logo
Hardware + Software · One On-Premise Stack

The BiltIQ AI Factory

Your Data. Your Premises. Your AI.

Private, full-stack enterprise AI — agents over your own CRM, documents, and support data — running on office-grade hardware, with nothing leaving the building.

50–500×
Cheaper / token vs cloud API
240 W
Per node — no data center
~500 GB
On-prem model memory
Zero
Cloud · water · data egress
DPIIT RecognisedNVIDIA Inception PartnerDPDP · HIPAA Aligned
See the AI Factory on Your Premises
30-min call · We map your workload and recommend the right on-premise stack.
Your data stays private. We never share your information.
The Reframe

India’s AI conversation is about models. The AI Factory is about the layer around them — the operating system that turns a model into a system you own and that can be inspected. The model is the easy part; this is the rest.

01Why Now

Token subscriptions are an unbounded tax on your own success

Enterprises have realized that token subscriptions are an unbounded tax on their own activity — the more their staff use AI, the more they pay, forever — and that their operational alpha (workflows, customer knowledge, institutional memory) should compound inside their own walls, not inside a vendor relationship. Frontier labs promise contractually that they won’t train on your data; we make the question irrelevant. When the models, the retrieval index, and the memory all live on your premises, there is nothing to trust and nothing to audit remotely: your data cannot train anyone else’s model because it never leaves your building.

Your intelligence stays yours, and it compounds for you.

Companies adopting AI hit the same three walls: runaway API token bills, per-seat pricing that scales with headcount, and customer data flowing to third-party clouds. Yet the work they actually need — assistants grounded in their own CRM, tickets, and documents — doesn’t require frontier-scale models. When retrieval (RAG) supplies the facts and MCP/tools supply the actions, a well-served 12B–35B open model does the job frontier APIs are sold for.

02The Platform

Two layers, one stack

We assess the workload, specify and install the right hardware fleet on your own premises, and layer software on top that turns it into a working AI factory — grounded in your data, safe by construction.

LAYER 1 — HARDWARE · REFERENCE FLEET LIVE

A data center that isn’t one

Five nodes on a standard office LAN, ~500 GB of total model memory, managed by our own Manthan cluster manager (auto-discovery, web-UI model lifecycle, self-healing systemd + watchdog, per-model ROI analytics). This is the reference fleet we run ourselves — proof of what office-grade hardware does; your fleet is sized to your workload. See the AI Factory Advisory tiers.

NodeHardwareServes
3× NVIDIA DGX Spark (GB10)128 GB unified each · 240 W maxDeep reasoning at long context · multimodal + voice · embeddings + reranking
1× RTX PRO 6000 Blackwell server96 GB + RTX 4070Full multimodal (text/image/audio/video), streaming ASR
1× RTX 5060 laptop8 GBRoaming demo node

Inference-time engineering — speculative decoding (~88% acceptance), KV-cache compression, MoE expert offloading, prefix caching — is what makes 240 W boxes serve at interactive speed.

Right-sized by design

The unit of deployment is not a data center — it’s a box on a desk. A single node comfortably serves ~5 concurrent requests: enough for a 20–30 person office, or a small hospital where 3–4 doctors are processing patients at the same time. Start with one node for a department, add nodes as demand grows — the same management plane runs one box or a fleet. Each box is silent, plugs into a normal wall socket, needs no water and no special cooling — the office AC you already run is the entire thermal plan.

ScenarioFleet drawAnnual cost
Typical serving~300–500 W~₹35,000 / yr (~$420)
Worst case, 24/7~1.7 kW~₹1.5 L / yr (~$1,800)
One 8×H100 server, for contrast10.2 kW~₹8.9 L + chilled-water cooling

Marginal cost: ~₹15–20 of electricity per million tokens, vs $2–15 per million on APIs. The whole fleet’s heat load is under one domestic AC unit. No liquid cooling, no raised floor, zero water (AI data centers evaporate ~1–2 L of water per kWh; we evaporate none).

LAYER 2 — SOFTWARE

The BiltIQ Platform

Three cores in one monorepo, compliance mode on_prem_required — no cloud egress, fail-closed, enforced in code, not policy.

MANTHAN · RAG + MEMORY

30+ format parsers (PDF/DOCX/XLSX/PPTX/images/nested ZIPs) with layout-aware chunking; hybrid retrieval (BM25F + dense + RRF, re-scored by a multimodal reranker) so answers are grounded with citations; five-layer memory from working memory through knowledge graph to an immutable audit trail.

AGENT OS

Teams of AI specialists defined in YAML, run in ReAct loops with budgets and circuit breakers; 30+ tools plus MCP (files, browser, code, vision, TTS/STT, SSH, IoT); CRM/ERP connectors syncing business systems into isolated per-project scopes; delivery where staff already work — Slack, WhatsApp, Teams, Telegram, email; Ed25519-signed agent-to-agent delegation; multi-tenant isolation at five layers.

PRIVACY FILTER

Presidio + NER detection extended for Indian identifiers (Aadhaar, PAN, mobile); HMAC pseudonymisation so originals never reach a model but records stay linkable; DPDP 2023, GDPR, HIPAA, CCPA validators scoring every document; a tamper-evident SHA-256 audit hash chain. Published as an MIT library on PyPI.

user / channelprivacy filterinjection scannerlocal vLLM modeloutput validatoraudit

How the layers connect

The software’s model registry maps every fleet endpoint by role — generation, reasoning, vision, medical, embedding, rerank — so Agent OS routes each task to the right-sized model automatically. One fleet key, health-checked endpoints, native systemd deployment end to end. The same build ships as SaaS or fully self-hosted.

03Proof

Proof, not promises

ATC Quest

LIVE · TRL 8
atcquest.com

Our AI-native enterprise LMS, live with paying customers: five specialised agents, deployable on-premise, cloud, or hybrid, with custom fine-tuning for institutions.

Certification Exam Pipeline

PRODUCTION
Brazil

Our multi-agent RAG pipeline running in production internationally, generating and managing certification-exam questions for financial-market credentials — the same core, carrying a domain-critical workload end to end.

ATC CommandCenter / LeadFlow

ALPHA
internal

Our AI-native business operating system: a CRM core with an AI agent wielding 57 tools across 15 modules, all on local vLLM.

Act Now

RETROFIT PATH
ERP platform

A horizontal ERP platform built to compete with Zoho and Odoo — the first showcase of the retrofit path: the moment our AI layer switches on, existing data and workflows become conversational, grounded, and automated, with no rewrite and no migration.

04Why This Wins

Why this wins

RAG + MCP shrink the model requirement

Facts come from your data, actions from tools; the model orchestrates, it doesn’t memorize the internet.

Unified-memory hardware shrinks the box requirement

128 GB mini-PCs and one workstation GPU replace 8-GPU servers and InfiniBand.

The privacy filter unlocks regulated data

Pseudonymised in, validated out, hash-chained audit.

The economics invert

A year of electricity costs less than a month of comparable API spend, and the marginal token is effectively free.

Your alpha compounds at home

Every conversation, document, and workflow enriches your knowledge graph and retrieval index on your hardware. Nothing feeds a vendor’s next model.

A complete AI factory — models, agents, retrieval, privacy, and audit — that fits in an office rack, cools with the AC you already own, and by construction cannot leak your data.

05FAQ

Frequently asked questions

What is an AI Factory?

An AI Factory is a complete AI production capability that runs inside your own building — the models, the retrieval system that grounds them in your data, the agents that take action, the privacy controls, and the audit trail, all on hardware you own. Your documents, tickets, and workflows go in; grounded, auditable output comes out. Unlike a cloud AI subscription, the capability is an asset you own rather than access you rent, and your data never leaves your premises.

How is an AI Factory different from using cloud AI APIs?

A cloud API gives you metered access to someone else's infrastructure, with your data sent to their premises for processing. An AI Factory runs the models on your own hardware, on your own network. The practical differences are cost structure, data control, and ownership: API cost rises with usage forever, while an AI Factory's marginal cost approaches electricity; your data never leaves your building; and at the end of the term you retain the hardware and the accumulated capability rather than nothing.

When is cloud AI the better choice?

When your workload is low-volume, non-sensitive, exploratory, or genuinely requires frontier-model capability that open models don't yet match. A team of twelve running a few thousand queries a month on public data should not buy hardware. We classify workloads into four tiers and tell clients directly when they fall into Tier D — stay on cloud — including referring them elsewhere. On premise is the right answer for regulated data and for sustained high volume, not for everything.

What does an AI Factory cost compared to cloud APIs over three years?

Modelled at one million queries per month, a mid-size on-premise deployment totals roughly ₹49–78 lakh over three years against ₹58 lakh to ₹1.15 crore in cloud API spend, with the crossover typically falling at 8–14 months. Entry deployments serving 5–50 users start at ₹8–18 lakh. The structural difference matters more than the totals: the next million tokens costs ₹15–20 in electricity on owned hardware versus $2–15 on an API, and at the end of the term you still own the hardware.

Do I need a data centre to run an AI Factory?

No. The unit of deployment is a box on a desk, not a rack in a data centre. Our reference fleet is five nodes on a standard office LAN with about 500 GB of total model memory, drawing 300–500 watts in typical use. Each node is silent, plugs into a normal wall socket, and needs no liquid cooling, no raised floor, and no water — the office air conditioning you already run is the entire thermal plan.

Can smaller open models really replace frontier models for enterprise work?

For most enterprise workloads, yes — because the model isn't doing the work alone. Retrieval supplies the facts from your own documents and tools supply the actions, so the model orchestrates rather than memorises the internet. That shifts the requirement: a well-served 12B–35B open model handles work that frontier APIs are commonly sold for. Inference-time engineering — speculative decoding, KV-cache compression, expert offloading, prefix caching — is what keeps 240-watt nodes serving at interactive speed.

How long does deployment take?

A Discovery Sprint runs one day and produces a written recommendation. A full Architecture Blueprint takes four to six weeks. Build and deployment move in weeks rather than quarters because the platform already exists — the work is configuration and integration against your systems, not development. Each phase is a legitimate stopping point with its own deliverable, so you are never committing to the whole programme on a proposal.

Who operates the AI Factory after it is deployed?

Three options, decided at Blueprint stage. We can operate it as a managed service, handling monitoring, model updates, patching, and capacity planning. We can co-manage, with the Architect Retainer keeping us embedded for architecture and escalation while your team runs daily operations. Or we train your team during the Blueprint and hand over fully — the Blueprint deliverable documents exactly what that handover requires.

How does an AI Factory satisfy DPDP, HIPAA, and RBI requirements?

The platform runs in a compliance mode where no cloud egress path exists and failures are fail-closed, enforced in code rather than policy. A privacy filter detects and pseudonymises personal data — including Aadhaar, PAN, and mobile numbers — before any model sees it, using HMAC pseudonymisation so originals never reach a model while records stay linkable. Every request is written to a tamper-evident SHA-256 hash chain, giving auditors evidence rather than assertion. The filter is published as an MIT-licensed library your security team can read.

What happens when better models are released — does the hardware go obsolete?

The architecture deliberately reduces the model-size requirement, so new open models in the same 12B–35B class load onto the same hardware and deliver capability gains without a hardware cycle. Unified-memory nodes are also generous on memory relative to compute, which is the dimension that constrains model loading. We size for a three-year useful life and state that explicitly in the Blueprint rather than leaving it implied.

How does an AI Factory scale as usage grows?

By adding nodes. The same management plane runs one box or a full fleet, so growth is incremental rather than a re-platforming exercise. A single node comfortably serves around five concurrent requests — enough for a 20–30 person office, or a small hospital with three or four doctors working simultaneously. Most deployments start with one department, prove the workload, then expand.

What is the exit path if we stop working with you?

You own the hardware and your data, the models are open-weight, and the privacy filter is MIT-licensed. If the relationship ends, the AI Factory keeps running — what you lose is our operational support, not your capability. This is the structural difference from a subscription, where ending a contract ends your access to the capability entirely.