A pre-configured enterprise AI bundle can go live on your premises in days, not quarters — because the hardware is desk-size, the models are pre-integrated, and the sizing is done against your measured workload rather than a vendor's fixed specification. That deployment speed is real, it is scoped (custom agentic systems typically take 4–12 weeks, and we'll say exactly why), and this post walks through how the build actually happens.
Why we won't publish a fixed spec
Ask most AI infrastructure vendors for a quote and you receive a fixed specification: a rack, a server class, a bill of materials, a price. We deliberately don't publish one, and the reason is arithmetic, not mystery.
The sizing unit of the whole system is this: one DGX Spark node — 128 GB unified memory, 240 W — serves roughly five concurrent requests. In practice that's a 20–30 person office, or a small hospital with three or four doctors processing patients at once. Concurrency, not headcount, is what hardware must be sized to — and concurrency varies enormously between a law firm where everyone queries at 9 a.m. and a plant where shifts stagger around the clock.
A fixed spec is how a buyer ends up paying for capacity they never use. So instead of a spec, we run an assessment — users, concurrency, modalities, compliance regime — and size a base configuration around what your organisation will actually do. Growth afterwards means adding a node over the office LAN: no InfiniBand fabric, no re-architecture, ever. The same management plane runs one box or an enterprise fleet.
The hardware vocabulary (not a bill of materials)
| Hardware class | Memory · power | Typical role |
|---|---|---|
| DGX Spark (GB10 Grace-Blackwell) | 128 GB unified · 240 W | General generation, deep reasoning, retrieval + medical |
| RTX PRO 6000 workstation (Blackwell) | 96 GB + RTX 4070 | Multimodal input, streaming ASR |
| Roaming / edge-class | 8 GB | Demos, small pilots |
This table is vocabulary, not a quote — your own base configuration comes out of the assessment. What turns these boxes into a system is the Manthan cluster manager: bare metal to serving in one command, auto-discovery of new nodes, web-UI model lifecycle management, self-healing supervision, and per-model ROI analytics drawn from live telemetry. When a fleet spans more than one box, multi-node tensor-parallel inference runs over a dedicated QSFP fabric.
And it is a factory, not just an inference host: dataset generation, LoRA fine-tunes, and quantisation run on the same fleet, so the hardware that serves answers by day can improve models on the same premises.
The deployment timeline, honestly scoped
Because "how fast?" is the first question every buyer asks, here is the answer with its boundaries attached:
- Pre-configured bundles: days, not quarters to live. These are the standardised configurations — hardware classes above, pre-integrated models, Manthan and the privacy rail installed as shipped. The assessment determines which bundle; installation, network integration, and initial document ingestion follow a rehearsed runbook.
- Custom agentic systems: typically 4–12 weeks, depending on integration complexity. When the engagement includes bespoke agent workflows, deep CRM/ERP integration, or sector-specific compliance packs, the calendar is dominated by integration and validation work, not hardware.
Both numbers are true; they describe different products. A vendor quoting one number for both is describing neither.
The software layer: what makes hardware a factory
Manthan — retrieval and memory. Thirty-plus format parsers with layout-aware chunking, so the system ingests what your organisation actually has — scanned contracts, engineering drawings, spreadsheets — rather than assuming clean text. Hybrid keyword + semantic retrieval is re-scored by a multimodal reranker so every answer carries citations, backed by five-layer memory ending in an immutable audit trail.
Agent OS — the workforce. Teams of AI specialists run in ReAct loops under iteration, token, and wall-clock budgets — bounded by design, because unbounded loops waste a fixed-cost fleet. The platform ships 49 tool modules and 8 delivery channels, plus MCP support, CRM and ERP connectors, a DAG workflow engine with cron and event triggers, and signed agent-to-agent delegation, so an auditor can verify not just what an agent did but on whose cryptographically-verifiable authority.
The privacy rail — on every model call. PII/PHI detection extended for Aadhaar and PAN; pseudonymisation so originals never reach a model; DPDP, GDPR, HIPAA, and CCPA validators; a tamper-evident audit hash chain. Underneath it all: deny-by-default egress, sandboxed execution, and row-level tenancy — compliance in architecture, not policy.
The platform carries 2,664 automated tests across 109 modules, and the privacy filter is published open-source — verifiable, not asserted.
What a deployment week actually looks like
For a pre-configured bundle, the days, not quarters break down roughly as: assessment already done (it precedes the order); day one, nodes on desks, LAN integration, Manthan cluster manager brings bare metal to serving; days two and three, document corpus ingestion through the format parsers, retrieval quality checks, privacy-rail validation against your data classes (Aadhaar/PAN detection tuned to your documents); days four and five, staff onboarding, delivery-channel setup (chat, and whichever of the 8 channels your teams use), and handover with the audit trail live from the first query.
The reason this is days rather than months is architectural: nothing waits on a data-center buildout, a cloud migration, or a model training run. The models arrive ready; the knowledge arrives by ingestion; the compliance arrives as configuration of a rail that's already in the stack.
Proof it runs — at the right confidence tier
Deployment claims deserve receipts, so here are ours, labelled at their actual confidence level. In production with paying customers: ATC Quest LMS, plus an international production client in Brazil. Independently validated: 1st place in the NHA PM-JAY AI Challenge 2026. And methodologically: every performance number BiltIQ publishes — throughput, concurrency, power, context length — is measured from our own deployed fleet, not vendor datasheets.
Credentials the procurement checklist will ask about: ISO/IEC 27001:2022 certified, ISO 9001:2015 certified, DPIIT Recognised (DIPP239966), NVIDIA Inception Partner, DPDP Act 2023 compliant.
Two ways to buy it — the same stack either way
01 · The hardware layer, sized to you. We assess the workload, then specify, install, and operate the right fleet on your own infrastructure under our cluster software. Pre-configured bundles go live in days, not quarters; growth means adding a node, never a re-architecture.
02 · The software layer, on what you already run. Retrieval, connectors, and agent tools wrap the CRM, ERP, ticketing, and records systems already in place and turn them into AI-native software — no rewrite, no migration. It runs on rented GPUs for buyers who won't sign a hardware purchase order, at ₹1.3–3.6 lakh per year, with an unchanged path to on-premises hardware later.
A data center that isn't one — sized to your workload, installed in your building.
Get a workload assessment and a scoped deployment plan: [email protected] · +91 89868 60088 · www.biltiq.ai
Frequently asked questions
How fast can an enterprise AI system be deployed on-premises?
Pre-configured BiltIQ bundles go live in days, not quarters, covering hardware installation, document ingestion, privacy-rail validation, and staff onboarding. Custom agentic systems — bespoke workflows and deep CRM/ERP integration — typically take 4–12 weeks depending on integration complexity; the two timelines describe different products.
How is the hardware actually sized?
Sizing is driven by concurrency, not headcount: one DGX Spark node serves roughly five concurrent requests, suiting a 20–30 person office or a small hospital with three to four doctors working simultaneously. An assessment of users, concurrency, modalities, and compliance regime produces a base configuration, and growth afterwards means adding one node over the office LAN.
Why doesn't BiltIQ publish a fixed specification or bill of materials?
Because a fixed spec is how a buyer ends up paying for capacity they never use — concurrency patterns differ radically between organisations of identical headcount. The published hardware classes are vocabulary; the actual base configuration is quoted after a workload assessment.
What software runs on the fleet?
Three layers: Manthan (30+ format parsers, hybrid cited retrieval, five-layer memory to an immutable audit trail), Agent OS (ReAct agent teams under iteration/token/wall-clock budgets, 49 tool modules, 8 delivery channels, DAG workflows, signed delegation), and a privacy rail on every model call (Aadhaar/PAN-extended PII/PHI detection, pseudonymisation, DPDP/GDPR/HIPAA/CCPA validators, hash-chained audit). The platform carries 2,664 automated tests across 109 modules.
Can the same system run without buying hardware?
Yes — the identical software stack runs on rented GPUs at ₹1.3–3.6 lakh per year at any tier, wrapping the CRM, ERP, and records systems already in place with no rewrite and no migration. Moving to on-premises hardware later requires no re-architecture.
What proof exists that the platform runs in production?
ATC Quest LMS operates in production with paying customers, alongside an international production client in Brazil, and the stack took 1st place in the NHA PM-JAY AI Challenge 2026. Every performance figure BiltIQ publishes is measured from its own deployed fleet, and the company holds ISO/IEC 27001:2022, ISO 9001:2015, DPIIT recognition, NVIDIA Inception partnership, and DPDP Act 2023 compliance.
