Nexus On-Prem AI Engine (document-intelligence backend for "DocuMind")
a state government (India) — GovTech / land-records & revenue administration
An air-gapped, on-prem GenAI appliance that turns noisy OCR'd Hindi land records into grounded, citation-backed summaries on commodity CPU-only hardware — with zero data leaving the government firewall.
- Engagement
- R&D + product build + platform + solution advisory
- Delivery model
- On-prem · air-gapped (data sovereignty)
- Team
- Small core team — solution architecture, GenAI/LLM engineering, backend/API, DevOps & release engineering, UX
- Industry
- GovTech / e-governance; public-records digitisation
PROJECT PROFILE — Nexus On-Prem AI Engine
1. Snapshot
A government revenue department needed to summarise scanned, OCR'd Hindi land records — sensitive citizen data that, by law and policy, cannot leave the department's premises. Agentryx designed, built and deployed Nexus, an air-gapped, on-premise Generative-AI appliance that reads noisy OCR text and returns grounded, citation-backed summaries in Hindi, running entirely inside the government firewall on commodity, CPU-only hardware (12 cores / 24 GB RAM, no GPU). Nexus serves an OpenAI-compatible API to the client's existing "DocuMind" document application, wraps a quantised Mistral 7B model in a hardened control plane, and ships as a single 401 MB offline installer with backup and rollback. The hard, foundational problem — useful GenAI with zero data egress on minimal hardware — is delivered and proven in production; a phased roadmap then scales it across hardware, models and applications.
2. The Challenge
The department's records are decades of hand-written and printed land registers (Roznamcha Rapat, Patwar-circle entries) digitised through OCR. The material is hard in three ways at once, under constraints that rule out every mainstream cloud option:
- Data sovereignty / air-gap. Citizen land and revenue records are sensitive government data; the deployment must run on-prem, air-gapped, with no internet egress and no third-party AI API. This removes OpenAI/Azure/Bedrock and any hosted model from consideration.
- Minimal hardware. The available server is a 12-CPU / 24 GB commodity box with no GPU. The solution had to deliver useful language-model performance within that envelope — a regime where most LLM stacks are simply assumed off the table.
- Noisy OCR input. Scans produce misread characters (
1Ofor10), impossible dates (40 Jan), and unreadable spans — fed to the model as-is. Output must be Hindi/Devanagari, faithful to the source. - Zero tolerance for hallucination. In a legal/records context, an invented fact is worse than no answer. The system must summarise only what the source supports, and say so when it can't.
- Drop-in integration. An existing application ("DocuMind") already expected an OpenAI-style chat endpoint; the AI had to integrate without changing the consuming app.
3. What We Built (the solution)
Agentryx delivered the complete on-prem stack, not just a model. The shipped foundation comprises:
- A four-service AI appliance —
inference(the model engine),gateway(control plane & API),ui(admin + chat), andnginx(reverse proxy) — orchestrated as a single Docker Compose stack on one host. - A CPU-tuned inference engine — a bundled llama.cpp build serving Mistral 7B in quantised GGUF format, with runtime CPU-feature dispatch so one binary runs optimally across Intel generations (SSE4.2 → Sapphire Rapids).
- The Nexus Gateway — a ~1,900-line FastAPI control plane that fronts the engine with an OpenAI-compatible
/v1API, injects the grounding/OCR system prompt, formats prompts per model template (Mistral / Llama-3 / Gemma / ChatML), scales summary length to input size, manages models (load/switch), and exposes live metrics, insights and resource governance. - A grounded-summarisation pipeline — passage-labelled input (
[P1]…), citation-tagged output ((P1)), Hindi rendering, and an explicit "NO GROUNDED SUMMARY" hallucination guardrail that refuses rather than invents. - An OCR-correction prompt framework — date-sanity checks, misread-character repair, and
[unclear]marking, applied before summarisation. - A Next.js admin & chat experience — with operational panels (Log Explorer, Traffic Inspector, Resource Usage) and conversation history.
- A production-grade offline installer — a 401 MB, self-contained, image-based package with a 7-phase install, timestamped backup and one-command rollback, designed for an air-gapped server.
The platform matured across releases (v0.7.1 → v0.8.0): inference moved from a host service into a container, a reverse-proxy tier was added, deployment was hardened for air-gap, and backup/rollback was introduced — evidence of disciplined iteration toward production.
4. Architecture & How It Works
Request flow. An external request from DocuMind hits nginx on port 80, which routes /v1 API calls to the Gateway and serves the UI for browser users. The Gateway injects the accuracy/grounding system prompt, renders the conversation into the active model's exact prompt template, and calls the llama.cpp engine's completion endpoint. It then repackages the engine's response into OpenAI chat shape, records token usage and an activity log, and returns it. Long CPU generations are accommodated with aligned 600-second timeouts end-to-end so requests complete rather than fail.
Model management. Models live as GGUF files under a shared storage volume; an active symlink selects the live model, and switching repoints it and recreates the engine container — a deliberately simple, robust mechanism for a single-node appliance.
Offline delivery. The installer loads pre-built Docker images from disk (docker load), lays down configuration and scripts under /opt/nexus, links the model, and boots the stack — no registry, no internet.
Key design decisions, with rationale.
- On-prem & air-gapped — mandated by data sovereignty; the entire design assumes zero egress.
- CPU-only quantised GGUF on llama.cpp — to deliver language-model capability within a 12-core / 24 GB budget at low cost, with portable runtime CPU dispatch instead of GPU dependence.
- An OpenAI-compatible gateway in front of the engine — so the existing DocuMind app integrates unchanged, while the gateway adds grounding, OCR correction, model management and observability the raw engine lacks.
- Prompt-level grounding and refusal — accuracy and non-hallucination enforced in the request path, appropriate to a legal/government context.
- Single-host Docker Compose + offline installer — the simplest reproducible unit that fits an air-gapped server, with backup/rollback for safe field updates.
5. Technology Stack
- AI/ML models: Mistral 7B (quantised GGUF) served via llama.cpp; grounded RAG-style summarisation; prompt-engineered OCR correction; multi-template prompt formatting (Mistral/Llama-3/Gemma/ChatML); bundled llama.cpp toolchain for quantisation/benchmarking/fine-tune (
llama-quantize,llama-imatrix,llama-bench,llama-fit-params). - Backend: FastAPI, Python, OpenAI-compatible REST API, background jobs, system-prompt injection & response shaping.
- Frontend: Next.js (admin + chat), operational dashboards, conversation history.
- Infra & deployment: Docker & Docker Compose, nginx reverse proxy, self-contained offline/air-gapped installer (bash,
docker load, image bundles), backup & rollback, systemd (legacy inference path). - Security & compliance: on-prem data sovereignty, air-gapped operation, zero external data egress, single-host isolation.
- Hardware: commodity x86 server — 12 CPU / 24 GB RAM, CPU-only (no GPU); runtime CPU-feature dispatch (SSE4.2 → Alder Lake / Ice Lake / Sapphire Rapids).
6. Capabilities & Expertise Demonstrated (Capability → Evidence)
- On-prem / air-gapped GenAI delivery → a complete, internet-free appliance with a 401 MB offline installer and zero data egress, running inside a government firewall.
- LLM serving on constrained, CPU-only hardware → quantised Mistral 7B on llama.cpp delivering useful summarisation on a 12-core / 24 GB box with no GPU.
- Applied prompt engineering & grounding → citation-backed summaries with a "NO GROUNDED SUMMARY" refusal guardrail and OCR date/character correction.
- Multilingual document AI → Hindi/Devanagari summarisation of OCR'd land records.
- API & integration engineering → an OpenAI-compatible gateway that let the existing DocuMind app adopt the AI without code changes.
- Platform / control-plane engineering → a FastAPI control plane with model management, live metrics & insights, and runtime resource governance.
- Release & deployment engineering → a 7-phase, image-based offline installer with timestamped backup and one-command rollback for air-gapped field updates.
- Solution architecture & advisory → a phased path from a proven minimal-hardware foundation to a scaled multi-model platform, with explicit hardware/model/application axes.
7. Hard Problems Solved / Innovations
- Useful GenAI with zero egress on a no-GPU, 12-core box. Quantised GGUF + a CPU-tuned llama.cpp build + an OpenAI-compatible control plane together make a normally GPU-and-cloud-shaped workload run, usefully, inside an air-gapped government server. This foundation is the differentiator — and it is built and proven, not theorised.
- Grounding and refusal for a legal context. Rather than trust a 7B model to "be accurate", the system enforces grounding in the request path — passage citations plus an explicit refusal token — so it declines on insufficient evidence instead of fabricating, which is the correct failure mode for land records.
- Prompt-level OCR repair for Devanagari records. Date-sanity, misread-character correction and
[unclear]handling are applied as a first-class pre-summarisation step, tuned to the failure modes of scanned Hindi registers. - Air-gapped delivery as a product. Packaging the whole stack as a reproducible offline installer with backup/rollback turns a bespoke deployment into a repeatable, field-serviceable appliance — a reusable capability beyond this one client.
8. Outcomes & Impact
- Measured. Nexus is in production at the client on a 12-CPU / 24 GB CPU-only server. On Mistral 7B it sustains ~5 tokens/second, summarising a land-record passage in roughly 45–50 seconds end-to-end; the grounding guardrail was observed correctly refusing low-information inputs. Zero citizen data leaves the premises. The foundation deployed across iterative releases (v0.7.1 → v0.8.0) with backup/rollback.
- Strategic. The engagement de-risks scale-up: the difficult, novel part (on-prem, air-gapped, minimal-hardware GenAI) is proven, making subsequent scaling a largely standard exercise against a known baseline.
- Projected (roadmap targets, to validate). A Gemma-class "Flash" fast-path targets sub-30-second summaries (≈5–10× speed-up); quantisation (IQ3/IQ2) is modelled to run a summary model and a chat model co-resident within the 24 GB envelope; a GPU engine option is projected to unlock larger models and real concurrency. (These figures are modelled/targeted, not yet measured.)
9. Roadmap & Evolution
Phase I–III delivered the proven foundation: the offline appliance, single-model CPU serving, the grounded-summarisation + OCR pipeline, and the control plane. With the hard part established, the forward roadmap is a deliberately phased, standard scale-up — stabilise & secure → scale foundation → then scale outward — across three axes:
- Hardware. A GPU-accelerated engine option (layer offload) for a step-change in throughput and larger models, alongside CPU scale-up with NUMA tuning and IQ-quantisation to run multiple models within RAM.
- AI Models (Phase IV — the scaled multi-model architecture). Move from one active model to a multi-model serving layer with per-request routing: a fast Gemma "Flash" path for standard summaries, a stronger model for hard documents, and a small model for chat — served concurrently, with a static KV-cache for the fixed grounding prompt and a warm model pool to remove switch latency. A fine-tuning pipeline (CSV/JSONL → adapter training, run headless in low-traffic windows) lets the platform specialise on the department's own corpus.
- Applications. A first-class, owned UI and a multi-tenant, authenticated, quota'd API so new government applications onboard with a key rather than a code change — turning Nexus from one app's backend into an on-prem AI platform. High-availability orchestration and centralised observability complete the picture.
Phase IV and beyond are planned/projected; this profile presents them as roadmap, distinct from the delivered-and-measured foundation in §3 and §8.
Related projects
Connected through shared technologies and capabilities.
DocuMind — Government Document AI Pipeline
In productiona state government department of personnel (India); integrates with the state e-archive
An air-gapped government document AI pipeline that turns inbound departmental PDFs into searchable, summarised, archive-ready records — with self-service ops controls that replaced terminal SQL with audit-logged UI actions.
AI-Powered Fabric Sourcing Platform
Proof of conceptone of India's largest vertically-integrated textile manufacturers (supplies global apparel brands)
An on-prem AI sourcing platform that turns a fabric image into the right match across a 70,000-SKU catalogue in seconds, powered by a proprietary, explainable 4-pillar matching engine.
Drishti AI — Grievance Intelligence Platform
Proof of conceptChief Minister's Office, Haryana (CMO Haryana) — State Government of India
An AI-powered grievance intelligence platform that cross-references citizen complaints, government Action Taken Reports and recorded call transcripts to score genuine resolution quality, surface false closures and rank cases for supervisory action — giving the CMO a real-time audit layer over the state grievance system.
Facing a similar problem?
Talk to Agentryx