DocuMind — Government Document AI Pipeline
a state government department of personnel (India); integrates with the state e-archive
An air-gapped government document AI pipeline that turns inbound departmental PDFs into searchable, summarised, archive-ready records — with self-service ops controls that replaced terminal SQL with audit-logged UI actions.
- Engagement
- product build + platform + AI/ML integration + release & ops engineering
- Delivery model
- Air-gapped on-prem (PROD on the state government's internal network); 4-environment workflow (DEV · STG · test-ring · PROD) with USB-only delivery
- Team
- ~2–3 across solution architecture, backend / pipeline, DevOps & release, product/UX
- Industry
- GovTech / e-governance; departmental document processing & archive integration
PROJECT PROFILE — DocuMind: Government Document AI Pipeline
1. Snapshot
A state government department of personnel (India) receives a continuous flow of departmental PDFs that must be OCR'd, summarised in English and Hindi, and pushed to the state e-archive inside a strictly air-gapped environment. Agentryx designed, built and operates DocuMind: an on-prem document AI pipeline pairing Tesseract + PaddleOCR with a self-hosted Mistral 7B LLM, fronted by an explainable Resolution Queue that lets a single operator triage unprocessable scans via templated overrides. By mid-2026 the platform had processed 3,491 documents at 100% pipeline completion with 7,479 successful e-archive pushes and every operator action audit-logged. Releases follow a strict DEV → STG → PROD discipline with USB-only delivery and idempotent installers.
2. The Challenge
Hundreds of PDFs each month — many scanned from aged paper, many in Hindi / Devanagari, some encoded with legacy Krutidev fonts poppler cannot rasterise — must reach the state e-archive with a clean structured summary even when the source is unsummarisable. Hard constraints: air-gapped PROD (no outbound internet, no scp; only USB transfer of sha256-verified tarballs); data sovereignty (no content leaves the state government network); a single non-engineer operator runs the platform — any cleanup that requires psql is a scaling ceiling; Mistral 7B's 4,096-token context cannot hold a long Hindi circular in one pass. Two systemic defects observed pre-engagement: a phantom-phenom where pushes succeeded but local row updates silently dropped, and a synthetic duplicate-link defect polluting the operator dashboard.
3. What We Built
- Eight-stage pipeline — PDF Ingest → OCR (Tesseract + PaddleOCR fallback) → Text Extract → Chunk → Embeddings (pgvector) → Semantic Summary → LLM Summary (English / Hindi) → e-archive webhook push.
- Ops Intelligence dashboard — completion funnel, drift-vs-baseline alert, 3rd-party delivery status, push timeline, document investigator, daily report.
- Resolution Queue UI (v0.9.15) — left-rail of stuck docs, embedded PDF, OCR / Current LLM / Override tabs, 5 confidence-scored templates auto-suggested by an in-process diagnoser, batch auto-resolve with operator-configurable threshold.
- Self-service Ops Acknowledgement controls (v0.9.15.1) — admin-gated UI buttons on every deficit panel (Drift alert, each funnel row) writing to
ops_stat_adjustmentsvia an audit-triggered DB row, replacing what used to be ad-hoc terminal SQL. - Operator override system — 5 templates (
OCR_LOW_YIELD,NUMERICAL_TABULAR,RENDER_FAILURE,HALLUC_GATE_TRIGGERED,LANG_DEFERRED) seeded inoverride_templates; every override captured inoperator_override_audit. - Idempotent air-gapped installer — topology auto-detect (stackctl vs systemd), preship ack gate, dry-run mode, sha256 pack integrity, per-ship rollback recipe.
4. Architecture & How It Works
Three Linux services behind a single nginx host. documind-api (FastAPI, :8000) owns admin endpoints and document mutation; documind-ask-api (FastAPI, :9010) owns the Ops Intelligence + Ask surfaces; a unified background worker drives OCR + chunk + embed + LLM + push. State lives in PostgreSQL with pgvector, with audit triggers on every operationally-mutable table so every state change is durably traceable across releases. The React + Vite + TypeScript frontend is bundled at build time with VERSION baked in. The LLM stage calls the state government's internal Mistral 7B endpoint — data never leaves the firewall. nginx regex-routes specific admin paths to documind-api and /ask/api/... to ask-api; this asymmetric routing was added in v0.9.15.1 after a 500 traced to the wrong upstream.
Releases follow a strict DEV → STG → PROD discipline on a single source tree, with per-environment differences expressed through env vars rather than forks. The build VM hosts DEV plus a PROD-shape test ring; STG mirrors PROD with e-archive push disabled; PROD is air-gapped. A PreToolUse shell hook blocks tar, install.sh, and rsync.*opt/dms until a daily acknowledgement file exists — making it impossible to ship without re-reading the checklist that day.
Key decisions, each evidence-led: on-prem LLM for data sovereignty and predictable cost; audit triggers in the DB so the audit survives rollouts and rollbacks; in-app self-service buttons over psql to make the operator scalable; idempotent migrations so partial ships re-run safely; VERSION as single source of truth for backend /api/health and frontend bundle, eliminating display-vs-running drift.
5. Technology Stack
- Frontend — React, Vite, TypeScript, TailwindCSS, wouter routing, hash-based bundle caching.
- Backend — FastAPI, Python, asyncio, psycopg2, in-process background workers.
- Data — PostgreSQL +
pgvector; schema-versioned migrations (0985–0990+); audit triggers onops_stat_adjustments,operator_override_audit,pipeline_settings,ingest_audit_log. - AI/ML models — Mistral 7B (self-hosted via llama.cpp on the state government's internal endpoint), embedding model server, semantic summarisation.
- Vision / OCR — Tesseract (English), PaddleOCR (Hindi / Devanagari) with cross-fallback orchestration.
- Integrations — state e-archive webhook, the state government's internal LLM endpoint, nginx topology-specific admin routing.
- Infra & deployment — nginx,
systemdper-service (DEV),stackctl.sh(STG/PROD), air-gapped USB delivery, sha256-verified packs, idempotent installers with topology auto-detect and dry-run. - Security & compliance — X-Admin-Secret RBAC, data sovereignty (air-gapped PROD), audit log on every state-changing endpoint, daily preship-ack gate.
- QA & testing — 7-step ship checklist, DEV-first-test-then-deliver, STG freeze workflow re-anchored after every PROD ship, per-ship rollback recipe.
6. Capabilities & Expertise Demonstrated
- GovTech / air-gapped delivery → 4 PROD releases shipped in one working day (v0.9.14.1.0 → v0.9.15.1.1) on a strictly air-gapped network using USB transfer + sha256 packs, zero downtime.
- Multi-modal AI pipeline engineering → OCR + chunk + embed + semantic + LLM pipeline processing 3,491 documents at 100% completion.
- On-prem LLM operations → integration with the state government's Mistral 7B endpoint including the 4,096-token-context constraint and the Hindi placeholder path.
- Self-service ops engineering → 6 routine cleanup operations replaced by audit-logged UI buttons; founder closed the full PROD dashboard RED → GREEN without typing a single line of SQL.
- Audit-by-design schema → every operationally-mutable table fronted by a trigger writing to a parallel
_audittable with JSONBold_row/new_rowdiff and actor captured viaapplication_name. - Release engineering → idempotent install pipeline with topology auto-detect, daily preship hook, dry-run, named rollback recipe per ship.
- Hindi NLP / Devanagari handling →
LANG_DEFERREDtemplate for context-overflow docs plus a Devanagari-aware OCR fallback (PaddleOCR). - Forensic debugging → diagnosed and shipped fixes for two systemic defects: phantom-phenom (silent UPDATE drop) caught via audit-log reconciliation, and a synthetic duplicate-link defect closed by a SQL filter shipped in v0.9.14.2.0 / v0.9.15.0.0.
7. Hard Problems Solved / Innovations
- Hindi LLM context overflow. Mistral 7B's 4,096-token context cannot hold a full Hindi government circular. Rather than block ingestion, the doc reaches the e-archive with a templated placeholder (
LANG_DEFERRED), the deferral is logged inoperator_override_audit, and the doc surfaces in the Resolution Queue — keeping the archive complete while marking deferred items for human review. - Self-service ops over terminal SQL. Every new operational deficit historically created a new
psqlsnippet for the operator. v0.9.15.1 inverted this with one React component (OpsAckButton) + existing admin endpoints + an audit trigger — giving every deficit panel a one-click reset that writes a full audit row. The pattern is now reused for every new alert surface added. - Phantom-phenom (silent UPDATE drops). Pushes left the system, the e-archive accepted them, but the local
documentsrow update silently dropped — leaving the e-archive ahead of DocuMind. Solved by audit triggers on every state table (the discrepancy is now mechanically detectable) plus an explicit reconciliation pass driven byingest_audit_logas the source of truth for incoming POSTs. - Idempotent air-gapped delivery.
install.shdetects topology (stackctl vs systemd), timestamped-backs-up everything it will touch, skips already-applied patches (nginx, migration), runs a real verification step (curl/api/healthand grep-assert the new version), prints the rollback recipe at end. Re-running a ship is a no-op; partial ships safely resume.
8. Outcomes & Impact
Measured (PROD dashboard, 16 June 2026):
- 3,491 documents through the pipeline end-to-end.
- 100% pipeline completion across all eight stages.
- 7,479 successful e-archive webhook pushes (zero
failed). - 99.94% docs with proper LLM summary; 99.99% page-weighted coverage.
- 3 PROD releases shipped in one working day, zero downtime, zero rollback.
- ~80 → 0 synthetic duplicate-link rows removed from the "Documents NOT pushed" panel via the v0.9.15.0.0 filter.
- 6 stale ops adjustments self-served via UI in a single operator session — no
psqlopened.
Operational impact (qualitative): the operator scaling ceiling moved from "minutes of psql per incident" to "one click + a reason string" for the routine cases; founder time recovered for v1.0 hardening rather than incident response.
9. Roadmap & Evolution
- v0.9.x (current) — production-stable foundation: Resolution Queue, Ops Acknowledgement controls, audit-by-design.
- v0.9.15.2 (next ship) — Resolution Queue filter refinement, override-status cascade, audit history page at
/ops/adjustments-history. - v1.0 (hardening) — long-context Hindi LLM upgrade, forensic alerting on phantom-phenom recurrence, per-department configurable resolution thresholds.
- Phase 2 — multi-department rollout beyond Personnel; additional Indic languages.
- Phase 3 — operator analytics feeding back into the diagnoser confidence model.
Related projects
Connected through shared technologies and capabilities.
Nexus On-Prem AI Engine (document-intelligence backend for "DocuMind")
In productiona state government (India) — GovTech / land-records & revenue administration
An air-gapped, on-prem GenAI appliance that turns noisy OCR'd Hindi land records into grounded, citation-backed summaries on commodity CPU-only hardware — with zero data leaving the government firewall.
HireStream — Overseas Placement Portal
Staging-verifieda state electronics & IT corporation (India), under the State's Department of Information Technology
A configuration-driven, multi-role overseas-placement platform that takes a candidate from the State from registration to verified placement abroad through KYB-licensed agencies and employers — wrapped in an embedded testing framework with a pre-merge confidence gate that calibrates release risk on every push.
State Tourism eServices Portal
In productiona state government department of tourism (India)
A configuration-driven e-governance platform that runs the full tourism-licensing lifecycle — application → multi-officer review → State payment → verifiable digital certificate — across 8 service verticals and all 12 districts of the State.
Facing a similar problem?
Talk to Agentryx