AGENTRYX
← All work
In production2026 · in delivery (pilot production stable at v0.9.15.x; v1.0 hardening)

DocuMind — Government Document AI Pipeline

a state government department of personnel (India); integrates with the state e-archive

An air-gapped government document AI pipeline that turns inbound departmental PDFs into searchable, summarised, archive-ready records — with self-service ops controls that replaced terminal SQL with audit-logged UI actions.

Engagement
product build + platform + AI/ML integration + release & ops engineering
Delivery model
Air-gapped on-prem (PROD on the state government's internal network); 4-environment workflow (DEV · STG · test-ring · PROD) with USB-only delivery
Team
~2–3 across solution architecture, backend / pipeline, DevOps & release, product/UX
Industry
GovTech / e-governance; departmental document processing & archive integration

PROJECT PROFILE — DocuMind: Government Document AI Pipeline

1. Snapshot

A state government department of personnel (India) receives a continuous flow of departmental PDFs that must be OCR'd, summarised in English and Hindi, and pushed to the state e-archive inside a strictly air-gapped environment. Agentryx designed, built and operates DocuMind: an on-prem document AI pipeline pairing Tesseract + PaddleOCR with a self-hosted Mistral 7B LLM, fronted by an explainable Resolution Queue that lets a single operator triage unprocessable scans via templated overrides. By mid-2026 the platform had processed 3,491 documents at 100% pipeline completion with 7,479 successful e-archive pushes and every operator action audit-logged. Releases follow a strict DEV → STG → PROD discipline with USB-only delivery and idempotent installers.

2. The Challenge

Hundreds of PDFs each month — many scanned from aged paper, many in Hindi / Devanagari, some encoded with legacy Krutidev fonts poppler cannot rasterise — must reach the state e-archive with a clean structured summary even when the source is unsummarisable. Hard constraints: air-gapped PROD (no outbound internet, no scp; only USB transfer of sha256-verified tarballs); data sovereignty (no content leaves the state government network); a single non-engineer operator runs the platform — any cleanup that requires psql is a scaling ceiling; Mistral 7B's 4,096-token context cannot hold a long Hindi circular in one pass. Two systemic defects observed pre-engagement: a phantom-phenom where pushes succeeded but local row updates silently dropped, and a synthetic duplicate-link defect polluting the operator dashboard.

3. What We Built

  • Eight-stage pipeline — PDF Ingest → OCR (Tesseract + PaddleOCR fallback) → Text Extract → Chunk → Embeddings (pgvector) → Semantic Summary → LLM Summary (English / Hindi) → e-archive webhook push.
  • Ops Intelligence dashboard — completion funnel, drift-vs-baseline alert, 3rd-party delivery status, push timeline, document investigator, daily report.
  • Resolution Queue UI (v0.9.15) — left-rail of stuck docs, embedded PDF, OCR / Current LLM / Override tabs, 5 confidence-scored templates auto-suggested by an in-process diagnoser, batch auto-resolve with operator-configurable threshold.
  • Self-service Ops Acknowledgement controls (v0.9.15.1) — admin-gated UI buttons on every deficit panel (Drift alert, each funnel row) writing to ops_stat_adjustments via an audit-triggered DB row, replacing what used to be ad-hoc terminal SQL.
  • Operator override system — 5 templates (OCR_LOW_YIELD, NUMERICAL_TABULAR, RENDER_FAILURE, HALLUC_GATE_TRIGGERED, LANG_DEFERRED) seeded in override_templates; every override captured in operator_override_audit.
  • Idempotent air-gapped installer — topology auto-detect (stackctl vs systemd), preship ack gate, dry-run mode, sha256 pack integrity, per-ship rollback recipe.

4. Architecture & How It Works

Three Linux services behind a single nginx host. documind-api (FastAPI, :8000) owns admin endpoints and document mutation; documind-ask-api (FastAPI, :9010) owns the Ops Intelligence + Ask surfaces; a unified background worker drives OCR + chunk + embed + LLM + push. State lives in PostgreSQL with pgvector, with audit triggers on every operationally-mutable table so every state change is durably traceable across releases. The React + Vite + TypeScript frontend is bundled at build time with VERSION baked in. The LLM stage calls the state government's internal Mistral 7B endpoint — data never leaves the firewall. nginx regex-routes specific admin paths to documind-api and /ask/api/... to ask-api; this asymmetric routing was added in v0.9.15.1 after a 500 traced to the wrong upstream.

Releases follow a strict DEV → STG → PROD discipline on a single source tree, with per-environment differences expressed through env vars rather than forks. The build VM hosts DEV plus a PROD-shape test ring; STG mirrors PROD with e-archive push disabled; PROD is air-gapped. A PreToolUse shell hook blocks tar, install.sh, and rsync.*opt/dms until a daily acknowledgement file exists — making it impossible to ship without re-reading the checklist that day.

Key decisions, each evidence-led: on-prem LLM for data sovereignty and predictable cost; audit triggers in the DB so the audit survives rollouts and rollbacks; in-app self-service buttons over psql to make the operator scalable; idempotent migrations so partial ships re-run safely; VERSION as single source of truth for backend /api/health and frontend bundle, eliminating display-vs-running drift.

5. Technology Stack

  • Frontend — React, Vite, TypeScript, TailwindCSS, wouter routing, hash-based bundle caching.
  • Backend — FastAPI, Python, asyncio, psycopg2, in-process background workers.
  • Data — PostgreSQL + pgvector; schema-versioned migrations (0985–0990+); audit triggers on ops_stat_adjustments, operator_override_audit, pipeline_settings, ingest_audit_log.
  • AI/ML models — Mistral 7B (self-hosted via llama.cpp on the state government's internal endpoint), embedding model server, semantic summarisation.
  • Vision / OCR — Tesseract (English), PaddleOCR (Hindi / Devanagari) with cross-fallback orchestration.
  • Integrations — state e-archive webhook, the state government's internal LLM endpoint, nginx topology-specific admin routing.
  • Infra & deployment — nginx, systemd per-service (DEV), stackctl.sh (STG/PROD), air-gapped USB delivery, sha256-verified packs, idempotent installers with topology auto-detect and dry-run.
  • Security & compliance — X-Admin-Secret RBAC, data sovereignty (air-gapped PROD), audit log on every state-changing endpoint, daily preship-ack gate.
  • QA & testing — 7-step ship checklist, DEV-first-test-then-deliver, STG freeze workflow re-anchored after every PROD ship, per-ship rollback recipe.

6. Capabilities & Expertise Demonstrated

  • GovTech / air-gapped delivery → 4 PROD releases shipped in one working day (v0.9.14.1.0 → v0.9.15.1.1) on a strictly air-gapped network using USB transfer + sha256 packs, zero downtime.
  • Multi-modal AI pipeline engineering → OCR + chunk + embed + semantic + LLM pipeline processing 3,491 documents at 100% completion.
  • On-prem LLM operations → integration with the state government's Mistral 7B endpoint including the 4,096-token-context constraint and the Hindi placeholder path.
  • Self-service ops engineering → 6 routine cleanup operations replaced by audit-logged UI buttons; founder closed the full PROD dashboard RED → GREEN without typing a single line of SQL.
  • Audit-by-design schema → every operationally-mutable table fronted by a trigger writing to a parallel _audit table with JSONB old_row / new_row diff and actor captured via application_name.
  • Release engineering → idempotent install pipeline with topology auto-detect, daily preship hook, dry-run, named rollback recipe per ship.
  • Hindi NLP / Devanagari handling → LANG_DEFERRED template for context-overflow docs plus a Devanagari-aware OCR fallback (PaddleOCR).
  • Forensic debugging → diagnosed and shipped fixes for two systemic defects: phantom-phenom (silent UPDATE drop) caught via audit-log reconciliation, and a synthetic duplicate-link defect closed by a SQL filter shipped in v0.9.14.2.0 / v0.9.15.0.0.

7. Hard Problems Solved / Innovations

  • Hindi LLM context overflow. Mistral 7B's 4,096-token context cannot hold a full Hindi government circular. Rather than block ingestion, the doc reaches the e-archive with a templated placeholder (LANG_DEFERRED), the deferral is logged in operator_override_audit, and the doc surfaces in the Resolution Queue — keeping the archive complete while marking deferred items for human review.
  • Self-service ops over terminal SQL. Every new operational deficit historically created a new psql snippet for the operator. v0.9.15.1 inverted this with one React component (OpsAckButton) + existing admin endpoints + an audit trigger — giving every deficit panel a one-click reset that writes a full audit row. The pattern is now reused for every new alert surface added.
  • Phantom-phenom (silent UPDATE drops). Pushes left the system, the e-archive accepted them, but the local documents row update silently dropped — leaving the e-archive ahead of DocuMind. Solved by audit triggers on every state table (the discrepancy is now mechanically detectable) plus an explicit reconciliation pass driven by ingest_audit_log as the source of truth for incoming POSTs.
  • Idempotent air-gapped delivery. install.sh detects topology (stackctl vs systemd), timestamped-backs-up everything it will touch, skips already-applied patches (nginx, migration), runs a real verification step (curl /api/health and grep-assert the new version), prints the rollback recipe at end. Re-running a ship is a no-op; partial ships safely resume.

8. Outcomes & Impact

Measured (PROD dashboard, 16 June 2026):

  • 3,491 documents through the pipeline end-to-end.
  • 100% pipeline completion across all eight stages.
  • 7,479 successful e-archive webhook pushes (zero failed).
  • 99.94% docs with proper LLM summary; 99.99% page-weighted coverage.
  • 3 PROD releases shipped in one working day, zero downtime, zero rollback.
  • ~80 → 0 synthetic duplicate-link rows removed from the "Documents NOT pushed" panel via the v0.9.15.0.0 filter.
  • 6 stale ops adjustments self-served via UI in a single operator session — no psql opened.

Operational impact (qualitative): the operator scaling ceiling moved from "minutes of psql per incident" to "one click + a reason string" for the routine cases; founder time recovered for v1.0 hardening rather than incident response.

9. Roadmap & Evolution

  • v0.9.x (current) — production-stable foundation: Resolution Queue, Ops Acknowledgement controls, audit-by-design.
  • v0.9.15.2 (next ship) — Resolution Queue filter refinement, override-status cascade, audit history page at /ops/adjustments-history.
  • v1.0 (hardening) — long-context Hindi LLM upgrade, forensic alerting on phantom-phenom recurrence, per-department configurable resolution thresholds.
  • Phase 2 — multi-department rollout beyond Personnel; additional Indic languages.
  • Phase 3 — operator analytics feeding back into the diagnoser confidence model.
GovTeche-governanceair-gappedon-prem LLMMistral 7BTesseractPaddleOCRpgvectorFastAPIReactPostgreSQLaudit logoperator-overrideHindi NLPrelease engineering

Related projects

Connected through shared technologies and capabilities.

Facing a similar problem?

Talk to Agentryx