Application note · Specialised production

Application Note: Sovereign RAG and LLM Serving Behind the Firewall

Permission-aware retrieval, efficient inference, evaluation, monitoring and controlled hybrid routing for confidential enterprise knowledge.

This design serves a private assistant over internal policies and technical knowledge. It assumes source permissions must be enforced before retrieval and that model output is advisory, cited and observable.

Model and platform fit

Benchmark representative prompts at intended context length and concurrency. Use AxiRelay for four smaller replicas/services, AxiAnvil for one 96GB model tier and AxiTitan for multiple tenants, large distributed models, fine-tuning and maintenance headroom. Quantisation changes quality as well as memory and must be evaluated.

Ingestion and permissions

Extract approved documents with owner, revision, effective date, classification and group ACL. Chunk deterministically and index embeddings plus metadata. At query time, filter candidates by the caller before text enters the prompt. Delete/rebuild index entries when documents or access change.

Serving plane

Run vLLM or TensorRT-LLM behind an authenticated gateway; use Triton where its repository, batching and metrics fit the service portfolio. Apply per-user quotas, maximum context, timeouts and cancellation. Keep a versioned model endpoint so releases can be canaried and rolled back.

Evaluation and guardrails

Test retrieval recall, citation correctness, grounded answer quality, refusal, sensitive-data leakage and prompt injection. NeMo Guardrails can help structure conversational controls, but tool permissions and data access must be enforced in application code. Log policy outcomes and model versions while minimising retained confidential text.

Operations

Monitor time-to-first-token, generation rate, queue time, VRAM, cache, error and user feedback. Define overload behaviour and an offline maintenance page. A hybrid gateway may route approved non-sensitive tasks externally only after classification and explicit policy; protected prompts fail closed locally.

Primary technical references

References are provided for software architecture and implementation planning. Validate the versions, licences, support matrix and regulated-use requirements applicable to the final deployment.

From technical concept to production system

Apply this technology through an Axiotech engineering work package

Axiotech can connect the compute, AI or analytics platform to the machine controls, data contracts, validation evidence and lifecycle-support model required for industrial use.

Relevant Hyperion platforms

Hyperion X1440

AxiRelay

A dense inference, RAG, CI and multi-service node where several independent GPU workloads must run concurrently.

One 4U node; 96 AMD EPYC cores; 768GB ECC DDR5; four NVIDIA RTX PRO 4000 24GB GPUs providing 96GB aggregate VRAM.

Hyperion X1160

AxiAnvil

A large-memory single-GPU platform for private LLM inference, retrieval-augmented generation, medical imaging and large scientific models.

One 4U node; 96 AMD EPYC cores; 768GB ECC DDR5; one NVIDIA RTX PRO 6000 96GB GPU; enterprise NVMe.

Hyperion X4460

AxiTitan

The flagship sovereign AI and research-computing platform for organisation-scale shared services and the largest Hyperion workloads.

Four 4U nodes; 512 AMD EPYC cores; 4.6TB ECC DDR5; sixteen NVIDIA RTX PRO 6000 96GB GPUs; more than 1.5TB aggregate VRAM; 100GbE compute and dedicated storage fabrics.

Configuration and quotation

Validate this workload on Hyperion

Final architecture and price depend on representative code and data, concurrency, storage, networking, site infrastructure, component availability and export compliance.

Request Formal Quotation