This design serves a private assistant over internal policies and technical knowledge. It assumes source permissions must be enforced before retrieval and that model output is advisory, cited and observable.
Model and platform fit
Benchmark representative prompts at intended context length and concurrency. Use AxiRelay for four smaller replicas/services, AxiAnvil for one 96GB model tier and AxiTitan for multiple tenants, large distributed models, fine-tuning and maintenance headroom. Quantisation changes quality as well as memory and must be evaluated.
Ingestion and permissions
Extract approved documents with owner, revision, effective date, classification and group ACL. Chunk deterministically and index embeddings plus metadata. At query time, filter candidates by the caller before text enters the prompt. Delete/rebuild index entries when documents or access change.
Serving plane
Run vLLM or TensorRT-LLM behind an authenticated gateway; use Triton where its repository, batching and metrics fit the service portfolio. Apply per-user quotas, maximum context, timeouts and cancellation. Keep a versioned model endpoint so releases can be canaried and rolled back.
Evaluation and guardrails
Test retrieval recall, citation correctness, grounded answer quality, refusal, sensitive-data leakage and prompt injection. NeMo Guardrails can help structure conversational controls, but tool permissions and data access must be enforced in application code. Log policy outcomes and model versions while minimising retained confidential text.
Operations
Monitor time-to-first-token, generation rate, queue time, VRAM, cache, error and user feedback. Define overload behaviour and an offline maintenance page. A hybrid gateway may route approved non-sensitive tasks externally only after classification and explicit policy; protected prompts fail closed locally.
Primary technical references
- vLLM documentation
- vLLM parallelism and scaling
- Triton dynamic batching guide
- NVIDIA TensorRT-LLM
- NVIDIA NeMo Guardrails documentation
References are provided for software architecture and implementation planning. Validate the versions, licences, support matrix and regulated-use requirements applicable to the final deployment.
From technical concept to production system
Apply this technology through an Axiotech engineering work package
Axiotech can connect the compute, AI or analytics platform to the machine controls, data contracts, validation evidence and lifecycle-support model required for industrial use.
Related Axiotech engineering services