HPC & AI foundation · Advanced operations

Operating Shared Hyperion Platforms: Slurm, Kubernetes, RDMA, Storage and Security

An operations blueprint for schedulers, GPU allocation, identity, quotas, fabrics, monitoring, backups and controlled change on multi-user Hyperion systems.

A cluster becomes useful when it provides a dependable service to people, not when its nodes merely pass a hardware test. Define ownership, admission, software lifecycle, data protection and incident response before the first shared workload arrives.

Choose the operating model

Slurm is a natural fit for queued batch jobs, MPI and allocation-based research computing. Kubernetes is a natural fit for long-running services, APIs and declarative application operations. Some organisations run both with clear node or time boundaries, but competing control planes must never believe they own the same GPU simultaneously.

Write service objectives: supported hours, maintenance windows, queue policy, recovery targets and who can approve privileged workloads.

Make resources explicit

Schedulers should allocate whole GPUs or supported partitioning mechanisms, CPU cores, memory and local scratch together. Kubernetes discovers accelerators through device plugins; Slurm uses generic resources and cgroups. Prevent untracked interactive work from bypassing the scheduler because it defeats capacity planning and can corrupt performance results.

Fabric and storage operations

Monitor RDMA link state, errors, pause behaviour, retransmits and topology. Keep compute and storage networks logically distinct where the design provides both. GPUDirect and CUDA-aware MPI require a compatible end-to-end stack; include them in post-maintenance validation.

Set project quotas and scratch expiry. Back up authoritative datasets, configuration, identity data, code registries and experiment metadata; do not spend the same protection budget on reproducible temporary files.

Identity, isolation and supply chain

Integrate central identity, use least privilege and separate administrators from ordinary users. Protect container registries, sign images and restrict untrusted privileged containers. Segment management interfaces from user and production networks. Store secrets in a dedicated secrets system, not environment files in shared home directories.

For regulated or export-controlled work, map datasets and model artefacts to authorised users and destinations, keep auditable approvals and apply retention policies.

Observability and capacity management

Collect node health, temperatures, power, ECC events, CPU/RAM/GPU usage, GPU errors, filesystem capacity, queue wait and job outcomes. Prometheus can collect time-series metrics; scheduler accounting provides allocation history. Alert on service-impacting conditions, not every transient fluctuation.

Review demand by project and resource. A high queue does not always mean more GPUs: it may reveal memory, licence, storage or policy constraints.

Change and recovery drill

  1. Snapshot configuration and record firmware/driver versions.
  2. Drain workloads and establish a rollback point.
  3. Apply one controlled layer of change.
  4. Run hardware, storage, RDMA, MPI/NCCL and application smoke tests.
  5. Return capacity gradually and observe error/performance baselines.
  6. Practise restoring scheduler/controller configuration and one representative dataset.

Primary technical references

References are provided for software architecture and implementation planning. Validate the versions, licences, support matrix and regulated-use requirements applicable to the final deployment.

Relevant Hyperion platforms

Hyperion X2160

AxiVector

A tightly coupled two-node simulation and mathematical-computing platform for MPI, CFD, FEA and optimisation.

Two 4U nodes; 256 AMD EPYC cores; 2.3TB aggregate ECC DDR5; two 96GB GPUs; direct 100GbE RoCEv2 RDMA between nodes.

Hyperion X3140

AxiLattice

A shared departmental CPU-compute facility for research pipelines, simulation, genomics and queued multi-user work.

Three 4U nodes; 384 AMD EPYC cores; 3.4TB aggregate ECC DDR5; three professional GPUs; switched 100GbE RDMA, 25GbE storage and Slurm.

Hyperion X2260

AxiCrucible

A two-node multi-user AI workgroup for clinical research, production intelligence, training and resilient service placement.

Two 4U nodes; four NVIDIA RTX PRO 6000 96GB GPUs; 100GbE RDMA; shared protected storage; KVM and Kubernetes-ready infrastructure.

Hyperion X3460

AxiBastion

A sustained departmental AI/HPC platform for pharmaceutical modelling, large training campaigns and shared accelerated research.

Three 4U nodes; 384 AMD EPYC cores; 3.4TB ECC DDR5; twelve NVIDIA RTX PRO 6000 96GB GPUs; 100GbE compute and 25GbE storage fabrics.

Hyperion X4460

AxiTitan

The flagship sovereign AI and research-computing platform for organisation-scale shared services and the largest Hyperion workloads.

Four 4U nodes; 512 AMD EPYC cores; 4.6TB ECC DDR5; sixteen NVIDIA RTX PRO 6000 96GB GPUs; more than 1.5TB aggregate VRAM; 100GbE compute and dedicated storage fabrics.

Configuration and quotation

Validate this workload on Hyperion

Final architecture and price depend on representative code and data, concurrency, storage, networking, site infrastructure, component availability and export compliance.

Request Formal Quotation