HPC & AI foundation · Foundation to programme design

Academic and Research Computing with Hyperion: Reproducibility, Teaching and Shared Facilities

How to turn departmental HPC and AI hardware into a sustainable research and teaching service with fair access, reproducible environments and publishable evidence.

Research computing has three simultaneous jobs: enable discovery, teach people safely and preserve enough evidence to reproduce results. Hyperion systems can start as a principal-investigator node and grow into a scheduled departmental facility while keeping a consistent server, software and support architecture.

Match the platform to the research pattern

AxiFoundry supports method development and GPU teaching; AxiKeystone serves memory-rich CPU workloads; AxiAnvil and AxiForge address large models and accelerated science; AxiVector supports two-node MPI; AxiLattice provides a shared CPU facility; AxiCrucible, AxiBastion and AxiTitan support multi-user AI and mixed services.

Buy against named research codes and datasets. A benchmark appendix is stronger evidence for a grant or capital case than a theoretical peak number.

Create a reproducible software catalogue

Offer versioned modules or Spack environments for compilers, MPI and libraries, plus Apptainer images for complete workflows. Publish supported combinations and deprecation dates. Let projects bring containers through a review path instead of giving every user administrator rights.

Capture job manifests and encourage repositories containing workflow definitions, parameters and environment recipes. Persistent identifiers for datasets and images make papers and student work easier to reconstruct.

Teaching without destabilising research

Use partitions, quotas or reservations for taught sessions. Provide small datasets and bounded jobs that demonstrate scheduling, parallelism and GPU profiling without monopolising the platform. Resettable container environments let students experiment while protecting the host.

Teach performance literacy: serial baseline, speed-up, efficiency, memory fit, I/O and energy. Students should learn why more cores or GPUs can make a badly partitioned workload slower.

Fair access and research governance

Define project sponsorship, user onboarding, storage quota, queue limits, priority and data classification. Publish the rules and make exceptions reviewable. Sensitive clinical, commercial or export-controlled research may require dedicated projects, network controls and audit evidence.

Scheduler accounting can support transparent usage reporting, grant attribution and future capacity proposals without treating utilisation alone as scientific value.

Data management and preservation

Separate high-speed scratch, active project storage and preservation. Link outputs to the inputs, code and environment that created them. Retain enough metadata to rerun important results even when raw temporary intermediates are deleted. Test restore paths and decide who owns data when researchers leave.

A staged facility roadmap

  1. Start with one workload-led node and a documented environment.
  2. Add identity, backup, monitoring and a lightweight allocation process.
  3. Introduce Slurm when users or queues compete.
  4. Add RDMA-connected nodes only after scaling tests demonstrate value.
  5. Formalise service levels, training, governance and lifecycle funding.
  6. Review scientific outcomes and unmet demand before each expansion.

Primary technical references

References are provided for software architecture and implementation planning. Validate the versions, licences, support matrix and regulated-use requirements applicable to the final deployment.

Relevant Hyperion platforms

Hyperion X1140

AxiFoundry

A professional development and proof-of-concept node for CUDA, AI, data science and software engineering.

One 4U node; 96 AMD EPYC cores; 384GB ECC DDR5; one NVIDIA RTX PRO 4000 24GB; enterprise NVMe and remote management.

Hyperion X1160

AxiAnvil

A large-memory single-GPU platform for private LLM inference, retrieval-augmented generation, medical imaging and large scientific models.

One 4U node; 96 AMD EPYC cores; 768GB ECC DDR5; one NVIDIA RTX PRO 6000 96GB GPU; enterprise NVMe.

Hyperion X1260

AxiForge

A professional two-GPU training, visualisation and applied-research server for larger datasets and models.

One 4U node; 96 AMD EPYC cores; 768GB ECC DDR5; two NVIDIA RTX PRO 6000 96GB GPUs; 192GB aggregate VRAM and high-capacity NVMe scratch.

Hyperion X1140-C

AxiKeystone

A CPU- and memory-led HPC platform for simulation, genomics, analytics, compilation and workloads with large in-memory working sets.

One 4U node; 128 AMD EPYC cores; 1.15TB ECC DDR5; one NVIDIA RTX PRO 4000 24GB; enterprise U.2 NVMe.

Hyperion X2160

AxiVector

A tightly coupled two-node simulation and mathematical-computing platform for MPI, CFD, FEA and optimisation.

Two 4U nodes; 256 AMD EPYC cores; 2.3TB aggregate ECC DDR5; two 96GB GPUs; direct 100GbE RoCEv2 RDMA between nodes.

Hyperion X3140

AxiLattice

A shared departmental CPU-compute facility for research pipelines, simulation, genomics and queued multi-user work.

Three 4U nodes; 384 AMD EPYC cores; 3.4TB aggregate ECC DDR5; three professional GPUs; switched 100GbE RDMA, 25GbE storage and Slurm.

Hyperion X2260

AxiCrucible

A two-node multi-user AI workgroup for clinical research, production intelligence, training and resilient service placement.

Two 4U nodes; four NVIDIA RTX PRO 6000 96GB GPUs; 100GbE RDMA; shared protected storage; KVM and Kubernetes-ready infrastructure.

Hyperion X3460

AxiBastion

A sustained departmental AI/HPC platform for pharmaceutical modelling, large training campaigns and shared accelerated research.

Three 4U nodes; 384 AMD EPYC cores; 3.4TB ECC DDR5; twelve NVIDIA RTX PRO 6000 96GB GPUs; 100GbE compute and 25GbE storage fabrics.

Hyperion X4460

AxiTitan

The flagship sovereign AI and research-computing platform for organisation-scale shared services and the largest Hyperion workloads.

Four 4U nodes; 512 AMD EPYC cores; 4.6TB ECC DDR5; sixteen NVIDIA RTX PRO 6000 96GB GPUs; more than 1.5TB aggregate VRAM; 100GbE compute and dedicated storage fabrics.

Configuration and quotation

Validate this workload on Hyperion

Final architecture and price depend on representative code and data, concurrency, storage, networking, site infrastructure, component availability and export compliance.

Request Formal Quotation