A shared departmental CPU-compute facility for research pipelines, simulation, genomics and queued multi-user work.
This brief explains where the baseline fits, the software and operating model it supports, and the evidence Axiotech should use to validate a final configuration. Indicative specifications remain subject to component availability, export compliance and formal quotation.
Baseline architecture
Three 4U nodes; 384 AMD EPYC cores; 3.4TB aggregate ECC DDR5; three professional GPUs; switched 100GbE RDMA, 25GbE storage and Slurm.
The current compute node is based on the Supermicro AS-4125GS-TNRT / CSE-418G2TS 4U rack platform, integrated with matched rails, power, management, networking and rack infrastructure as required.
Best-fit workloads
Departmental genomics, CPU-led pharmaceutical pipelines, CFD/FEA queues, research software and shared memory-rich compute across three Slurm-managed nodes.
Three nodes create useful scheduling flexibility: one large job, several independent jobs or a mixture of CPU and GPU-assisted workflows.
Recommended software stack
Slurm, OpenMPI/UCX, Apptainer, Environment Modules or Spack, Nextflow, OpenFOAM/PETSc, monitoring, identity integration and protected shared storage.
Pin host drivers and infrastructure separately from versioned application containers. Record source, image, dataset/model and hardware allocation with each benchmark or production release.
Deployment pattern
Integrate central identity, fair-share/project accounting, Apptainer/Spack or modules, 100GbE compute fabric, 25GbE storage path and protected NAS with explicit scratch/retention policy.
Define monitoring, identity, backup, change control and workload ownership at the same time as compute. Multi-user platforms require resource allocation and quotas; production services require health, overload and rollback behaviour.
Sizing boundary
Shared service operation introduces identity, quota, accounting, backup and software-environment responsibilities that must be designed with the hardware.
Final sizing should use representative code, data, concurrency and service objectives. Aggregate core, RAM or VRAM figures do not by themselves predict application performance.
Commissioning and acceptance
Test scheduler allocation, multi-user contention, one- to three-node MPI, workflow arrays, storage throughput, quotas, accounting, node drain/return and restoration of scheduler configuration.
Axiotech should retain the resulting configuration, firmware/driver baseline, environment manifest, benchmark data and recovery procedure as the system acceptance pack.
Primary technical references
- NVIDIA Container Toolkit overview
- Prometheus overview
- Slurm Quick Start Administrator Guide
- NVIDIA NCCL documentation
References are provided for software architecture and implementation planning. Validate the versions, licences, support matrix and regulated-use requirements applicable to the final deployment.