A sustained departmental AI/HPC platform for pharmaceutical modelling, large training campaigns and shared accelerated research.
This brief explains where the baseline fits, the software and operating model it supports, and the evidence Axiotech should use to validate a final configuration. Indicative specifications remain subject to component availability, export compliance and formal quotation.
Baseline architecture
Three 4U nodes; 384 AMD EPYC cores; 3.4TB ECC DDR5; twelve NVIDIA RTX PRO 6000 96GB GPUs; 100GbE compute and 25GbE storage fabrics.
The current compute node is based on the Supermicro AS-4125GS-TNRT / CSE-418G2TS 4U rack platform, integrated with matched rails, power, management, networking and rack infrastructure as required.
Best-fit workloads
Departmental pharmaceutical AI, molecular modelling campaigns, multi-team training, replicated inference and sustained accelerated research across twelve 96GB GPUs.
Twelve high-memory GPUs across three nodes support multiple teams, replicated services and distributed jobs with a dedicated storage path.
Recommended software stack
Kubernetes or Slurm, NVIDIA GPU Operator, CUDA/NCCL, PyTorch/JAX, BioNeMo/GROMACS or domain stacks, MLflow, object/shared storage, Prometheus and central identity.
Pin host drivers and infrastructure separately from versioned application containers. Record source, image, dataset/model and hardware allocation with each benchmark or production release.
Deployment pattern
Treat the system as a managed service: scheduler/control plane, identity, quotas, 100GbE compute fabric, 25GbE storage, registry, experiment metadata, observability and formal maintenance windows.
Define monitoring, identity, backup, change control and workload ownership at the same time as compute. Multi-user platforms require resource allocation and quotas; production services require health, overload and rollback behaviour.
Sizing boundary
At this scale, scheduler policy, data locality, power/cooling, monitoring, security and capacity management become first-class design inputs.
Final sizing should use representative code, data, concurrency and service objectives. Aggregate core, RAM or VRAM figures do not by themselves predict application performance.
Commissioning and acceptance
Run multi-user soak tests, one-/multi-node scaling, RDMA/NCCL health, scheduler fairness, storage saturation, node maintenance and recovery, plus scientific regression for named production workloads.
Axiotech should retain the resulting configuration, firmware/driver baseline, environment manifest, benchmark data and recovery procedure as the system acceptance pack.
Primary technical references
- NVIDIA Container Toolkit overview
- Prometheus overview
- Slurm Quick Start Administrator Guide
- NVIDIA CUDA C++ Best Practices Guide
- NVIDIA NCCL documentation
References are provided for software architecture and implementation planning. Validate the versions, licences, support matrix and regulated-use requirements applicable to the final deployment.