A tightly coupled two-node simulation and mathematical-computing platform for MPI, CFD, FEA and optimisation.
This brief explains where the baseline fits, the software and operating model it supports, and the evidence Axiotech should use to validate a final configuration. Indicative specifications remain subject to component availability, export compliance and formal quotation.
Baseline architecture
Two 4U nodes; 256 AMD EPYC cores; 2.3TB aggregate ECC DDR5; two 96GB GPUs; direct 100GbE RoCEv2 RDMA between nodes.
The current compute node is based on the Supermicro AS-4125GS-TNRT / CSE-418G2TS 4U rack platform, integrated with matched rails, power, management, networking and rack infrastructure as required.
Best-fit workloads
Two-node CFD/FEA, PETSc and MPI, distributed mathematical workloads, optimisation campaigns and high-fidelity engineering simulation that benefits from 100GbE RDMA.
It is sized for applications that genuinely benefit from a second NUMA domain and low-latency inter-node communication without a larger switched fabric.
Recommended software stack
Linux, OpenMPI or vendor MPI, UCX, OpenFOAM, PETSc, Slurm, Apptainer, CUDA-aware libraries, parallel filesystems/NAS clients and performance profilers.
Pin host drivers and infrastructure separately from versioned application containers. Record source, image, dataset/model and hardware allocation with each benchmark or production release.
Deployment pattern
Operate with Slurm or disciplined MPI launch, UCX/RDMA validation, shared project storage and node-local scratch. Establish a one-node reference before enabling distributed runs.
Define monitoring, identity, backup, change control and workload ownership at the same time as compute. Multi-user platforms require resource allocation and quotas; production services require health, overload and rollback behaviour.
Sizing boundary
Poorly partitioned or memory-latency-bound codes may not scale across nodes. Establish a one-node baseline and verify MPI efficiency before production sizing.
Final sizing should use representative code, data, concurrency and service objectives. Aggregate core, RAM or VRAM figures do not by themselves predict application performance.
Commissioning and acceptance
Run numerical regression, fixed-problem and growing-problem scaling tests, inspect RDMA counters, verify checkpoint/restart and document the problem sizes at which two nodes outperform one.
Axiotech should retain the resulting configuration, firmware/driver baseline, environment manifest, benchmark data and recovery procedure as the system acceptance pack.
Primary technical references
- NVIDIA Container Toolkit overview
- Prometheus overview
- Slurm Quick Start Administrator Guide
- NVIDIA CUDA C++ Best Practices Guide
- NVIDIA NCCL documentation
References are provided for software architecture and implementation planning. Validate the versions, licences, support matrix and regulated-use requirements applicable to the final deployment.