Application note · Specialised

Application Note: Production Variant Calling and Long-Read Assembly

A reproducible Nextflow/Slurm pattern for sample throughput, memory-heavy assembly, accelerated stages and governed genomic data.

This application pattern covers a mixed genomics service: routine short-read germline processing, occasional cohort operations and long-read assembly. It deliberately separates throughput scheduling from memory-heavy jobs and treats workflow provenance as part of the deliverable.

Reference architecture

Use AxiKeystone for method development and memory-heavy single-node work; AxiLattice for the production Slurm queue; and AxiForge for validated GPU-accelerated or imaging/protein stages. Protected NAS holds raw reads, references and approved outputs. Per-job NVMe holds sorting, temporary BAM/CRAM and assembler scratch.

Workflow stages

  1. Validate sample sheet, checksums and metadata.
  2. Run read QC and contamination/identity checks.
  3. Align, sort and mark/process reads in immutable containers.
  4. Call per-sample variants, then joint genotype where required.
  5. Apply approved filters and annotation.
  6. For long reads, reserve a high-memory node for assembly/polishing.
  7. Publish results with a run manifest and QC report.

Scheduler profile

Create resource labels for standard CPU, high-memory and GPU stages. Let Nextflow request the correct Slurm partition and memory, rather than embedding hostnames. Cap concurrency for metadata-heavy stages and pre-stage reference bundles to node-local storage. Use retry only for clearly transient failures; scientific tool errors need review.

Acceptance evidence

Run recognised truth or internal reference samples and compare sensitivity/precision, genotype concordance, coverage, contamination and structural metrics. Record wall time, samples/day, peak RAM and scratch. Validate restore, interrupted-run resume and reference/version rollback.

Operational controls

Encrypt governed storage, audit access and restrict exports. Do not include identifiers in generic monitoring labels. Separate research findings from clinically reportable outputs. Revalidate when changing caller, reference, container, major driver/runtime or workflow logic.

Primary technical references

References are provided for software architecture and implementation planning. Validate the versions, licences, support matrix and regulated-use requirements applicable to the final deployment.

Relevant Hyperion platforms

Hyperion X1260

AxiForge

A professional two-GPU training, visualisation and applied-research server for larger datasets and models.

One 4U node; 96 AMD EPYC cores; 768GB ECC DDR5; two NVIDIA RTX PRO 6000 96GB GPUs; 192GB aggregate VRAM and high-capacity NVMe scratch.

Hyperion X1140-C

AxiKeystone

A CPU- and memory-led HPC platform for simulation, genomics, analytics, compilation and workloads with large in-memory working sets.

One 4U node; 128 AMD EPYC cores; 1.15TB ECC DDR5; one NVIDIA RTX PRO 4000 24GB; enterprise U.2 NVMe.

Hyperion X3140

AxiLattice

A shared departmental CPU-compute facility for research pipelines, simulation, genomics and queued multi-user work.

Three 4U nodes; 384 AMD EPYC cores; 3.4TB aggregate ECC DDR5; three professional GPUs; switched 100GbE RDMA, 25GbE storage and Slurm.

Configuration and quotation

Validate this workload on Hyperion

Final architecture and price depend on representative code and data, concurrency, storage, networking, site infrastructure, component availability and export compliance.

Request Formal Quotation