HPC Security · NIST SP 800-223 and SP 800-234
Specialist security for high-performance computing, from the gap assessment and the zoned architecture through to running it as a managed service.
The problem
Genetic data, patient data, export-controlled research, models worth more than the building. All of it runs on a system that was built from day one to be shared, fast and open.
On the same login machines, queuing through the same scheduler.
Open research and regulated data sitting on the same storage.
Research code, compiled on the box, run by the person who wrote it.
The funding ends. The login usually does not.
Your controls assume a host you administer, software you approved, and a machine you can reboot.
A cluster gives you none of the three.
When teams bring us in
Each one starts with a sentence we hear almost word for word. If one of them is yours, here is what we do about it.
An audit is coming
An auditor is coming and this machine has never been assessed.”
A funding condition, a data use agreement, or a sponsor who now wants NIST 800-171 or 800-53 evidence. The cluster has run for years and nobody has ever mapped a control to it.
What we do
We assess it on the machine itself, hand you the gaps with the evidence behind each one, then design the architecture that closes them without costing the machine its throughput.
A dataset with conditions attached
We have data that must not leak, and it is on the open cluster.”
Patient records, controlled-access genomes, export-controlled research, a partner's models. The conditions arrived with the data. The cluster was not built for them.
What we do
We design the enclave architecture, so one dataset does not drag the whole machine into scope for the strictest rule you carry.
Security and the HPC team disagree
Security wants an agent on every node and the HPC team says no.”
Both of them are right, which is why it has not moved in six months. The catalogue assumes a host you administer and can reboot. This one is neither.
What we do
We translate. Every control gets an architectural equivalent that satisfies the auditor without taking the throughput the machine exists to deliver.
A new cluster nobody owns yet
We have just bought a GPU cluster and nobody owns its security.”
It arrived as an AI project rather than an infrastructure one. Its own network, its own storage, its own admins, and it is not in anybody's security model yet.
What we do
We bring it into the estate on purpose rather than by accident, with an identity, segmentation and logging architecture designed for how a cluster is actually run.
Every one of these needs both halves of the problem at once.
That is the part that is hard to buy.
Why this is unsolved
None of these people is being careless. Until the overlay went final on 4 May 2026, there was no authoritative answer to which controls apply here.
You get a gap report against 800-53.
Put an agent on every node, encrypt everything, patch inside thirty days. Every line of it is correct on paper, and the HPC team will reject all three inside a week.
You get a machine that is very fast.
Which is what you bought. Security was not in the acceptance criteria, and the honest answer back is that it is a policy question rather than a design one.
You get a documented risk acceptance.
They know the standard and they know the auditor. Nobody has ever handed them a pattern for a scheduler, a parallel filesystem or an interconnect, so the risk gets accepted instead.
Decades of supercomputing
Every site worked it out alone.
9 Feb 2024 · SP 800-223
The architecture. What the machine is, and its threats.
4 May 2026 · SP 800-234
The control overlay. Sixty controls, tailored.
Decades of supercomputing
Every site worked it out alone.
Sixty of your 800-53 controls, tailored for this machine, final on 4 May 2026.
The control catalogue
Part one: the controls that work exactly as designed, and cost you the throughput the cluster exists to deliver.
EDR agent on every host
It wakes on its own cadence and desynchronises tightly-coupled jobs, and both licence and telemetry multiply by node count.
Petrini, Kerbyson & Pakin, “The Case of the Missing Supercomputer Performance”, SC 2003.
eBPF and sampled instrumentation, out-of-band collection, telemetry drawn from the fabric and the scheduler.
Inline firewall, IPS, microsegmentation
RDMA bypasses the kernel, so the appliance either inspects traffic that routes around it or reimposes the latency the fabric was built to remove.
Segmentation designed into the topology and the scheduler, enforced at the zone edge rather than the hot path.
Encrypt everywhere, inline DLP
Per-operation crypto and content inspection tax the exact I/O path that strong scaling depends on.
Controller-level and self-encrypting media, project-scoped access, root-squash on exports, and a disciplined staging design.
Per-host SIEM forwarding
Log volume multiplied by node count swamps the pipeline, the network and the licence.
Sampled and aggregated at the fabric and the scheduler, shipped out of band.
Authenticated vulnerability scanning
Scanning thousands of identical nodes is redundant, and the scanner traffic lands on the interconnect.
Scan the golden image, attest the node at boot, and detect drift from it.
EDR agent on every host
It wakes on its own cadence and desynchronises tightly-coupled jobs, and both licence and telemetry multiply by node count.
Petrini, Kerbyson & Pakin, “The Case of the Missing Supercomputer Performance”, SC 2003.
eBPF and sampled instrumentation, out-of-band collection, telemetry drawn from the fabric and the scheduler.
Inline firewall, IPS, microsegmentation
RDMA bypasses the kernel, so the appliance either inspects traffic that routes around it or reimposes the latency the fabric was built to remove.
Segmentation designed into the topology and the scheduler, enforced at the zone edge rather than the hot path.
Encrypt everywhere, inline DLP
Per-operation crypto and content inspection tax the exact I/O path that strong scaling depends on.
Controller-level and self-encrypting media, project-scoped access, root-squash on exports, and a disciplined staging design.
Per-host SIEM forwarding
Log volume multiplied by node count swamps the pipeline, the network and the licence.
Sampled and aggregated at the fabric and the scheduler, shipped out of band.
Authenticated vulnerability scanning
Scanning thousands of identical nodes is redundant, and the scanner traffic lands on the interconnect.
Scan the golden image, attest the node at boot, and detect drift from it.
Part two: these do not fail on performance. They fail on the operating model: who administers the host, who installs the software, and when you are allowed to reboot.
MFA and SSO on every login
Hundreds of users share login nodes and one filesystem, and static SSH keys have historically travelled freely between sites.
Short-lived SSH certificates, federated research identity, and multi-factor enforced at the access zone.
Patch inside 30 days
You cannot reboot a node under a running job, and driver, firmware and fabric stacks are certified together as a set.
Rolling drain-and-patch through the scheduler, image-based reprovisioning, maintenance tied to allocation cycles.
Least privilege per user
Access is granted to projects rather than people, on a shared filesystem, and research software frequently expects to build and run as root.
Project-scoped ACLs, enclave separation for regulated work, and no shared root anywhere.
Back everything up
The scratch tier is not backed up, deliberately. Its economics do not permit it and its contents are meant to be regenerable.
A tiering policy that states what is recoverable and what is reproducible, and campaign storage treated differently from scratch.
Change control and a software CAB
Researchers install their own stacks with Spack and containers, on the cadence the science demands rather than a change window.
Curated base images, a container policy on Apptainer, and provenance on what actually ran.
MFA and SSO on every login
Hundreds of users share login nodes and one filesystem, and static SSH keys have historically travelled freely between sites.
Short-lived SSH certificates, federated research identity, and multi-factor enforced at the access zone.
Patch inside 30 days
You cannot reboot a node under a running job, and driver, firmware and fabric stacks are certified together as a set.
Rolling drain-and-patch through the scheduler, image-based reprovisioning, maintenance tied to allocation cycles.
Least privilege per user
Access is granted to projects rather than people, on a shared filesystem, and research software frequently expects to build and run as root.
Project-scoped ACLs, enclave separation for regulated work, and no shared root anywhere.
Back everything up
The scratch tier is not backed up, deliberately. Its economics do not permit it and its contents are meant to be regenerable.
A tiering policy that states what is recoverable and what is reproducible, and campaign storage treated differently from scratch.
Change control and a software CAB
Researchers install their own stacks with Spack and containers, on the cadence the science demands rather than a change window.
Curated base images, a container policy on Apptainer, and provenance on what actually ran.
Performance-aware by design
Securing HPC is rarely a question of whether a control is worth having. It is a question of where it goes. Every control has a place on the critical path, and the work is to choose the ones that protect the system without standing between the compute and its throughput. We baseline the machine first, design the controls to sit off the hot path, and prove the throughput held afterwards. If a control costs you strong-scaling, it is either the wrong control or it is in the wrong place.
The reference architecture
NIST SP 800-223 is the reference architecture the sector has settled on. We design to it, because a control that belongs at the front door will wreck the network the jobs run over.
Login nodes, data-transfer nodes, and the science portals. The exposed surface, designed as a Science DMZ so the bulk data path stays fast while the front door itself moves to multi-factor, short-lived SSH certificates, and federated research identity.
Provisioning, scheduling, monitoring, and identity, plus the out-of-band BMC and Redfish plane that can power and reimage the whole machine. We isolate it, harden the SLURM control path and its MUNGE trust, and keep it off any network a running job can reach.
The compute nodes and the high-speed interconnect, where throughput is the whole point. Segmentation by design rather than inline appliance, per-job isolation through the scheduler, and confidential-computing isolation where a sensitive workload genuinely needs it.
The parallel and campaign storage that keeps the processors fed. Protection that stays out of the I/O path: controller-level and self-encrypting-media encryption, project-scoped access, root-squash on the exports, and staging that keeps regulated data where it belongs.
The enclave
The usual mistake is to apply your strictest rule to the whole machine. It costs a fortune, it slows everyone down, and it does not make the sensitive data any safer.
A separate roster you can produce on request, not a subset of everyone.
Regulated data never lands on the open store, and cannot leave by accident.
What can leave, to where, approved by whom. The export question, answered by design.
A scheduler partition, or separate hardware where the classification demands it.
Who runs on it
Five kinds of data, five different rules, and one architecture underneath all of them.
Human-subjects and genomic
Universities and research hospitals
Ethics approvals and controlled-access data use agreements, with conditions attached to every copy.
Patient and clinical
Health systems and their research partners
The health privacy regime it was collected under, which travels with the data and does not stop at your perimeter.
Export-controlled research
National labs and defence suppliers
ITAR and EAR, plus the deemed-export problem a foreign national on a shared cluster creates.
Models and training data
AI teams inside ordinary companies
No statute at all. A customer contract, the training data, and the worst day the company has if it leaks.
Compound libraries, seismic, simulation
Pharma, energy, manufacturing
Commercial agreements with the same consequences as a regulation, and none of the guidance.
Human-subjects and genomic
Universities and research hospitals
Ethics approvals and controlled-access data use agreements, with conditions attached to every copy.
Patient and clinical
Health systems and their research partners
The health privacy regime it was collected under, which travels with the data and does not stop at your perimeter.
Export-controlled research
National labs and defence suppliers
ITAR and EAR, plus the deemed-export problem a foreign national on a shared cluster creates.
Models and training data
AI teams inside ordinary companies
No statute at all. A customer contract, the training data, and the worst day the company has if it leaks.
Compound libraries, seismic, simulation
Pharma, energy, manufacturing
Commercial agreements with the same consequences as a regulation, and none of the guidance.
We fit the controls to where the data actually is,
not the whole machine for one dataset.
How the work runs
HPC security is the same four delivery tracks as the rest of the practice, with a performance baseline built into every one. Each stage ends on a decision that stays yours: carry on with us, take it in-house, or stop where you are.
We benchmark the workloads that matter and walk the four zones looking for the soft spots: the flat management network, the static keys, the data-transfer nodes nobody segmented. You get a threat model, a performance baseline, and a ranked list of what to fix first.
We design the target architecture on NIST SP 800-223: the zone boundaries, the identity model, the scheduler isolation, and the research enclave for regulated work, with every control placed against its cost on the critical path.
We implement alongside the people who run the cluster, with throughput as a release gate: a control ships when the benchmark confirms the science still runs as fast. Nothing is marked done on a closed ticket alone.
Allocations turn over, nodes get added, new workloads arrive. We keep the monitoring jitter-aware and out of band, re-benchmark on a schedule, and re-test the zone boundaries so the posture does not drift as the cluster grows.
Who does the work
Which is why it is hard to buy, and why it stalls between two teams who are each half right.
This is where the work is
Architects and engineers who have built and run machines like yours, not career consultants. The people advising you have operated a cluster and answered to an auditor for it.
The work behind this
That is about as hard as this work gets. The controls have to satisfy a federal auditor, on a machine where nobody will accept security that slows the science down. The same people do the work described on this page.
Named references on request.
We publish no anonymous or invented quotes.
A national-scale system, a campus research cluster, or a cloud-burst pipeline: tell us what it runs and what it has to protect. We scope the work in writing, with the performance baseline built in, before anyone touches a node.