ALVOR
Platform
Advisory
PricingBlog
Get Demo
ALVOR
Platform
Advisory
PricingBlog
Get Demo
AlvorAdvisory
Advisory/HPC Security

HPC Security · NIST SP 800-223 and SP 800-234

Securing the machine without slowing the science.

Specialist security for high-performance computing, from the gap assessment and the zoned architecture through to running it as a managed service.

Book a consultationHow an engagement runs
ManagementSLURMAccessLogin · DTNRDMA fabricdispatchCompute · MPI ranksone barrier, all ranksParallel filesystemOut ofbandoff thehot pathThroughput on the hot path · security on the side
NIST SP 800-223 reference architectureSLURM · MPI · RDMA fabricStrong-scaling preserved

The problem

The most sensitive data you hold runs on the most shared machine you own.

Genetic data, patient data, export-controlled research, models worth more than the building. All of it runs on a system that was built from day one to be shared, fast and open.

01

Hundreds of users

On the same login machines, queuing through the same scheduler.

02

One filesystem

Open research and regulated data sitting on the same storage.

03

Software they install themselves

Research code, compiled on the box, run by the person who wrote it.

04

Accounts that outlive the project

The funding ends. The login usually does not.

Your controls assume a host you administer, software you approved, and a machine you can reboot.

A cluster gives you none of the three.

When teams bring us in

Most HPC teams call us in one of four situations.

Each one starts with a sentence we hear almost word for word. If one of them is yours, here is what we do about it.

  • An audit is coming

    “An auditor is coming and this machine has never been assessed.”

    A funding condition, a data use agreement, or a sponsor who now wants NIST 800-171 or 800-53 evidence. The cluster has run for years and nobody has ever mapped a control to it.

    What we do

    We assess it on the machine itself, hand you the gaps with the evidence behind each one, then design the architecture that closes them without costing the machine its throughput.

  • A dataset with conditions attached

    “We have data that must not leak, and it is on the open cluster.”

    Patient records, controlled-access genomes, export-controlled research, a partner's models. The conditions arrived with the data. The cluster was not built for them.

    What we do

    We design the enclave architecture, so one dataset does not drag the whole machine into scope for the strictest rule you carry.

  • Security and the HPC team disagree

    “Security wants an agent on every node and the HPC team says no.”

    Both of them are right, which is why it has not moved in six months. The catalogue assumes a host you administer and can reboot. This one is neither.

    What we do

    We translate. Every control gets an architectural equivalent that satisfies the auditor without taking the throughput the machine exists to deliver.

  • A new cluster nobody owns yet

    “We have just bought a GPU cluster and nobody owns its security.”

    It arrived as an AI project rather than an infrastructure one. Its own network, its own storage, its own admins, and it is not in anybody's security model yet.

    What we do

    We bring it into the estate on purpose rather than by accident, with an identity, segmentation and logging architecture designed for how a cluster is actually run.

Every one of these needs both halves of the problem at once.

That is the part that is hard to buy.

Why this is unsolved

You have probably asked around. Here is what comes back.

None of these people is being careless. Until the overlay went final on 4 May 2026, there was no authoritative answer to which controls apply here.

01 You ask a security firm

You get a gap report against 800-53.

Put an agent on every node, encrypt everything, patch inside thirty days. Every line of it is correct on paper, and the HPC team will reject all three inside a week.

02 You ask your integrator

You get a machine that is very fast.

Which is what you bought. Security was not in the acceptance criteria, and the honest answer back is that it is a policy question rather than a design one.

03 You ask your own team

You get a documented risk acceptance.

They know the standard and they know the auditor. Nobody has ever handed them a pattern for a scheduler, a parallel filesystem or an interconnect, so the risk gets accepted instead.

Decades of supercomputing

Every site worked it out alone.

9 Feb 2024 · SP 800-223

The architecture. What the machine is, and its threats.

4 May 2026 · SP 800-234

The control overlay. Sixty controls, tailored.

Decades of supercomputing

Every site worked it out alone.

9 Feb 2024 · SP 800-223

The architecture. What the machine is, and its threats.

4 May 2026 · SP 800-234

The control overlay. Sixty controls, tailored.

Sixty of your 800-53 controls, tailored for this machine, final on 4 May 2026.

The control catalogue

Your control catalogue, on a machine that breaks it.

Part one: the controls that work exactly as designed, and cost you the throughput the cluster exists to deliver.

The enterprise controlWhy it breaks hereWhat replaces it
01

EDR agent on every host

It wakes on its own cadence and desynchronises tightly-coupled jobs, and both licence and telemetry multiply by node count.

Petrini, Kerbyson & Pakin, “The Case of the Missing Supercomputer Performance”, SC 2003.

eBPF and sampled instrumentation, out-of-band collection, telemetry drawn from the fabric and the scheduler.

02

Inline firewall, IPS, microsegmentation

RDMA bypasses the kernel, so the appliance either inspects traffic that routes around it or reimposes the latency the fabric was built to remove.

Segmentation designed into the topology and the scheduler, enforced at the zone edge rather than the hot path.

03

Encrypt everywhere, inline DLP

Per-operation crypto and content inspection tax the exact I/O path that strong scaling depends on.

Controller-level and self-encrypting media, project-scoped access, root-squash on exports, and a disciplined staging design.

04

Per-host SIEM forwarding

Log volume multiplied by node count swamps the pipeline, the network and the licence.

Sampled and aggregated at the fabric and the scheduler, shipped out of band.

05

Authenticated vulnerability scanning

Scanning thousands of identical nodes is redundant, and the scanner traffic lands on the interconnect.

Scan the golden image, attest the node at boot, and detect drift from it.

The enterprise control

EDR agent on every host

Why it breaks here

It wakes on its own cadence and desynchronises tightly-coupled jobs, and both licence and telemetry multiply by node count.

Petrini, Kerbyson & Pakin, “The Case of the Missing Supercomputer Performance”, SC 2003.

What replaces it

eBPF and sampled instrumentation, out-of-band collection, telemetry drawn from the fabric and the scheduler.

The enterprise control

Inline firewall, IPS, microsegmentation

Why it breaks here

RDMA bypasses the kernel, so the appliance either inspects traffic that routes around it or reimposes the latency the fabric was built to remove.

What replaces it

Segmentation designed into the topology and the scheduler, enforced at the zone edge rather than the hot path.

The enterprise control

Encrypt everywhere, inline DLP

Why it breaks here

Per-operation crypto and content inspection tax the exact I/O path that strong scaling depends on.

What replaces it

Controller-level and self-encrypting media, project-scoped access, root-squash on exports, and a disciplined staging design.

The enterprise control

Per-host SIEM forwarding

Why it breaks here

Log volume multiplied by node count swamps the pipeline, the network and the licence.

What replaces it

Sampled and aggregated at the fabric and the scheduler, shipped out of band.

The enterprise control

Authenticated vulnerability scanning

Why it breaks here

Scanning thousands of identical nodes is redundant, and the scanner traffic lands on the interconnect.

What replaces it

Scan the golden image, attest the node at boot, and detect drift from it.

And the controls that assume an operating model you do not have.

Part two: these do not fail on performance. They fail on the operating model: who administers the host, who installs the software, and when you are allowed to reboot.

The enterprise controlWhy it breaks hereWhat replaces it
01

MFA and SSO on every login

Hundreds of users share login nodes and one filesystem, and static SSH keys have historically travelled freely between sites.

Short-lived SSH certificates, federated research identity, and multi-factor enforced at the access zone.

02

Patch inside 30 days

You cannot reboot a node under a running job, and driver, firmware and fabric stacks are certified together as a set.

Rolling drain-and-patch through the scheduler, image-based reprovisioning, maintenance tied to allocation cycles.

03

Least privilege per user

Access is granted to projects rather than people, on a shared filesystem, and research software frequently expects to build and run as root.

Project-scoped ACLs, enclave separation for regulated work, and no shared root anywhere.

04

Back everything up

The scratch tier is not backed up, deliberately. Its economics do not permit it and its contents are meant to be regenerable.

A tiering policy that states what is recoverable and what is reproducible, and campaign storage treated differently from scratch.

05

Change control and a software CAB

Researchers install their own stacks with Spack and containers, on the cadence the science demands rather than a change window.

Curated base images, a container policy on Apptainer, and provenance on what actually ran.

The enterprise control

MFA and SSO on every login

Why it breaks here

Hundreds of users share login nodes and one filesystem, and static SSH keys have historically travelled freely between sites.

What replaces it

Short-lived SSH certificates, federated research identity, and multi-factor enforced at the access zone.

The enterprise control

Patch inside 30 days

Why it breaks here

You cannot reboot a node under a running job, and driver, firmware and fabric stacks are certified together as a set.

What replaces it

Rolling drain-and-patch through the scheduler, image-based reprovisioning, maintenance tied to allocation cycles.

The enterprise control

Least privilege per user

Why it breaks here

Access is granted to projects rather than people, on a shared filesystem, and research software frequently expects to build and run as root.

What replaces it

Project-scoped ACLs, enclave separation for regulated work, and no shared root anywhere.

The enterprise control

Back everything up

Why it breaks here

The scratch tier is not backed up, deliberately. Its economics do not permit it and its contents are meant to be regenerable.

What replaces it

A tiering policy that states what is recoverable and what is reproducible, and campaign storage treated differently from scratch.

The enterprise control

Change control and a software CAB

Why it breaks here

Researchers install their own stacks with Spack and containers, on the cadence the science demands rather than a change window.

What replaces it

Curated base images, a container policy on Apptainer, and provenance on what actually ran.

Performance-aware by design

A control you cannot benchmark is a control you cannot trust.

Securing HPC is rarely a question of whether a control is worth having. It is a question of where it goes. Every control has a place on the critical path, and the work is to choose the ones that protect the system without standing between the compute and its throughput. We baseline the machine first, design the controls to sit off the hot path, and prove the throughput held afterwards. If a control costs you strong-scaling, it is either the wrong control or it is in the wrong place.

  • 1We benchmark before and after. A control ships when the regression on your real workloads is measured and accepted, never assumed away.
  • 2Security moves off the critical path: DPU and SmartNIC offload, out-of-band collection, encryption at the controller rather than in the I/O loop.
  • 3The scheduler does the isolating it already knows how to do, through cgroups, per-job constraints, and hardened prolog and epilog, rather than an agent fighting the kernel on every node.

The reference architecture

The machine has four parts. Each one needs different security.

NIST SP 800-223 is the reference architecture the sector has settled on. We design to it, because a control that belongs at the front door will wreck the network the jobs run over.

01

Access zone

The front door

Login nodes, data-transfer nodes, and the science portals. The exposed surface, designed as a Science DMZ so the bulk data path stays fast while the front door itself moves to multi-factor, short-lived SSH certificates, and federated research identity.

SSH certificate authorityCILogonOpen OnDemandGlobus / DTN
02

Management zone

The crown jewels

Provisioning, scheduling, monitoring, and identity, plus the out-of-band BMC and Redfish plane that can power and reimage the whole machine. We isolate it, harden the SLURM control path and its MUNGE trust, and keep it off any network a running job can reach.

SLURM control pathMUNGEBMC · IPMI · RedfishProvisioning
03

Compute zone

The hot path

The compute nodes and the high-speed interconnect, where throughput is the whole point. Segmentation by design rather than inline appliance, per-job isolation through the scheduler, and confidential-computing isolation where a sensitive workload genuinely needs it.

InfiniBand · Slingshot · RoCEcgroup job isolationSEV-SNP · TDXNode attestation
04

Data zone

What feeds it

The parallel and campaign storage that keeps the processors fed. Protection that stays out of the I/O path: controller-level and self-encrypting-media encryption, project-scoped access, root-squash on the exports, and staging that keeps regulated data where it belongs.

LustreIBM Storage ScaleBeeGFSWEKA · VAST

The enclave

Wall off the regulated work. Leave the rest alone.

The usual mistake is to apply your strictest rule to the whole machine. It costs a fortune, it slows everyone down, and it does not make the sensitive data any safer.

The open estateThe enclave, one gate
01

Its own logins

A separate roster you can produce on request, not a subset of everyone.

02

Its own storage

Regulated data never lands on the open store, and cannot leave by accident.

03

Its own way out

What can leave, to where, approved by whom. The export question, answered by design.

04

Its own nodes, when it needs them

A scheduler partition, or separate hardware where the classification demands it.

Who runs on it

A university, a drug company and an AI team run the same machine.

Five kinds of data, five different rules, and one architecture underneath all of them.

The dataWho is holding itWhat binds it
01

Human-subjects and genomic

Universities and research hospitals

Ethics approvals and controlled-access data use agreements, with conditions attached to every copy.

02

Patient and clinical

Health systems and their research partners

The health privacy regime it was collected under, which travels with the data and does not stop at your perimeter.

03

Export-controlled research

National labs and defence suppliers

ITAR and EAR, plus the deemed-export problem a foreign national on a shared cluster creates.

04

Models and training data

AI teams inside ordinary companies

No statute at all. A customer contract, the training data, and the worst day the company has if it leaks.

05

Compound libraries, seismic, simulation

Pharma, energy, manufacturing

Commercial agreements with the same consequences as a regulation, and none of the guidance.

The data

Human-subjects and genomic

Who is holding it

Universities and research hospitals

What binds it

Ethics approvals and controlled-access data use agreements, with conditions attached to every copy.

The data

Patient and clinical

Who is holding it

Health systems and their research partners

What binds it

The health privacy regime it was collected under, which travels with the data and does not stop at your perimeter.

The data

Export-controlled research

Who is holding it

National labs and defence suppliers

What binds it

ITAR and EAR, plus the deemed-export problem a foreign national on a shared cluster creates.

The data

Models and training data

Who is holding it

AI teams inside ordinary companies

What binds it

No statute at all. A customer contract, the training data, and the worst day the company has if it leaks.

The data

Compound libraries, seismic, simulation

Who is holding it

Pharma, energy, manufacturing

What binds it

Commercial agreements with the same consequences as a regulation, and none of the guidance.

We fit the controls to where the data actually is,

not the whole machine for one dataset.

How the work runs

Wherever your cluster is, we pick it up from there.

HPC security is the same four delivery tracks as the rest of the practice, with a performance baseline built into every one. Each stage ends on a decision that stays yours: carry on with us, take it in-house, or stop where you are.

01Assess

Baseline the machine, then map the gaps.

We benchmark the workloads that matter and walk the four zones looking for the soft spots: the flat management network, the static keys, the data-transfer nodes nobody segmented. You get a threat model, a performance baseline, and a ranked list of what to fix first.

HPC threat modelPerformance baselinePrioritised gap register
Explore Assess
02Architect

Design the zones before anyone reconfigures a node.

We design the target architecture on NIST SP 800-223: the zone boundaries, the identity model, the scheduler isolation, and the research enclave for regulated work, with every control placed against its cost on the critical path.

Zoned target architectureControl-to-critical-path mapResearch enclave design
Explore Architect
03Build

Stand it up with your HPC team, not around them.

We implement alongside the people who run the cluster, with throughput as a release gate: a control ships when the benchmark confirms the science still runs as fast. Nothing is marked done on a closed ticket alone.

Controls implementedThroughput regression gateValidated on the benchmark
Explore Build
04Operate

Keep it secure as the machine and the science change.

Allocations turn over, nodes get added, new workloads arrive. We keep the monitoring jitter-aware and out of band, re-benchmark on a schedule, and re-test the zone boundaries so the posture does not drift as the cluster grows.

Jitter-aware monitoringScheduled re-benchmarkBoundary re-testing
Explore Operate

Who does the work

Two kinds of expertise, in the same people.

Which is why it is hard to buy, and why it stalls between two teams who are each half right.

The control catalogue

  • NIST SP 800-53
  • NIST SP 800-171
  • ISO 27001
  • Export control

The machine

  • The scheduler
  • The fabric
  • The parallel filesystem
  • The login plane

This is where the work is

The control catalogue

  • NIST SP 800-53
  • NIST SP 800-171
  • ISO 27001
  • Export control

This is where the work is

The machine

  • The scheduler
  • The fabric
  • The parallel filesystem
  • The login plane

Architects and engineers who have built and run machines like yours, not career consultants. The people advising you have operated a cluster and answered to an auditor for it.

The work behind this

We build the entire export control and
NIST 800-53 HIGH baseline controls program
for a world top 20 supercomputer.

That is about as hard as this work gets. The controls have to satisfy a federal auditor, on a machine where nobody will accept security that slows the science down. The same people do the work described on this page.

Scope
Export control and the full NIST SP 800-53 HIGH baseline
Environment
A world top 20 supercomputer
Status
Live engagement, delivered by us

Named references on request.

We publish no anonymous or invented quotes.

AlvorAdvisory

Bring us your hardest cluster.

A national-scale system, a campus research cluster, or a cloud-burst pipeline: tell us what it runs and what it has to protect. We scope the work in writing, with the performance baseline built in, before anyone touches a node.

Book a consultationSee the four tracks
ALVOR

Security architecture management and compliance: connected into one source of truth.

Security,
Simplified.

Platform

  • Overview
  • AI Assistant
  • Secure by Design
  • Asset Management
  • Risk Management
  • Compliance
  • Policy
  • Security Management
  • Third-Party Risk Management
  • Business Continuity

Capabilities

  • Security Architecture
  • Security Design Review
  • Threat Modeling
  • Dependency Mapping
  • Data Governance
  • Components & SBOM
  • System Security Plan
  • Deployment models

Solutions

  • All solutions
  • CISO
  • Security architect
  • GRC lead
  • Engineering leader
  • Startups
  • Mid-Market
  • Enterprise
  • Regulated & Sovereign
  • Australia

Frameworks

  • ISO 27001
  • SOC 2
  • NIST CSF
  • HIPAA
  • GDPR
  • ISM
  • IRAP
  • Essential Eight
  • ASD Essentials
  • SABSA
  • PCI DSS
  • CMMC
  • FedRAMP
  • Control alignment

Advisory

  • Advisory overview
  • Assess
  • Architect
  • Build
  • Operate
  • All engagements

Company

  • About
  • Blog
  • Learn
  • Security
  • Pricing
  • Compare Alvor

© 2026 Alvor Pty Ltd · ABN 40 700 022 546 · All rights reserved.

PrivacyTermsCookie PolicyVulnerability Disclosure