Back to DashboardHardware InfrastructurePhase 6 · Infrastructure & Performanceadvanced60 min

PQC Hardware Acceleration

How post-quantum signatures are accelerated on CPUs, GPUs, FPGAs, ASICs and NPUs — backed by our own measurements on three generations of Arm cores and an FPGA.

Why this matters: ML-DSA and SLH-DSA spend their time in completely different places, so the same accelerator can give +29% on one and more than 40× on the other. Choosing hardware, parameter sets and hash variants for a device — or judging a vendor’s acceleration claim — needs an honest picture of where the time goes and what moving it costs.

Start here: Click through the seven acceleration models: each opens with an everyday analogy, then an animated diagram of the real mechanism, then how well it fits ML-DSA, SLH-DSA, single operations and big batches.

Practice in the Simulation

Where the Time Goes

“Accelerating PQC” is not one problem. Each algorithm spends its time in a different place, and an accelerator only helps if it speeds up that place — and if moving the work there costs less than it saves.

(lattice)

Polynomial arithmetic mod 8,380,417 via the NTT, plus a lot of SHAKE to grow the public matrix and sample secrets, inside a retry loop that runs about 4–5 times per signature on average. Already fast: milliseconds on a small core.

SLH-DSA (hash-based)

Almost nothing but hashing — about two million Keccak permutations for one SHAKE-128s signature — built into large Merkle trees. Slow: seconds on a small core for the compact “s” parameter sets.

RSA / ECC (classical)

Big-integer multiplication with Montgomery reduction. Still in every hybrid deployment during the migration, so still worth accelerating.

Measured by PQC Today
On the i.MX 95, ML-DSA-65 signs 5,800 times per second and SLH-DSA-SHAKE-128s 0.94 times — a gap of more than 6,000× (same engine, 2026-09-27). The same accelerator effort therefore buys very different results: see “ML-DSA vs SLH-DSA” below.

The Five Building Blocks

Hardware rarely accelerates “an algorithm”. It accelerates the primitives underneath:

  • Montgomery reduction — computes “mod N” with multiplications and a shift instead of a slow division (Montgomery, 1985; FIPS 204 Appendix A). Inside RSA, ECC and every ML-DSA coefficient multiply.
  • NTT (Number-Theoretic Transform) — turns polynomial multiplication from 65,536 products into about 3,300 for 256 coefficients. ML-DSA runs a full 8-layer NTT; ML-KEM’s modulus 3,329 only allows 7 layers.
  • Keccak-f[1600] — a 24-round shuffle of a 1,600-bit state using only XOR, AND, NOT and rotate (FIPS 202). Constant-time by nature and cheap in hardware.
  • SHA-3 — Keccak in a “sponge”: absorb the message into the 1,088-bit rate, keep a 512-bit capacity sealed, output a fixed 256 bits (SHA3-256).
  • SHAKE — the same sponge with an output tap you can leave open. SHAKE128 (capacity 256) grows ML-DSA’s and ML-KEM’s matrices; SHAKE256 is every hash in SLH-DSA-SHAKE.

Workshop step 2 explains each one in plain English, with a worked Montgomery example you can check by hand and an animated NTT butterfly network.

Seven Models of Acceleration

From closest-to-the-CPU to furthest away: scalar code; SIMD vector units (Arm NEON/SVE, x86 AVX2/AVX-512); dedicated crypto instructions (AES, SHA-256, SHA-512, SHA-3); GPUs; FPGAs; ASICs and secure elements; and NPUs. The further from the CPU, the more raw speed is available — and the more you pay to move data there and back.

FeatureArmIntel / AMD x86Where PQC uses it
Wide SIMD
Same operation on many numbers at once
NEON 128-bit (all Armv8); SVE/SME on newer coresAVX2 256-bit; AVX-512 512-bitML-DSA / ML-KEM NTT and polynomial arithmetic; several SLH-DSA hash chains side by side
[1] · [2]
AES
One AES round per instruction
FEAT_AES (optional from Armv8.0)AES-NI; VAES (AES on AVX/AVX-512 registers)Symmetric encryption after a PQC key exchange; AES-based schemes
[1] · [2]
SHA-256
SHA-256 rounds in hardware
FEAT_SHA256 (optional from Armv8.0) — on A53, A55, Apple M-seriesSHA-NI (SHA-1 and SHA-256 only)SLH-DSA-SHA2 (all of category 1; F and PRF at every category), LMS/HSS
[1] · [2] · [3] · [4]
SHA-512
SHA-512 rounds in hardware
FEAT_SHA512 (Armv8.2 extension) — NOT on A53/A55; on Apple M4VSHA512* — only Arrow Lake-S / Lunar Lake and laterSLH-DSA-SHA2 categories 3 and 5 (H_msg, PRF_msg, H, T)
[1] · [2] · [3]
SHA-3 / Keccak
EOR3, RAX1, XAR, BCAX fuse Keccak’s θ, ρ and χ steps
FEAT_SHA3 (Armv8.2 extension) — NOT on A53/A55; on Apple M4None — no SHA-3 or Keccak instruction exists; SIMD runs 4–8 Keccak states in parallelEvery SHAKE call: ML-KEM, ML-DSA sampling, SLH-DSA-SHAKE
[1] · [2] · [3]
Big-integer multiply
Wide multiply-accumulate for bignum math
Scalar MUL/UMULH; NEON UMULLMULX/ADX; AVX-512 IFMA (eight 52-bit multiplies per instruction)Classical RSA/ECC during hybrid migration — not needed by lattice or hash PQC
Published, not measured by us
Instruction availability per Arm’s feature list and core manuals, Intel’s SHA Extensions note and LLVM’s CPU definitions. Note the asymmetry for PQC: x86 has no SHA-3/Keccak instruction at all, and hardware SHA-512 reached x86 only with Arrow Lake-S and Lunar Lake (2024); Arm has had optional SHA-3 and SHA-512 instructions since the Armv8.2 extension — but the Cortex-A53 and A55 do not implement them.

Arm Across Three Generations — Our Boards

We run the same Rust PKCS#11 engine on three Arm platforms: an Apple M4 Pro laptop (Armv9.2-A, with SHA-3, SHA-512 and SME2), an NXP i.MX 95 appliance (6× Cortex-A55, Armv8.2-A) and an AMD Kria KV260 appliance (4× Cortex-A53, Armv8.0-A, plus FPGA fabric). The A55 and A53 both have AES and SHA-256 instructions and neither has SHA-3 or SHA-512.

Measured by PQC Today
  • ML-DSA-65 signing: M4 Pro 72,776/s · i.MX 95 5,800/s · KV260 3,268/s (same engine, 4 workers, 2026-09-27).
  • The instruction set decides which variant wins. On the A55 and A53, SLH-DSA-SHA2-128s signs about 11× faster than SLH-DSA-SHAKE-128s on today’s engine (hardware SHA-256 plus SHA-2-specific software work vs software Keccak). Digest throughput at 16 KiB through the PKCS#11 engine on the boards, 2026-09-26. Cortex-A5x has SHA-256 instructions but not SHA-512 ones, so the usual software ordering inverts. (SHA-256 396–470 MB/s vs SHA-512 136–143 MB/s.)
  • Same instructions, newer core: one worker signing ML-DSA-44, the A55 does 255.6/s and the A53 186/s (+37%). A newer microarchitecture and a higher clock lift everything; a new instruction lifts only what uses it.
Two Apple generations: M4 Pro and M5 Max

Newer is not simply “more of the same”. The M4 Pro has 10 big cores and 4 small efficiency cores. The M5 Max drops the efficiency cores: it has 6 top-tier “super cores” and 12 new mid-tier performance cores. A single signature runs on one core, so its speed depends on which kind of core it lands on; a batch spread over all cores depends on how many cores of each kind there are.

Published, not measured by us
Apple M4 Pro: up to 14 CPU cores — 10 performance and 4 efficiency. Apple M5 Max: 18 cores — 6 “super cores” (M5’s next-generation performance core, with more front-end bandwidth, a new cache hierarchy and better branch prediction) and 12 “all-new performance cores” tuned for power-efficient multithreaded work; no efficiency cores. Apple quotes up to 15% higher multithreaded performance for M5 Max over M4 Max and gives no single-thread figure. Third-party measurements put the M4 Pro’s big cores near 4.5 GHz, the M5’s super cores near 4.6 GHz and its new performance cores near 4.4 GHz.

Workshop step 3 lets you compare every algorithm across the three platforms.

What We Built on Arm — AES, ML-DSA, SLH-DSA, AWS-LC

Every acceleration below is real work in our engine or appliance images, with the measured before → after. The coloured tag says which acceleration model it is — and notice how many of the biggest wins are plain software engineering, not new hardware.

AES

What we didModelMeasuredStatus
Switch on the Armv8 AES and PMULL instructions
Build flags for the Rust AES crates (they were silently off)
Crypto instr.
AES-128-CBC 16 KiB: KV260 7.4×, i.MX 95 7.8×, M4 Pro 10.8×; AES-256-CBC up to 12.9×; GCM 3–6×
All three platforms, A-B-A-B, 2026-09-26
Shipped
cacp #36/#38, sandbox #83/#84
Move to the newer AES crates that use the hardware by default
Runtime CPU detection instead of a build flag
Crypto instr.
CBC unchanged (0.98–0.99×), GCM 1.15–1.19× — same speed, no flag to forget
M4 Pro, 2026-09-27
Merged, next image
hsm #283
Process GCM a whole block at a time instead of byte by byte
Whole-block counter keystream and GHASH
Software
AES-128-GCM 16 KiB: KV260 board 6.1–7.6× (5,455 → 35,603/s), i.MX 95 Pro board 6.5–7.0× (11,743 → 76,682/s), M4 Pro 5.6–6.5×
KV260 + i.MX 95 Pro on-board A/B, M4 Pro, 2026-09-27
Shipped
hsm #291
Stop cores queuing on shared engine state
Per-session state sharded 64 ways; AES key schedule cached
Software
AES-128-CBC 64 B: KV260 board 3.2–3.3× (46k → 150k/s), i.MX 95 Pro board 4.1–4.4× (up to 5.0× at 6 workers, 87k → 435k/s). The old engine did not scale past 1 worker; the new one scales near-linearly. Single worker 5–18% lower (known trade-off, follow-up planned)
KV260 + i.MX 95 Pro on-board A/B, M4 Pro, 2026-09-27
Shipped
hsm #298
Compile the appliance engine at full optimisation
opt-level 3 instead of size-optimised
Software
AES-CBC 16 KiB 1.40–1.53×, 64 B 1.23×
M4 Pro, 2026-09-27
Shipped
cacp recipes

ML-DSA

What we didModelMeasuredStatus
Run ML-DSA and ML-KEM on AWS-LC’s hand-written NEON assembly
mldsa-native / mlkem-native AArch64 code inside AWS-LC
SIMD
On/off on the M4 Pro (same engine, AWS-LC path disabled vs enabled): ML-DSA-65 sign 18.2k → 72.7k/s (4.0×), verify 4.0×, keygen 7.7×; ML-DSA-87 sign 3.4×; ML-KEM-768 decapsulate 1.5×. Earlier container test with SHA-3 masked (A53/A55-like): ML-DSA-65 sign 3.8×, verify 3.3×
M4 Pro A-B-A-B, 2026-09-27; container, 2026-09-24
Shipped
hsm #254 (98b67e63)
Decode each private key and expand its matrix once, not on every signature
Per-key expanded signing context, cached
Software
Host CPU per FPGA-assisted signature 0.252 → 0.041 of a CPU-only signature
arm64 container model
Shipped
hsm #254 (e2246faf)
Offload whole ML-DSA-65 signatures to two FPGA signers
DMA + one system call per signature + interrupt
FPGA
404 → 523 sign/s at 4 workers (+29%); later 515 → 975 at 8 threads with the host-path work
KV260 board
Shipped
hsm #254, cacp images

SLH-DSA

What we didModelMeasuredStatus
Use the Armv8 SHA-256/SHA-512 instructions for the SHA-2 variants
Rust sha2 crate “asm” feature, Arm only
Crypto instr.
SHA2-128s sign 568 → 102 ms; SHAKE control unchanged (49.0 → 47.4 ms)
arm64 container on a Mac
Shipped
hsm 490988a3
Hash the public seed block once per signature, not once per hash call
SHA-256/512 midstate cache
Software
SHA2-128s sign 98.0 → 79.6 ms, SHA2-192s 229.9 → 153.6 ms
arm64 container on a Mac
Shipped
hsm f2e195fa
Build the independent subtrees of one signature on all cores
Multi-threaded signing with one shared core budget
Software
SHA2-128s 76.9 → 19.8 ms (4 threads); SHAKE-128s 735.8 → 195.0 ms
arm64 container on a Mac
Shipped
hsm 42291340
Use the Armv8.2 SHA-3 instructions for SHAKE (M4-class cores)
Rust keccak crate “asm” feature — tested as an A/B build, not yet in the engine
Crypto instr.
M4 Pro: SHAKE 1.17–1.20×, SLH-DSA-SHAKE-128s sign 14.5 → 17.5/s (1.19×), SHA3-256 digest 1.14–1.15×; SHA-2 controls 1.00×. A55/A53 lack the instructions, so no effect there
M4 Pro A-B-A-B, 2026-09-27 (M5 Max partial run agrees: 1.17–1.22×)
Planned
A/B build of hsm 476f97d1
Offload whole SLH-DSA-SHAKE signatures to a 4-lane Keccak engine
4 Keccak lanes at 240 MHz in the KV260 fabric
FPGA
Same board, FPGA on vs off: SLH-DSA-SHAKE-128s sign 0.46 → 13.0/s (28.4×), 192s 30.6×, 256s 26.4×, keygen 14–15×; SHAKE-f and SHA-2 sets unchanged (not routed to the fabric)
KV260 board, 2026-09-27 (re-proves the 09-25 result: 2.12 s → 69 ms)
Shipped
hsm #254, cacp #34

AWS-LC (RSA)

What we didModelMeasuredStatus
Move RSA (and NIST-curve ECDH) onto AWS-LC
aws-lc-rs: optimised, constant-time big-number code
Software
RSA-2048 sign 49.4 → 68.6/s (1 thread), 235.6 → 326.8/s (6 threads)
i.MX 95 board
Shipped
hsm #239
Stop re-parsing the RSA key on every operation
Parsed-key cache
Software
RSA-2048 sign 80.8 → 105.4/s (1 thread), 356 → 481/s (6 threads)
i.MX 95 board, pinned core
Shipped
hsm #240
Answer the decrypt “how big is the output?” query without decrypting
Return the size from the key instead of doing a private-key operation
Software
RSA-OAEP-2048 decrypt: KV260 board 167 → 347/s (2.05–2.11×), i.MX 95 Pro board 321 → 642/s (2.00×), M4 Pro 4,144 → 8,295/s
KV260 + i.MX 95 Pro on-board A/B, M4 Pro, 2026-09-27
Shipped
hsm #289
Use the scalar Montgomery kernel on Cortex-A5x cores
Patch AWS-LC’s kernel choice (it picks a server-tuned NEON kernel that is slower on A5x)
Software
On-board, stock vs patched engine from one commit: RSA-OAEP decrypt 2048/3072/4096 — i.MX 95 (A55) 1.27× / 1.60× / 1.21–1.26×, KV260 (A53) 1.28–1.39× / 1.64–1.69× / 1.22–1.25×. RSA-PSS sign unaffected (pure Rust)
i.MX 95 + KV260 boards, 2026-09-27
In review
AWS-LC dispatch patch (hsm PR in review)

Container and M4 Pro figures isolate one change at a time; the boards only have bundled before/after figures (e.g. KV260 SLH-DSA-SHA2-128s 9.4 s → 191 ms, SHAKE-128s 12.3 s → 2.12 s, all CPU changes together). Items marked “merged” reach the boards with the next image build.

ML-DSA vs SLH-DSA: Different Bottlenecks

ML-DSA — vectorise it

256 coefficients all get the same treatment, which is exactly what SIMD does best, inside the CPU with no transfer cost. Offload is hard: the retry loop, key decoding and SHAKE sampling all bounce between stages.

SLH-DSA — accelerate the hash

Millions of small, independent hash calls. Hash instructions help directly (if the variant matches the core’s hash), SIMD can run several tree paths side by side, and a hash engine that takes a whole signature per call can shine.

Published, not measured by us
Reference C vs AVX2 on Skylake: ML-DSA’s NTT-based polynomial multiplication is about 4.5× faster, and a whole signature about 3.4× faster (1,157,783 vs 344,174 cycles, median). The SPHINCS+ AVX2 code computes 8 SHA-256 or 4 Keccak (SHAKE) hashes in parallel, filling the vector register with independent tree nodes. OpenTitan (Earl Grey) verifies boot images with ECDSA-P256 plus SLH-DSA-SHA2-128s running on its SHA-256/HMAC block; its separate KMAC block implements SHA-3, SHAKE and cSHAKE with masking.
Measured by PQC Today
On the KV260 FPGA the two schemes had opposite outcomes. ML-DSA-65 with two whole-signature hardware signers: 404.45 → 522.5 signatures/s, only +29.2%. SLH-DSA-SHAKE-128s with a 4-lane Keccak engine: 12.3 s → 69.0 ms per signature. The i.MX 95’s best software path, SLH-DSA-SHA2-128s on its SHA-256 instructions, takes 95 ms.
Measured by PQC Today
SLH-DSA head-to-head, both boards after acceleration
ConfigurationPer signatureSignatures/s
i.MX 95 — CPU, SHAKE-128s1.06 s1.0/s
i.MX 95 — CPU, SHAKE-128s, 6 callers1.59–1.68 s0.93–1.0/s
i.MX 95 — CPU, SHA2-128s (SHA-256 instructions)95 ms9.7–9.9/s
KV260 — CPU only, SHAKE-128s2.12 s0.5/s
KV260 — FPGA engine, SHAKE-128s69 ms14.3–14.5/s
KV260 FPGA vs i.MX 95 on SHAKE-128s: ~15× lower latency, ~14× the throughput. Even against the i.MX 95’s best case (SHA2-128s on its SHA-256 instructions) the KV260 is 1.4× faster per signature and 1.45× in throughput — although the KV260’s own CPU is the slower of the two. Same KV260, same image, FPGA on vs off (2026-09-27): SLH-DSA-SHAKE-128s sign 13.0 vs 0.46/s (28.4×), 192s 30.6×, 256s 26.4×. With the FPGA off, the A53 signs SHAKE-s about 2× slower than the A55 — the whole lead is the fabric. Confirmed in the 2026-09-26 bench (4 workers): SHAKE-128s KV260 12.97/s vs i.MX 95 0.96/s; SHAKE-256s 8.04/s vs 0.61/s.
Measured by PQC Today
Then software caught up. After our engine moved ML-DSA onto hand-written NEON assembly, the KV260’s A53 alone signed ML-DSA-65 at 2,675/s (2026-09-26) — about five times the FPGA-assisted figure on the older engine. The FPGA’s place is the hash-heavy workload: in the same run the KV260 signed SLH-DSA-SHAKE-128s 12.97 times per second on its fabric while the faster i.MX 95 managed 0.96 in software.

FPGA Limits: Area, Clock and the Bus

An FPGA is a grid of small programmable parts: LUTs (6-input truth tables that make up all logic), flip-flops (1-bit registers), DSP slices (hard multipliers) and block RAM. The KV260 has 117,120 LUTs, 1,248 DSPs and 144 block-RAM tiles. Four limits decide what it can do:

  1. Area. Our ML-DSA design used 81% of the LUTs and 95% of the block RAM. Each SLH-DSA Keccak lane cost about 14k LUTs after a redesign, so eight lanes (~130k) would not fit on the whole chip.
  2. Clock. The clock can tick only as fast as the longest logic path allows, and that path grows as the design grows. Our 4-lane engine missed 250 MHz by 10 picoseconds and ships at 240 MHz.
  3. The bus. Data must be copied into the fabric and back, with cache flushes on both sides. Our first Keccak engine cost 1,481 µs per call before doing any work.
  4. Sharing. One engine serves every core. Four threads got the same 14.3/s as one, and threads that fell back to ARM pushed p99 latency to 2.7 s.
Measured by PQC Today
Where the FPGA wins: SLH-DSA-SHAKE signing, today’s engine on every platform
Signatures/sKV260 + FPGAi.MX 95 (CPU)M4 Pro (CPU)KV260 ÷ i.MX 95KV260 FPGA on ÷ off
SLH-DSA-SHAKE-128s13.000.9413.4913.8×28.4× (13.00 vs 0.46)
SLH-DSA-SHAKE-192s8.060.538.3315.2×30.6× (8.05 vs 0.26)
SLH-DSA-SHAKE-256s8.060.629.6113.0×26.4× (8.05 vs 0.31)
Narrow, but powerful — that is what an FPGA is for. On the same KV260, switching the FPGA on raises SLH-DSA-SHAKE-128s signing from 0.46 to 13.0 signatures per second (28×, 4 workers); measured one at a time, a signature takes 2.12 seconds on the A53 and 69 milliseconds on the fabric — the difference between a device that can sign on demand and one that cannot. A 4-core Cortex-A53 board with a 4-lane Keccak engine in its fabric signs the SHAKE “s” sets 13–15× faster than the 6-core Cortex-A55 i.MX 95, and reaches 84–97% of a whole Apple M4 Pro CPU (all 14 cores; about 80% once the M4’s SHA-3 instructions are used, which our engine does not do yet — a measured +19% on this signing). This is one FPGA engine against a laptop processor on one workload: for SHA-2 SLH-DSA and ML-DSA the M4 Pro is 20–25× faster. Both halves are the lesson: pick the one workload that is actually stuck, build hardware for exactly that, and a small, low-power board can stand next to a laptop processor. Switching the FPGA off on the same board drops it 26–31× — the whole lead is the fabric. SHA-2 sets and the fast “f” sets are not routed to the fabric and run on the CPU. Measured by PQC Today, 2026-09-27, hsm 37de2892 on every platform: sign operations per second, 4 concurrent workers, 2 s measurement windows (0.5 s warm-up, at least 20 operations). KV260 on its hashsig FPGA profile (4 Keccak lanes, 240 MHz); i.MX 95 and M4 Pro on CPU.
Measured by PQC Today
Why ML-DSA gained only +29.2%: the two signers could reach about 1,316 signatures/s on their own, but the ARM cores still decode the key, expand the matrix, copy data in and out and run PKCS#11 for every signature. An earlier NTT-only offload was worse: 0.68 ms per attempt vs 0.44 ms on the CPU, of which 0.27 ms was conversion and DMA — making whole signatures slower.

Reprogramming. Unlike an ASIC, the fabric can be rewritten in the field. Our KV260 ships one complete image; we also tested swapping whole images at runtime (“profiles”: ML-DSA signers ⇄ SLH-DSA hash engine, live on the board, with the crypto services restarting). AMD’s modular alternative — partial reconfiguration, which swaps one slot while the rest keeps running — we documented but did not use.

Published, not measured by us
AMD’s Dynamic Function eXchange (DFX, also called partial reconfiguration) splits the fabric into a static “shell” that stays loaded and one or more reconfigurable slots. A new accelerator is loaded into a slot at runtime while the shell — and anything in the other slots — keeps running. AMD publishes a 2-slot DFX reference design for the Kria K26 and the packaging (shell, module firmware, device-tree overlay) that its dfx-mgr loader uses.

Workshop steps 4 and 5 turn these into interactive labs.

Crypto Agility: Accelerating Many Algorithms

A real system never runs one algorithm. It runs ML-KEM and ML-DSA for new sessions, SLH-DSA for long-lived signatures, RSA and ECC for everything not yet migrated, and AES and SHA-2 for the traffic — and the list will change as standards evolve. Acceleration has to follow that list without being rebuilt every time.

Capacity forces choices — our example: two FPGA profiles

An FPGA has a fixed amount of logic and memory. Our ML-DSA image alone uses 81% of the KV260’s logic and 95% of its on-chip memory; the SLH-DSA image uses 55% of the logic. Together they would need roughly 159k LUTs on a 117k-LUT chip — so they cannot be loaded at the same time. Every extra algorithm you want in hardware makes this worse, which is what forces a loading model: swap whole images (profiles), swap slots (partial reconfiguration), or keep only small shared building blocks resident.

ML-DSA profile (2 signers + monitor)94,630 LUT (81%) · 137/144 block RAM
SLH-DSA profile (4 Keccak lanes + monitor)64,600 LUT (55%)
Both at once~159k LUT — more than the whole chip

So the KV260 carries two profiles and loads one at a time: mldsa (two whole-signature ML-DSA-65 signers) or hashsig (a 4-lane Keccak engine for SLH-DSA-SHAKE), both with the same behaviour monitor. One command switches them live — the crypto services stop, the image is checked and loaded, the services restart. Whatever the loaded profile does not cover runs on the ARM cores.

AlgorithmKeccak (SHA-3 / SHAKE)NTT polynomial arithmeticSHA-256 / SHA-512Big-number (Montgomery) arithmeticAESFloating-point FFT
ML-KEM (FIPS 203)
NTT over q = 3,329 (7 layers); SHAKE128 grows the public matrix, SHA3-256/512 and SHAKE256 hash the rest.
●○
ML-DSA (FIPS 204)
NTT over q = 8,380,417 (8 layers); SHAKE128/256 for matrix expansion, sampling and hashing, inside a retry loop.
●○
SLH-DSA-SHAKE (FIPS 205)
SHAKE256 for every hash — about two million Keccak permutations per 128s signature.
●
SLH-DSA-SHA2 (FIPS 205)
SHA-256 at category 1; SHA-512 for some functions at categories 3 and 5.
●
FN-DSA (FIPS 206, draft)
Floating-point FFT sampling for signing (hard to accelerate safely) plus SHAKE256 hashing.
○●
RSA / ECDSA / ECDH
Still in every hybrid deployment during migration.
○●
AES-GCM, SHA-2 hashing
Carries the traffic once a PQC key exchange has run.
○●

● dominant cost · ○ also used. Read down a column: an accelerator for that building block helps every algorithm with a mark in it. Keccak is the most shared block in post-quantum cryptography. Standards: NIST FIPS 202 · NIST FIPS 203 · NIST FIPS 204 · NIST FIPS 205 · FIPS 206

Accelerate shared building blocks, not whole algorithms

One Keccak engine serves ML-KEM, ML-DSA and SLH-DSA-SHAKE; one NTT engine serves both lattice schemes. A new algorithm built from the same blocks gets faster for free.

Our first plan did exactly this — but a per-hash Keccak engine lost to the CPU (1,481 µs fixed cost per call). Building blocks only pay off when each call carries a lot of work.

Offload whole operations only where the win is decisive

A whole-signature engine avoids the round trip and can be many times faster — but it only knows one algorithm and a few parameter sets.

Our SLH-DSA engine signs the SHAKE “s” sets 26–31× faster on the same board, and does nothing for the SHA-2 or “f” sets.

Always keep a software path for every algorithm

Hardware is an optimisation, never the only implementation: if the accelerator is missing, busy or does not know the parameter set, the CPU does the work.

Our engine falls back to ARM when the FPGA is busy or absent, and verification always stays on ARM. Every accelerated result is checked against the software path.

Route by configuration, not by code change

A routing table decides which algorithm goes to which accelerator, so a new standard or a broken parameter set is a configuration change.

Our engine prints its routing at start-up, e.g. “shake=Fpga sha2=Cpu”, and can be told to ignore the hardware entirely (a switch we use for on/off measurements).

Reprogram the hardware when the workload changes

An FPGA can swap whole images (profiles) or single slots (partial reconfiguration). An ASIC cannot — which is why ASIC PQC engines fix their parameter sets up front.

We swap ML-DSA and SLH-DSA profiles live on the KV260; Caliptra’s Adams Bridge ASIC supports ML-DSA-87 and ML-KEM-1024 only.

Prefer general instructions over fixed engines on the CPU

A SHA-3 or SIMD instruction helps every algorithm that uses the primitive, now and later — the most agile acceleration there is.

NEON code gave ML-DSA ~4×; SHA-3 instructions gave SHAKE ~20% — for every SHAKE user at once.

GPU, ASIC and NPU

GPUs win on throughput, not latency. Every launch has a fixed cost, lanes run in lock-step groups, and memory delays are hidden only when thousands of independent operations are in flight — perfect for signing 50,000 certificates, poor for one TLS handshake.

Measured by PQC Today
PQC Today Metal study, ML-DSA-65 key generation, Apple M5 Max 40-core GPU vs its own CPU cores, 2026-06-13. Millions of keys per second; every output bit-exact vs the reference. At a batch of 256 the GPU lost to the CPU; at 65,536 it was 9.3× the all-core CPU. Signing was harder: At a batch of 16,384 the first GPU signing version LOST (0.87× the all-core CPU) because one straggler signature needed 59 rejection rounds and every round re-launched the whole batch. Letting finished signatures exit early fixed it.
Published, not measured by us
cuPQC supports ML-KEM and ML-DSA (all parameter sets) plus SHA-2/SHA-3/SHAKE hashing. NVIDIA lists compute capabilities 8.0, 8.6, 8.7, 8.9 and 9.0 on x86_64 or Arm64 hosts (CUDA 12.8+) — i.e. Ampere data-centre and desktop GPUs (A100, RTX 30-series), the Ampere-based Jetson Orin family (8.7), Ada and Hopper. So it is not only a data-centre library: it runs on an embedded Jetson Orin module too. Its launch figures — 13.5 M ML-KEM-768 key generations/s, 175×/120×/135× over one AMD EPYC 7313P core for keygen/encaps/decaps — were measured on an H100 (Hopper). It ships as closed-source static libraries.

ASICs share the FPGA’s limits — bus transfers, one engine shared by many cores, a fixed area budget — and add their own: the design is frozen at tape-out, takes years and a large up-front investment, and cannot follow a new parameter set or a revised standard. Their reward is the highest clock and lowest power of any option.

Published, not measured by us
Caliptra 2.0’s Adams Bridge accelerator has a hardware NTT, a Keccak (SHAKE128/256) core, samplers and side-channel masking — and supports ML-DSA-87 and ML-KEM-1024 only. Choosing only the highest security level keeps the silicon small — and means ML-DSA-44/65 users get no help from it.
Capacity, size and power: from tiny FPGAs to custom silicon
FPGALogicOn-chip memory / DSPPackagePower, as published
Lattice iCE40 UltraPlus UP5K5,280 LUTs1,024 kb + 120 kb; 8 DSP2.15 × 2.50 mmstatic 75 µA; active 1–10 mA
Efinix Titanium Ti6062,016 logic elements2.6 Mb; 160 DSP3.5 × 3.4 mm (64-ball)not published as one figure
Lattice Certus-NXup to 65k logic cells3.3 Mb; 128 multipliersfrom 6 × 6 mm“up to 4× lower power vs. similar FPGAs” (relative only)
AMD Zynq UltraScale+ K26 (our KV260)117,120 LUTs (256,200 logic cells)5.1 Mb BRAM + 64 × 288 Kb URAM; 1,248 DSPmodule 77 × 60 mmno typical figure; sized with AMD’s power tool (5 V rail up to 4 A)
AMD Virtex UltraScale+ VU19P8.9M logic cells224 Mb; 3,840 DSP—not published
AMD Versal Premium VP19028,460k LUTs (18.5M logic cells)239 Mb BRAM + 619 Mb URAM; 6,864 DSP77.5 × 77.5 mmnot published
ASIC PQC engineNodeAreaPowerPerformance
Caliptra Adams Bridge (ML-DSA-87 + ML-KEM-1024, side-channel protected)5 nm0.1096 mm² (pre-dates a masking refactor — to be re-synthesised)not published600 MHz; ML-DSA phases 26 / 61 / 31 µs
Sapphire (MIT) — configurable lattice crypto processor40 nm0.28 mm²~8 mW at 72 MHz; Kyber-512 encapsulation 5.12 mW, 9.37 µJKyber, Dilithium, NewHope, Frodo, qTESLA on one chip
Saber ASIC (Purdue / KU Leuven / Intel)65 nm0.158 mm²334 µW at 10 MHz, 0.7 Vup to 160 MHz at 1.1 V

Read the two tables together. A complete, side-channel-protected ML-DSA + ML-KEM engine fits in about a tenth of a square millimetre of 5 nm silicon, and lattice ASICs run in milliwatts or less. The same engines in an FPGA need a mid-sized part like our K26 — far more silicon and power — because programmable logic pays for its flexibility:

Published, not measured by us
The classic measurement of the gap (Kuon & Rose, FPGA 2006): the same circuit in an FPGA needs on average about 40× more silicon area (about 21× when the FPGA’s hard multipliers and memories are used), runs 3–4× slower, and uses about 12× more dynamic power than in a custom ASIC. FPGA vendors do not publish one power figure per chip — power depends on the design, its clock and how often its signals switch — so compare designs, not part numbers. Kuon & Rose, “Measuring the gap between FPGAs and ASICs” (FPGA 2006)
Buying it instead of building it: commercial PQC hardware IP
Vendor / IPAlgorithmsSize (as published)PerformanceSide-channel claim
Rambus
QSE-IP-86
ML-KEM, ML-DSA, SLH-DSA (+ SHA-3/SHAKE)not publishedML-KEM-1024: 7,100 decaps / 13,500 encaps per s; ML-DSA-87: up to 1,400 signs per s (1 GHz)DPA protection sold as a separate variant (QSE-IP-86-DPA); CAVP
Synopsys
Agile PQC Public Key Accelerator
ML-KEM, ML-DSA, SLH-DSA, XMSS, LMSnot publishednot publishedconfigurable DPA / timing / fault countermeasures; primitives in hardware, algorithms in firmware
PQSecure
PQC hardware IP
ML-KEM, ML-DSA, SLH-DSAtiny → high-performance profilesnot publishedfirst-order masking + shuffling, constant-time datapaths; CAVP
Secure-IC
Securyzr PQC (lattice)
ML-KEM, ML-DSAnot publishednot publishedclaims SPA / DPA / DEMA / CPA / CEMA protection; hybrid hardware–software
CAST (engineered by KiviCore)
KiviPQC-KEM
ML-KEMFast: 83k gate-equivalents (7 nm, 600 MHz); Tiny: 33k (100 MHz)not publishedtiming only
Xiphera
XIP6110B (ML-KEM) / XIP6220B (ML-DSA)
ML-KEM, ML-DSAML-KEM under 10k LUTs“thousands of operations per second”constant-time (timing) only
BERTEN
MLKE-B135
ML-KEM9.0k LUT (7-series), 8.6k LUT (Versal)datasheet table at 300 MHztiming and simple power analysis (SPA); no DPA claim
IP Cores Inc.
PQC1
ML-KEM, ML-DSA~110k gatesML-DSA-44 sign 15,000/s; ML-KEM-512 keygen 95,000/snone stated
PQShield
PQPerform-Lattice
ML-KEM, ML-DSAnot published“high-throughput hardware accelerator” (NIST CAVP entry)not stated on the CAVP entry

Taken from the hardware entries of our Migrate product catalog and checked against each vendor’s own page. Note how differently “side-channel protected” is used: some cores only promise constant time (no timing leak), some add simple-power-analysis resistance, and only a few claim masking against differential power analysis — often sold as a separate, larger variant.

In plain English: constant-time vs masked

Constant-time means the circuit takes exactly the same time and follows the same steps whatever the secret key is, so a stopwatch learns nothing. It is cheap. Masking goes further: every secret value is split into random pieces (“shares”) that are processed separately and only recombine at the end, so the power drawn or radio noise emitted at any moment is unrelated to the key. It defeats power and electromagnetic analysis — but every share has to be carried, refreshed with fresh randomness and recombined carefully, which costs extra area, randomness and time, and the cost grows with the number of shares.

What side-channel protection costs
ImplementationProtectionCost vs unprotectedPlatform
Caliptra Adams Bridge, ML-KEM-1024 / ML-DSA-872-share masking + shuffling (two parallel NTT engines)Zero extra cycles (ML-KEM decapsulation 11,054 vs 11,056 cycles) — “the only cost of enabling masking is area”ASIC
OpenTitan KMAC (Keccak)First-order domain-oriented maskingKeccak round logic “more than twice” the area; SHA3-224 throughput 3.43 → 1.26 bytes/cycleASIC
Kyber-512 (Kamucheka et al.)Hiding, then hiding + maskingHiding: 1.83× cycles, 1.6× resources vs baseline; adding masking: a further 1.08× cycles, 1.06× resourcesFPGA (Virtex-7)
ML-DSA (Raj et al.)First-order masking1.127× LUTs, 1.2× flip-flops — but 378× execution timeFPGA (Kintex-7 / Zynq)
Kyber768 (Bos et al.)First, second, third orderFirst order 3.5× the cycles of the optimised unprotected code (3.1M vs 0.88M); second order 44.3M cycles; third order 115.5MArm Cortex-M4F
Dilithium3 signing (Azouaoui et al.)2 shares (first order)4.3× slower (randomised signing), 7.32× slower (deterministic)Arm Cortex-M4
Dilithium variant (Migliore et al.)Order 1 / 2 / 3About 5.6× / 11.6× / 28× slower than unmaskedgeneral-purpose processor
  • In software, first-order masking costs roughly 3–7× the time for Kyber and Dilithium, and the cost climbs steeply with each extra order (Kyber768: 3.5× at first order, over 100× at third).
  • In hardware the bill can be paid in area instead of time: Adams Bridge runs two NTT engines in parallel so masked decapsulation takes no extra cycles — while masked Keccak in OpenTitan more than doubles the round logic.
  • Masking a design that was not built for it can be disastrous: one first-order ML-DSA FPGA design added only 13–20% area but ran 378× slower.
  • Always ask which attacks a “protected” product covers: constant-time (timing), SPA, or DPA/masking — and at what order.

Published, not measured by us. Our own KV260 engines were built for speed and have not yet been evaluated against power or electromagnetic analysis — a planned later phase.

Can an ASIC be reprogrammed? Partly — it depends where the “recipe” lives

Many crypto engines are microcoded: fixed hardware (multipliers, Keccak, memories) follows a stored list of steps, like a kitchen following a recipe card. If that card is in writable memory, a signed update can change the step order, the parameter sets or fix a bug — but never add hardware the chip lacks. If the card is burned into ROM, nothing changes after manufacture. Adams Bridge is microcoded, and its microcode is in ROM.

ApproachIn plain EnglishCan change after manufactureCannot change
Hard-wired or ROM microcodeThe step-by-step recipe exists, but it is burned into read-only memory.nothing in the algorithmthe recipe, the parameter sets
Programmable crypto co-processor (e.g. OpenTitan OTBN)A small processor with its own instruction set; the host loads a new program.the software — RSA, ECC, even a PQC NTTits instructions and memory size — fast PQC needed new instructions and 8× more memory (a new chip)
Fixed building blocks, algorithm in firmware (e.g. Sapphire)Keccak, NTT and modular arithmetic in hardware; each scheme’s steps in updatable firmware.which lattice scheme runs, its parametersthe building blocks themselves
Embedded FPGA (eFPGA) inside an ASIC
Flex Logix press release, 7 Nov 2022: “all forthcoming updates and algorithm modifications can be supported by simply reprogramming eFPGA”
A patch of FPGA fabric built into the chip, reprogrammed with a new bitstream.whole algorithm circuitsthe size of the patch and the rest of the chip

The practical middle path for crypto agility in silicon: put the stable building blocks (Keccak, NTT, modular multiply) in hardware and keep each scheme’s steps in updatable firmware — most of the ASIC’s efficiency, much of the FPGA’s flexibility.

NPUs are arrays of 8-bit multiply-accumulate units fed a pre-compiled neural-network graph, designed to tolerate rounding. PQC needs exact arithmetic on 23-bit (ML-DSA) and 12-bit (ML-KEM) numbers, and Keccak is bitwise logic with no multiplication at all.

Our own evaluation (desk analysis, not measured)
We evaluated the i.MX 95’s eIQ Neutron NPU for PQC and ruled it out on paper (not prototyped): it is reached only through compiled model graphs, its int8 quantisation is lossy, every NTT stage would need a round trip back to the CPU, and the dominant cost in our traces is Keccak, which no multiply-accumulate array can express. We use the NPU to monitor the appliance’s crypto behaviour instead.
Published, not measured by us
The i.MX 95 integrates an eIQ Neutron N3-1024S NPU rated at 2.0 TOPS for 8-bit neural-network inference. TensorFHE recasts the NTT as matrix multiplications and splits each 32-bit value into four 8-bit pieces to use INT8 tensor cores (NVIDIA A100) — relying on the GPU’s general-purpose cores to recombine them.

Lessons: Owning an Instruction Is Not Using It

Measured by PQC Today
  • AES ran in software on every board until a compile flag was set. AES-128-CBC encrypt at 16 KiB, A-B-A-B, 2026-09-26. Same crate, with vs without the build flag that enables the ARMv8 AES instructions. Before the fix, AES ran in software on every board (peak 75 MB/s on the MX95). Speed-up at 16 KiB: M4 Pro 10.8–12.9×, i.MX 95 7.8–9.1×, KV260 7.4–10.0×.
  • A library chose the wrong kernel for our core. Cortex-A55, AWS-LC, 2026-09-16. AWS-LC picked a NEON/Karatsuba Montgomery kernel tuned for Neoverse cores; on the A55 it LOST to the plain scalar kernel. Forcing the scalar kernel: RSA-2048 +29%, RSA-3072 +62%, RSA-4096 +22%.
  • An instruction can be present, unused — and then only a modest win. Build audit, 2026-09-27: the M4 Pro reports FEAT_SHA3, but our engine builds the Rust `keccak` crate (used by SLH-DSA) without its `asm` feature, so SLH-DSA’s SHAKE runs the portable software permutation even there. (AWS-LC’s own Keccak, used by ML-DSA/ML-KEM, already uses the instructions.) Measured A/B on the M4 Pro, same engine with the feature switched on: SHAKE 1.17–1.20×, SLH-DSA-SHAKE-128s signing 1.19× — a real but modest gain, far below AES’s 7–13×, because Apple’s wide cores already run the software permutation well.
  • An accelerator slower than the CPU is not an accelerator. Our first Keccak FPGA engine needed 58.7 µs per hash where the A53 needed about 12 µs — we removed it, and dropped SHA-2 from the fabric for the same reason.

The practical checklist: find where the time actually goes; check which instructions the core really has and that the build uses them; prefer acceleration inside the CPU for small, latency-sensitive operations; offload only whole operations; and measure end to end, under concurrency, against the best CPU path — not the reference code.

Related modules

Seven acceleration models and five building blocks in plain English, then our own measurements and FPGA labs.

Check your understanding

8 questions on PQC Hardware Acceleration, each with its answer and the reason.

Take the quiz

Next step

Produce the artifact: Infrastructure Modernization Planner

This module belongs to phase 6 (Infrastructure & Performance); Infrastructure Modernization Planner produces a deliverable of that phase in the Command Center.

Learning module content can be inaccurate. Please double-check its information. Report inaccuracies in PQC Today GitHub Discussions.