PQC Hardware Acceleration
How post-quantum signatures are accelerated on CPUs, GPUs, FPGAs, ASICs and NPUs — backed by our own measurements on three generations of Arm cores and an FPGA.
Why this matters: ML-DSA and SLH-DSA spend their time in completely different places, so the same accelerator can give +29% on one and more than 40× on the other. Choosing hardware, parameter sets and hash variants for a device — or judging a vendor’s acceleration claim — needs an honest picture of where the time goes and what moving it costs.
Start here: Click through the seven acceleration models: each opens with an everyday analogy, then an animated diagram of the real mechanism, then how well it fits ML-DSA, SLH-DSA, single operations and big batches.
Where the Time Goes
“Accelerating PQC” is not one problem. Each algorithm spends its time in a different place, and an accelerator only helps if it speeds up that place — and if moving the work there costs less than it saves.
Polynomial arithmetic mod 8,380,417 via the NTT, plus a lot of SHAKE to grow the public matrix and sample secrets, inside a retry loop that runs about 4–5 times per signature on average. Already fast: milliseconds on a small core.
Almost nothing but hashing — about two million Keccak permutations for one SHAKE-128s signature — built into large Merkle trees. Slow: seconds on a small core for the compact “s” parameter sets.
Big-integer multiplication with Montgomery reduction. Still in every hybrid deployment during the migration, so still worth accelerating.
The Five Building Blocks
Hardware rarely accelerates “an algorithm”. It accelerates the primitives underneath:
- Montgomery reduction — computes “mod N” with multiplications and a shift instead of a slow division (Montgomery, 1985; FIPS 204 Appendix A). Inside RSA, ECC and every ML-DSA coefficient multiply.
- NTT (Number-Theoretic Transform) — turns polynomial multiplication from 65,536 products into about 3,300 for 256 coefficients. ML-DSA runs a full 8-layer NTT; ML-KEM’s modulus 3,329 only allows 7 layers.
- Keccak-f[1600] — a 24-round shuffle of a 1,600-bit state using only XOR, AND, NOT and rotate (FIPS 202). Constant-time by nature and cheap in hardware.
- SHA-3 — Keccak in a “sponge”: absorb the message into the 1,088-bit rate, keep a 512-bit capacity sealed, output a fixed 256 bits (SHA3-256).
- SHAKE — the same sponge with an output tap you can leave open. SHAKE128 (capacity 256) grows ML-DSA’s and ML-KEM’s matrices; SHAKE256 is every hash in SLH-DSA-SHAKE.
Workshop step 2 explains each one in plain English, with a worked Montgomery example you can check by hand and an animated NTT butterfly network.
Seven Models of Acceleration
From closest-to-the-CPU to furthest away: scalar code; SIMD vector units (Arm NEON/SVE, x86 AVX2/AVX-512); dedicated crypto instructions (AES, SHA-256, SHA-512, SHA-3); GPUs; FPGAs; ASICs and secure elements; and NPUs. The further from the CPU, the more raw speed is available — and the more you pay to move data there and back.
| Feature | Arm | Intel / AMD x86 | Where PQC uses it |
|---|---|---|---|
Wide SIMD Same operation on many numbers at once | NEON 128-bit (all Armv8); SVE/SME on newer cores | AVX2 256-bit; AVX-512 512-bit | ML-DSA / ML-KEM NTT and polynomial arithmetic; several SLH-DSA hash chains side by side |
AES One AES round per instruction | FEAT_AES (optional from Armv8.0) | AES-NI; VAES (AES on AVX/AVX-512 registers) | Symmetric encryption after a PQC key exchange; AES-based schemes |
SHA-256 SHA-256 rounds in hardware | FEAT_SHA256 (optional from Armv8.0) — on A53, A55, Apple M-series | SHA-NI (SHA-1 and SHA-256 only) | SLH-DSA-SHA2 (all of category 1; F and PRF at every category), LMS/HSS |
SHA-512 SHA-512 rounds in hardware | FEAT_SHA512 (Armv8.2 extension) — NOT on A53/A55; on Apple M4 | VSHA512* — only Arrow Lake-S / Lunar Lake and later | SLH-DSA-SHA2 categories 3 and 5 (H_msg, PRF_msg, H, T) |
SHA-3 / Keccak EOR3, RAX1, XAR, BCAX fuse Keccak’s θ, ρ and χ steps | FEAT_SHA3 (Armv8.2 extension) — NOT on A53/A55; on Apple M4 | None — no SHA-3 or Keccak instruction exists; SIMD runs 4–8 Keccak states in parallel | Every SHAKE call: ML-KEM, ML-DSA sampling, SLH-DSA-SHAKE |
Big-integer multiply Wide multiply-accumulate for bignum math | Scalar MUL/UMULH; NEON UMULL | MULX/ADX; AVX-512 IFMA (eight 52-bit multiplies per instruction) | Classical RSA/ECC during hybrid migration — not needed by lattice or hash PQC |
Arm Across Three Generations — Our Boards
We run the same Rust PKCS#11 engine on three Arm platforms: an Apple M4 Pro laptop (Armv9.2-A, with SHA-3, SHA-512 and SME2), an NXP i.MX 95 appliance (6× Cortex-A55, Armv8.2-A) and an AMD Kria KV260 appliance (4× Cortex-A53, Armv8.0-A, plus FPGA fabric). The A55 and A53 both have AES and SHA-256 instructions and neither has SHA-3 or SHA-512.
- ML-DSA-65 signing: M4 Pro 72,776/s · i.MX 95 5,800/s · KV260 3,268/s (same engine, 4 workers, 2026-09-27).
- The instruction set decides which variant wins. On the A55 and A53, SLH-DSA-SHA2-128s signs about 11× faster than SLH-DSA-SHAKE-128s on today’s engine (hardware SHA-256 plus SHA-2-specific software work vs software Keccak). Digest throughput at 16 KiB through the PKCS#11 engine on the boards, 2026-09-26. Cortex-A5x has SHA-256 instructions but not SHA-512 ones, so the usual software ordering inverts. (SHA-256 396–470 MB/s vs SHA-512 136–143 MB/s.)
- Same instructions, newer core: one worker signing ML-DSA-44, the A55 does 255.6/s and the A53 186/s (+37%). A newer microarchitecture and a higher clock lift everything; a new instruction lifts only what uses it.
Newer is not simply “more of the same”. The M4 Pro has 10 big cores and 4 small efficiency cores. The M5 Max drops the efficiency cores: it has 6 top-tier “super cores” and 12 new mid-tier performance cores. A single signature runs on one core, so its speed depends on which kind of core it lands on; a batch spread over all cores depends on how many cores of each kind there are.
Workshop step 3 lets you compare every algorithm across the three platforms.
What We Built on Arm — AES, ML-DSA, SLH-DSA, AWS-LC
Every acceleration below is real work in our engine or appliance images, with the measured before → after. The coloured tag says which acceleration model it is — and notice how many of the biggest wins are plain software engineering, not new hardware.
AES
| What we did | Model | Measured | Status |
|---|---|---|---|
Switch on the Armv8 AES and PMULL instructions Build flags for the Rust AES crates (they were silently off) | Crypto instr. | AES-128-CBC 16 KiB: KV260 7.4×, i.MX 95 7.8×, M4 Pro 10.8×; AES-256-CBC up to 12.9×; GCM 3–6× All three platforms, A-B-A-B, 2026-09-26 | Shipped cacp #36/#38, sandbox #83/#84 |
Move to the newer AES crates that use the hardware by default Runtime CPU detection instead of a build flag | Crypto instr. | CBC unchanged (0.98–0.99×), GCM 1.15–1.19× — same speed, no flag to forget M4 Pro, 2026-09-27 | Merged, next image hsm #283 |
Process GCM a whole block at a time instead of byte by byte Whole-block counter keystream and GHASH | Software | AES-128-GCM 16 KiB: KV260 board 6.1–7.6× (5,455 → 35,603/s), i.MX 95 Pro board 6.5–7.0× (11,743 → 76,682/s), M4 Pro 5.6–6.5× KV260 + i.MX 95 Pro on-board A/B, M4 Pro, 2026-09-27 | Shipped hsm #291 |
Stop cores queuing on shared engine state Per-session state sharded 64 ways; AES key schedule cached | Software | AES-128-CBC 64 B: KV260 board 3.2–3.3× (46k → 150k/s), i.MX 95 Pro board 4.1–4.4× (up to 5.0× at 6 workers, 87k → 435k/s). The old engine did not scale past 1 worker; the new one scales near-linearly. Single worker 5–18% lower (known trade-off, follow-up planned) KV260 + i.MX 95 Pro on-board A/B, M4 Pro, 2026-09-27 | Shipped hsm #298 |
Compile the appliance engine at full optimisation opt-level 3 instead of size-optimised | Software | AES-CBC 16 KiB 1.40–1.53×, 64 B 1.23× M4 Pro, 2026-09-27 | Shipped cacp recipes |
ML-DSA
| What we did | Model | Measured | Status |
|---|---|---|---|
Run ML-DSA and ML-KEM on AWS-LC’s hand-written NEON assembly mldsa-native / mlkem-native AArch64 code inside AWS-LC | SIMD | On/off on the M4 Pro (same engine, AWS-LC path disabled vs enabled): ML-DSA-65 sign 18.2k → 72.7k/s (4.0×), verify 4.0×, keygen 7.7×; ML-DSA-87 sign 3.4×; ML-KEM-768 decapsulate 1.5×. Earlier container test with SHA-3 masked (A53/A55-like): ML-DSA-65 sign 3.8×, verify 3.3× M4 Pro A-B-A-B, 2026-09-27; container, 2026-09-24 | Shipped hsm #254 (98b67e63) |
Decode each private key and expand its matrix once, not on every signature Per-key expanded signing context, cached | Software | Host CPU per FPGA-assisted signature 0.252 → 0.041 of a CPU-only signature arm64 container model | Shipped hsm #254 (e2246faf) |
Offload whole ML-DSA-65 signatures to two FPGA signers DMA + one system call per signature + interrupt | FPGA | 404 → 523 sign/s at 4 workers (+29%); later 515 → 975 at 8 threads with the host-path work KV260 board | Shipped hsm #254, cacp images |
SLH-DSA
| What we did | Model | Measured | Status |
|---|---|---|---|
Use the Armv8 SHA-256/SHA-512 instructions for the SHA-2 variants Rust sha2 crate “asm” feature, Arm only | Crypto instr. | SHA2-128s sign 568 → 102 ms; SHAKE control unchanged (49.0 → 47.4 ms) arm64 container on a Mac | Shipped hsm 490988a3 |
Hash the public seed block once per signature, not once per hash call SHA-256/512 midstate cache | Software | SHA2-128s sign 98.0 → 79.6 ms, SHA2-192s 229.9 → 153.6 ms arm64 container on a Mac | Shipped hsm f2e195fa |
Build the independent subtrees of one signature on all cores Multi-threaded signing with one shared core budget | Software | SHA2-128s 76.9 → 19.8 ms (4 threads); SHAKE-128s 735.8 → 195.0 ms arm64 container on a Mac | Shipped hsm 42291340 |
Use the Armv8.2 SHA-3 instructions for SHAKE (M4-class cores) Rust keccak crate “asm” feature — tested as an A/B build, not yet in the engine | Crypto instr. | M4 Pro: SHAKE 1.17–1.20×, SLH-DSA-SHAKE-128s sign 14.5 → 17.5/s (1.19×), SHA3-256 digest 1.14–1.15×; SHA-2 controls 1.00×. A55/A53 lack the instructions, so no effect there M4 Pro A-B-A-B, 2026-09-27 (M5 Max partial run agrees: 1.17–1.22×) | Planned A/B build of hsm 476f97d1 |
Offload whole SLH-DSA-SHAKE signatures to a 4-lane Keccak engine 4 Keccak lanes at 240 MHz in the KV260 fabric | FPGA | Same board, FPGA on vs off: SLH-DSA-SHAKE-128s sign 0.46 → 13.0/s (28.4×), 192s 30.6×, 256s 26.4×, keygen 14–15×; SHAKE-f and SHA-2 sets unchanged (not routed to the fabric) KV260 board, 2026-09-27 (re-proves the 09-25 result: 2.12 s → 69 ms) | Shipped hsm #254, cacp #34 |
AWS-LC (RSA)
| What we did | Model | Measured | Status |
|---|---|---|---|
Move RSA (and NIST-curve ECDH) onto AWS-LC aws-lc-rs: optimised, constant-time big-number code | Software | RSA-2048 sign 49.4 → 68.6/s (1 thread), 235.6 → 326.8/s (6 threads) i.MX 95 board | Shipped hsm #239 |
Stop re-parsing the RSA key on every operation Parsed-key cache | Software | RSA-2048 sign 80.8 → 105.4/s (1 thread), 356 → 481/s (6 threads) i.MX 95 board, pinned core | Shipped hsm #240 |
Answer the decrypt “how big is the output?” query without decrypting Return the size from the key instead of doing a private-key operation | Software | RSA-OAEP-2048 decrypt: KV260 board 167 → 347/s (2.05–2.11×), i.MX 95 Pro board 321 → 642/s (2.00×), M4 Pro 4,144 → 8,295/s KV260 + i.MX 95 Pro on-board A/B, M4 Pro, 2026-09-27 | Shipped hsm #289 |
Use the scalar Montgomery kernel on Cortex-A5x cores Patch AWS-LC’s kernel choice (it picks a server-tuned NEON kernel that is slower on A5x) | Software | On-board, stock vs patched engine from one commit: RSA-OAEP decrypt 2048/3072/4096 — i.MX 95 (A55) 1.27× / 1.60× / 1.21–1.26×, KV260 (A53) 1.28–1.39× / 1.64–1.69× / 1.22–1.25×. RSA-PSS sign unaffected (pure Rust) i.MX 95 + KV260 boards, 2026-09-27 | In review AWS-LC dispatch patch (hsm PR in review) |
Container and M4 Pro figures isolate one change at a time; the boards only have bundled before/after figures (e.g. KV260 SLH-DSA-SHA2-128s 9.4 s → 191 ms, SHAKE-128s 12.3 s → 2.12 s, all CPU changes together). Items marked “merged” reach the boards with the next image build.
ML-DSA vs SLH-DSA: Different Bottlenecks
256 coefficients all get the same treatment, which is exactly what SIMD does best, inside the CPU with no transfer cost. Offload is hard: the retry loop, key decoding and SHAKE sampling all bounce between stages.
Millions of small, independent hash calls. Hash instructions help directly (if the variant matches the core’s hash), SIMD can run several tree paths side by side, and a hash engine that takes a whole signature per call can shine.
| Configuration | Per signature | Signatures/s |
|---|---|---|
| i.MX 95 — CPU, SHAKE-128s | 1.06 s | 1.0/s |
| i.MX 95 — CPU, SHAKE-128s, 6 callers | 1.59–1.68 s | 0.93–1.0/s |
| i.MX 95 — CPU, SHA2-128s (SHA-256 instructions) | 95 ms | 9.7–9.9/s |
| KV260 — CPU only, SHAKE-128s | 2.12 s | 0.5/s |
| KV260 — FPGA engine, SHAKE-128s | 69 ms | 14.3–14.5/s |
FPGA Limits: Area, Clock and the Bus
An FPGA is a grid of small programmable parts: LUTs (6-input truth tables that make up all logic), flip-flops (1-bit registers), DSP slices (hard multipliers) and block RAM. The KV260 has 117,120 LUTs, 1,248 DSPs and 144 block-RAM tiles. Four limits decide what it can do:
- Area. Our ML-DSA design used 81% of the LUTs and 95% of the block RAM. Each SLH-DSA Keccak lane cost about 14k LUTs after a redesign, so eight lanes (~130k) would not fit on the whole chip.
- Clock. The clock can tick only as fast as the longest logic path allows, and that path grows as the design grows. Our 4-lane engine missed 250 MHz by 10 picoseconds and ships at 240 MHz.
- The bus. Data must be copied into the fabric and back, with cache flushes on both sides. Our first Keccak engine cost 1,481 µs per call before doing any work.
- Sharing. One engine serves every core. Four threads got the same 14.3/s as one, and threads that fell back to ARM pushed p99 latency to 2.7 s.
| Signatures/s | KV260 + FPGA | i.MX 95 (CPU) | M4 Pro (CPU) | KV260 ÷ i.MX 95 | KV260 FPGA on ÷ off |
|---|---|---|---|---|---|
| SLH-DSA-SHAKE-128s | 13.00 | 0.94 | 13.49 | 13.8× | 28.4× (13.00 vs 0.46) |
| SLH-DSA-SHAKE-192s | 8.06 | 0.53 | 8.33 | 15.2× | 30.6× (8.05 vs 0.26) |
| SLH-DSA-SHAKE-256s | 8.06 | 0.62 | 9.61 | 13.0× | 26.4× (8.05 vs 0.31) |
Reprogramming. Unlike an ASIC, the fabric can be rewritten in the field. Our KV260 ships one complete image; we also tested swapping whole images at runtime (“profiles”: ML-DSA signers ⇄ SLH-DSA hash engine, live on the board, with the crypto services restarting). AMD’s modular alternative — partial reconfiguration, which swaps one slot while the rest keeps running — we documented but did not use.
Workshop steps 4 and 5 turn these into interactive labs.
Crypto Agility: Accelerating Many Algorithms
A real system never runs one algorithm. It runs ML-KEM and ML-DSA for new sessions, SLH-DSA for long-lived signatures, RSA and ECC for everything not yet migrated, and AES and SHA-2 for the traffic — and the list will change as standards evolve. Acceleration has to follow that list without being rebuilt every time.
An FPGA has a fixed amount of logic and memory. Our ML-DSA image alone uses 81% of the KV260’s logic and 95% of its on-chip memory; the SLH-DSA image uses 55% of the logic. Together they would need roughly 159k LUTs on a 117k-LUT chip — so they cannot be loaded at the same time. Every extra algorithm you want in hardware makes this worse, which is what forces a loading model: swap whole images (profiles), swap slots (partial reconfiguration), or keep only small shared building blocks resident.
So the KV260 carries two profiles and loads one at a time: mldsa (two whole-signature ML-DSA-65 signers) or hashsig (a 4-lane Keccak engine for SLH-DSA-SHAKE), both with the same behaviour monitor. One command switches them live — the crypto services stop, the image is checked and loaded, the services restart. Whatever the loaded profile does not cover runs on the ARM cores.
| Algorithm | Keccak (SHA-3 / SHAKE) | NTT polynomial arithmetic | SHA-256 / SHA-512 | Big-number (Montgomery) arithmetic | AES | Floating-point FFT |
|---|---|---|---|---|---|---|
ML-KEM (FIPS 203) NTT over q = 3,329 (7 layers); SHAKE128 grows the public matrix, SHA3-256/512 and SHAKE256 hash the rest. | ● | ○ | ||||
ML-DSA (FIPS 204) NTT over q = 8,380,417 (8 layers); SHAKE128/256 for matrix expansion, sampling and hashing, inside a retry loop. | ● | ○ | ||||
SLH-DSA-SHAKE (FIPS 205) SHAKE256 for every hash — about two million Keccak permutations per 128s signature. | ● | |||||
SLH-DSA-SHA2 (FIPS 205) SHA-256 at category 1; SHA-512 for some functions at categories 3 and 5. | ● | |||||
FN-DSA (FIPS 206, draft) Floating-point FFT sampling for signing (hard to accelerate safely) plus SHAKE256 hashing. | ○ | ● | ||||
RSA / ECDSA / ECDH Still in every hybrid deployment during migration. | ○ | ● | ||||
AES-GCM, SHA-2 hashing Carries the traffic once a PQC key exchange has run. | ○ | ● |
● dominant cost · ○ also used. Read down a column: an accelerator for that building block helps every algorithm with a mark in it. Keccak is the most shared block in post-quantum cryptography. Standards: NIST FIPS 202 · NIST FIPS 203 · NIST FIPS 204 · NIST FIPS 205 · FIPS 206
One Keccak engine serves ML-KEM, ML-DSA and SLH-DSA-SHAKE; one NTT engine serves both lattice schemes. A new algorithm built from the same blocks gets faster for free.
Our first plan did exactly this — but a per-hash Keccak engine lost to the CPU (1,481 µs fixed cost per call). Building blocks only pay off when each call carries a lot of work.
A whole-signature engine avoids the round trip and can be many times faster — but it only knows one algorithm and a few parameter sets.
Our SLH-DSA engine signs the SHAKE “s” sets 26–31× faster on the same board, and does nothing for the SHA-2 or “f” sets.
Hardware is an optimisation, never the only implementation: if the accelerator is missing, busy or does not know the parameter set, the CPU does the work.
Our engine falls back to ARM when the FPGA is busy or absent, and verification always stays on ARM. Every accelerated result is checked against the software path.
A routing table decides which algorithm goes to which accelerator, so a new standard or a broken parameter set is a configuration change.
Our engine prints its routing at start-up, e.g. “shake=Fpga sha2=Cpu”, and can be told to ignore the hardware entirely (a switch we use for on/off measurements).
An FPGA can swap whole images (profiles) or single slots (partial reconfiguration). An ASIC cannot — which is why ASIC PQC engines fix their parameter sets up front.
We swap ML-DSA and SLH-DSA profiles live on the KV260; Caliptra’s Adams Bridge ASIC supports ML-DSA-87 and ML-KEM-1024 only.
A SHA-3 or SIMD instruction helps every algorithm that uses the primitive, now and later — the most agile acceleration there is.
NEON code gave ML-DSA ~4×; SHA-3 instructions gave SHAKE ~20% — for every SHAKE user at once.
GPU, ASIC and NPU
GPUs win on throughput, not latency. Every launch has a fixed cost, lanes run in lock-step groups, and memory delays are hidden only when thousands of independent operations are in flight — perfect for signing 50,000 certificates, poor for one TLS handshake.
ASICs share the FPGA’s limits — bus transfers, one engine shared by many cores, a fixed area budget — and add their own: the design is frozen at tape-out, takes years and a large up-front investment, and cannot follow a new parameter set or a revised standard. Their reward is the highest clock and lowest power of any option.
| FPGA | Logic | On-chip memory / DSP | Package | Power, as published |
|---|---|---|---|---|
| Lattice iCE40 UltraPlus UP5K | 5,280 LUTs | 1,024 kb + 120 kb; 8 DSP | 2.15 × 2.50 mm | static 75 µA; active 1–10 mA |
| Efinix Titanium Ti60 | 62,016 logic elements | 2.6 Mb; 160 DSP | 3.5 × 3.4 mm (64-ball) | not published as one figure |
| Lattice Certus-NX | up to 65k logic cells | 3.3 Mb; 128 multipliers | from 6 × 6 mm | “up to 4× lower power vs. similar FPGAs” (relative only) |
| AMD Zynq UltraScale+ K26 (our KV260) | 117,120 LUTs (256,200 logic cells) | 5.1 Mb BRAM + 64 × 288 Kb URAM; 1,248 DSP | module 77 × 60 mm | no typical figure; sized with AMD’s power tool (5 V rail up to 4 A) |
| AMD Virtex UltraScale+ VU19P | 8.9M logic cells | 224 Mb; 3,840 DSP | — | not published |
| AMD Versal Premium VP1902 | 8,460k LUTs (18.5M logic cells) | 239 Mb BRAM + 619 Mb URAM; 6,864 DSP | 77.5 × 77.5 mm | not published |
| ASIC PQC engine | Node | Area | Power | Performance |
|---|---|---|---|---|
| Caliptra Adams Bridge (ML-DSA-87 + ML-KEM-1024, side-channel protected) | 5 nm | 0.1096 mm² (pre-dates a masking refactor — to be re-synthesised) | not published | 600 MHz; ML-DSA phases 26 / 61 / 31 µs |
| Sapphire (MIT) — configurable lattice crypto processor | 40 nm | 0.28 mm² | ~8 mW at 72 MHz; Kyber-512 encapsulation 5.12 mW, 9.37 µJ | Kyber, Dilithium, NewHope, Frodo, qTESLA on one chip |
| Saber ASIC (Purdue / KU Leuven / Intel) | 65 nm | 0.158 mm² | 334 µW at 10 MHz, 0.7 V | up to 160 MHz at 1.1 V |
Read the two tables together. A complete, side-channel-protected ML-DSA + ML-KEM engine fits in about a tenth of a square millimetre of 5 nm silicon, and lattice ASICs run in milliwatts or less. The same engines in an FPGA need a mid-sized part like our K26 — far more silicon and power — because programmable logic pays for its flexibility:
| Vendor / IP | Algorithms | Size (as published) | Performance | Side-channel claim |
|---|---|---|---|---|
| Rambus QSE-IP-86 | ML-KEM, ML-DSA, SLH-DSA (+ SHA-3/SHAKE) | not published | ML-KEM-1024: 7,100 decaps / 13,500 encaps per s; ML-DSA-87: up to 1,400 signs per s (1 GHz) | DPA protection sold as a separate variant (QSE-IP-86-DPA); CAVP |
| Synopsys Agile PQC Public Key Accelerator | ML-KEM, ML-DSA, SLH-DSA, XMSS, LMS | not published | not published | configurable DPA / timing / fault countermeasures; primitives in hardware, algorithms in firmware |
| PQSecure PQC hardware IP | ML-KEM, ML-DSA, SLH-DSA | tiny → high-performance profiles | not published | first-order masking + shuffling, constant-time datapaths; CAVP |
| Secure-IC Securyzr PQC (lattice) | ML-KEM, ML-DSA | not published | not published | claims SPA / DPA / DEMA / CPA / CEMA protection; hybrid hardware–software |
| CAST (engineered by KiviCore) KiviPQC-KEM | ML-KEM | Fast: 83k gate-equivalents (7 nm, 600 MHz); Tiny: 33k (100 MHz) | not published | timing only |
| Xiphera XIP6110B (ML-KEM) / XIP6220B (ML-DSA) | ML-KEM, ML-DSA | ML-KEM under 10k LUTs | “thousands of operations per second” | constant-time (timing) only |
| BERTEN MLKE-B135 | ML-KEM | 9.0k LUT (7-series), 8.6k LUT (Versal) | datasheet table at 300 MHz | timing and simple power analysis (SPA); no DPA claim |
| IP Cores Inc. PQC1 | ML-KEM, ML-DSA | ~110k gates | ML-DSA-44 sign 15,000/s; ML-KEM-512 keygen 95,000/s | none stated |
| PQShield PQPerform-Lattice | ML-KEM, ML-DSA | not published | “high-throughput hardware accelerator” (NIST CAVP entry) | not stated on the CAVP entry |
Taken from the hardware entries of our Migrate product catalog and checked against each vendor’s own page. Note how differently “side-channel protected” is used: some cores only promise constant time (no timing leak), some add simple-power-analysis resistance, and only a few claim masking against differential power analysis — often sold as a separate, larger variant.
Constant-time means the circuit takes exactly the same time and follows the same steps whatever the secret key is, so a stopwatch learns nothing. It is cheap. Masking goes further: every secret value is split into random pieces (“shares”) that are processed separately and only recombine at the end, so the power drawn or radio noise emitted at any moment is unrelated to the key. It defeats power and electromagnetic analysis — but every share has to be carried, refreshed with fresh randomness and recombined carefully, which costs extra area, randomness and time, and the cost grows with the number of shares.
| Implementation | Protection | Cost vs unprotected | Platform |
|---|---|---|---|
| Caliptra Adams Bridge, ML-KEM-1024 / ML-DSA-87 | 2-share masking + shuffling (two parallel NTT engines) | Zero extra cycles (ML-KEM decapsulation 11,054 vs 11,056 cycles) — “the only cost of enabling masking is area” | ASIC |
| OpenTitan KMAC (Keccak) | First-order domain-oriented masking | Keccak round logic “more than twice” the area; SHA3-224 throughput 3.43 → 1.26 bytes/cycle | ASIC |
| Kyber-512 (Kamucheka et al.) | Hiding, then hiding + masking | Hiding: 1.83× cycles, 1.6× resources vs baseline; adding masking: a further 1.08× cycles, 1.06× resources | FPGA (Virtex-7) |
| ML-DSA (Raj et al.) | First-order masking | 1.127× LUTs, 1.2× flip-flops — but 378× execution time | FPGA (Kintex-7 / Zynq) |
| Kyber768 (Bos et al.) | First, second, third order | First order 3.5× the cycles of the optimised unprotected code (3.1M vs 0.88M); second order 44.3M cycles; third order 115.5M | Arm Cortex-M4F |
| Dilithium3 signing (Azouaoui et al.) | 2 shares (first order) | 4.3× slower (randomised signing), 7.32× slower (deterministic) | Arm Cortex-M4 |
| Dilithium variant (Migliore et al.) | Order 1 / 2 / 3 | About 5.6× / 11.6× / 28× slower than unmasked | general-purpose processor |
- In software, first-order masking costs roughly 3–7× the time for Kyber and Dilithium, and the cost climbs steeply with each extra order (Kyber768: 3.5× at first order, over 100× at third).
- In hardware the bill can be paid in area instead of time: Adams Bridge runs two NTT engines in parallel so masked decapsulation takes no extra cycles — while masked Keccak in OpenTitan more than doubles the round logic.
- Masking a design that was not built for it can be disastrous: one first-order ML-DSA FPGA design added only 13–20% area but ran 378× slower.
- Always ask which attacks a “protected” product covers: constant-time (timing), SPA, or DPA/masking — and at what order.
Published, not measured by us. Our own KV260 engines were built for speed and have not yet been evaluated against power or electromagnetic analysis — a planned later phase.
Many crypto engines are microcoded: fixed hardware (multipliers, Keccak, memories) follows a stored list of steps, like a kitchen following a recipe card. If that card is in writable memory, a signed update can change the step order, the parameter sets or fix a bug — but never add hardware the chip lacks. If the card is burned into ROM, nothing changes after manufacture. Adams Bridge is microcoded, and its microcode is in ROM.
| Approach | In plain English | Can change after manufacture | Cannot change |
|---|---|---|---|
| Hard-wired or ROM microcode | The step-by-step recipe exists, but it is burned into read-only memory. | nothing in the algorithm | the recipe, the parameter sets |
| Programmable crypto co-processor (e.g. OpenTitan OTBN) | A small processor with its own instruction set; the host loads a new program. | the software — RSA, ECC, even a PQC NTT | its instructions and memory size — fast PQC needed new instructions and 8× more memory (a new chip) |
| Fixed building blocks, algorithm in firmware (e.g. Sapphire) | Keccak, NTT and modular arithmetic in hardware; each scheme’s steps in updatable firmware. | which lattice scheme runs, its parameters | the building blocks themselves |
| Embedded FPGA (eFPGA) inside an ASIC Flex Logix press release, 7 Nov 2022: “all forthcoming updates and algorithm modifications can be supported by simply reprogramming eFPGA” | A patch of FPGA fabric built into the chip, reprogrammed with a new bitstream. | whole algorithm circuits | the size of the patch and the rest of the chip |
The practical middle path for crypto agility in silicon: put the stable building blocks (Keccak, NTT, modular multiply) in hardware and keep each scheme’s steps in updatable firmware — most of the ASIC’s efficiency, much of the FPGA’s flexibility.
NPUs are arrays of 8-bit multiply-accumulate units fed a pre-compiled neural-network graph, designed to tolerate rounding. PQC needs exact arithmetic on 23-bit (ML-DSA) and 12-bit (ML-KEM) numbers, and Keccak is bitwise logic with no multiplication at all.
Lessons: Owning an Instruction Is Not Using It
- AES ran in software on every board until a compile flag was set. AES-128-CBC encrypt at 16 KiB, A-B-A-B, 2026-09-26. Same crate, with vs without the build flag that enables the ARMv8 AES instructions. Before the fix, AES ran in software on every board (peak 75 MB/s on the MX95). Speed-up at 16 KiB: M4 Pro 10.8–12.9×, i.MX 95 7.8–9.1×, KV260 7.4–10.0×.
- A library chose the wrong kernel for our core. Cortex-A55, AWS-LC, 2026-09-16. AWS-LC picked a NEON/Karatsuba Montgomery kernel tuned for Neoverse cores; on the A55 it LOST to the plain scalar kernel. Forcing the scalar kernel: RSA-2048 +29%, RSA-3072 +62%, RSA-4096 +22%.
- An instruction can be present, unused — and then only a modest win. Build audit, 2026-09-27: the M4 Pro reports FEAT_SHA3, but our engine builds the Rust `keccak` crate (used by SLH-DSA) without its `asm` feature, so SLH-DSA’s SHAKE runs the portable software permutation even there. (AWS-LC’s own Keccak, used by ML-DSA/ML-KEM, already uses the instructions.) Measured A/B on the M4 Pro, same engine with the feature switched on: SHAKE 1.17–1.20×, SLH-DSA-SHAKE-128s signing 1.19× — a real but modest gain, far below AES’s 7–13×, because Apple’s wide cores already run the software permutation well.
- An accelerator slower than the CPU is not an accelerator. Our first Keccak FPGA engine needed 58.7 µs per hash where the A53 needed about 12 µs — we removed it, and dropped SHA-2 from the fabric for the same reason.
The practical checklist: find where the time actually goes; check which instructions the core really has and that the build uses them; prefer acceleration inside the CPU for small, latency-sensitive operations; offload only whole operations; and measure end to end, under concurrency, against the best CPU path — not the reference code.
Related modules
Seven acceleration models and five building blocks in plain English, then our own measurements and FPGA labs.
Related modules
Check your understanding
8 questions on PQC Hardware Acceleration, each with its answer and the reason.
Take the quizNext step
Produce the artifact: Infrastructure Modernization PlannerThis module belongs to phase 6 (Infrastructure & Performance); Infrastructure Modernization Planner produces a deliverable of that phase in the Command Center.
Learning module content can be inaccurate. Please double-check its information. Report inaccuracies in PQC Today GitHub Discussions.