DOI: 10.4218/etrij.2025-0418 ISSN: 1225-6463

Accelerating SMAUG‐T using the number theoretic transform on modern high‐end CPUs

WooHyung Ko, YoungBeom Kim, Seog Chung Seo

Abstract

We enhance the performance of SMAUG‐T, the winning algorithm of South Korea's Post‐Quantum Cryptography competition, by applying an NTT‐friendly ring homomorphism tailored to x86/64 and AVX2/AVX‐512. This paper makes two main contributions: (i) We integrate the NTT into the reference C(x86/64) implementation of SMAUG‐T and apply a more efficient NTT‐friendly ring homomorphism optimized for parallel processing. Previous attempts have also been made to apply the NTT to SMAUG‐T on AVX2; however, they have not addressed the methodology for selecting optimal primes in close conjunction with implementation optimization. (ii) We redesign several state‐of‐the‐art PQC optimization techniques, including layer merging, lazy reduction, shuffle, and negative wrapping convolution, to suit SMAUG‐T in parallel environments. Our C‐based x86/64 implementation achieves a performance improvement on matrix‐vector multiplication compared with the reference, whereas the AVX‐512 assembly implementation yields a improvement over AVX2 in the same operation.

More from our Archive