DOI: 10.3390/electronics15184300 ISSN: 2079-9292

Hardware Architectures for ECC Scalar Multiplication: A Trade-Off Between Performance and FPGA Resources

Natalya Glazyrina, Nikita Nakonechnyy, Renat Ibrayev, Olzhas Satiyev, Gulsipat Abisheva, Bibigul Razakhova

Elliptic curve cryptography (ECC) combines strong security with short keys, which makes it a good fit for embedded and IoT devices. In these systems, scalar point multiplication is usually the most demanding operation. Published hardware results are hard to compare because researchers work with different platforms, operand widths, and evaluation methods. DSP-free, logic-only implementations of GOST R 34.10-2012 over GF(p), especially those assessed by both cycle count and hardware cost, are uncommon. This work presents an open-source, reproducible RTL implementation and evaluates three architectures under identical conditions at 256 and 512 bits, using neither DSP blocks nor BRAM. Architecture A relies on a direct full-size modular multiplier and serves as a non-viable baseline. Architecture B carries out bitwise Montgomery multiplication, while Architecture C applies a regular LadderStep scheme. All three designs were written in SystemVerilog and checked against a reference software model using 100 test vectors. They were then synthesized on the Tang Primer 20K and Tang Mega 138K FPGAs. Architecture A had the lowest average cycle count, but its performance depended on the input and it exceeded the available resources. Architectures B and C both completed in constant time. Compared with Architecture B, Architecture C used 5.53–5.55× as many cycles, while cutting the Gowin-reported logic-element count by 4.88–7.13× and the flip-flop count by 2.56–3.13×. It also produced a smaller implementation area in the preliminary ASIC-oriented synthesis. That makes Architecture B the stronger option for resource-rich platforms where speed comes first; Architecture C fits better on resource-constrained FPGAs and in compact ASIC designs.