AI Chat Paper
Note: Please note that the following content is generated by AMiner AI. SciOpen does not take any responsibility related to this content.
{{lang === 'zh_CN' ? '文章概述' : 'Summary'}}
{{lang === 'en_US' ? '中' : 'Eng'}}
Chat more with AI
Article Link
Collect
Submit Manuscript
Show Outline
Outline
Show full outline
Hide outline
Outline
Show full outline
Hide outline
Cover Article

Homomorphic Processing Unit

State Key Laboratory of Processors, Institute of Computing Technology, Chinese Academy of Sciences, Beijing 100190, China
Show Author Information

Abstract

Fully homomorphic encryption (FHE) enables computation over encrypted data without decryption, ensuring data confidentiality throughout the processing pipeline. However, the complexity and heterogeneity of FHE schemes like CKKS, BFV, and TFHE pose challenges for unified hardware design. In this paper, we propose the Homomorphic Processing Unit (HPU), a general-purpose platform supporting multiple FHE schemes. Unlike fixed-function accelerators, HPU is implemented as a RISC-V extension and introduces a dedicated homomorphic instruction set architecture (H-ISA), comprising micro-instructions for minimal compute kernels and macro-instructions for integration with general-purpose models. Core operations like the number-theoretic transform (NTT) and automorphism are abstracted into a unified instruction layer, mapped onto specialized homomorphic compute units. HPU integrates a collaborative computation and storage design, utilizing a polynomial-level instruction set and a polynomial-granular memory management unit (MMU) to optimize memory and reduce data movement. By combining a complete ISA with a general-purpose computing core, HPU can achieve efficient computing, control, and scheduling. We evaluate the HPU prototype using four multi-FHE benchmarks on field-programmable gate array (FPGA) and 7 nm ASIC. The key results are as follows. 1) HPU achieves up to 1596x and 1174x speedup over CPU for CKKS-like and TFHE operations, respectively. 2) Compared with SOTA FPGA solutions for CKKS and TFHE, HPU achieves 1.36x and 1.2x improvement, respectively. 3) HPU's ASIC implementation achieves 1.36x speedup over state-of-the-art TFHE accelerators for long short-term memory (LSTM).

Electronic Supplementary Material

Download File(s)
JCST-2412-15125-Highlights.pdf (779.8 KB)

References

【1】
【1】
 
 
Journal of Computer Science and Technology
Pages 845-861

{{item.num}}

Comments on this article

Go to comment

< Back to all reports

Review Status: {{reviewData.commendedNum}} Commended , {{reviewData.revisionRequiredNum}} Revision Required , {{reviewData.notCommendedNum}} Not Commended Under Peer Review

Review Comment

Close
Close
Cite this article:
Yang Y-H, Xu X-C, Lu H, et al. Homomorphic Processing Unit. Journal of Computer Science and Technology, 2026, 41(3): 845-861. https://doi.org/10.1007/s11390-026-5125-0

6

Views

0

Crossref

0

Web of Science

0

Scopus

0

CSCD

Received: 06 March 2025
Accepted: 03 June 2026
Published: 01 May 2026
© Institute of Computing Technology, Chinese Academy of Sciences 2026