PrivTI: Efficient End-to-End Privacy-Preserving Inference for Transformer-based Models in MLaaS
Document Type
Article
Publication Date
2026
Abstract
Machine Learning as a Service (MLaaS) offers convenient access to high-quality transformer-based model inference without local deployment, but it requires uploading sensitive data to the server, raising significant privacy concerns. Recent advances in homomorphic encryption (HE) and secure multi-party computation (MPC) have shown great potential for enabling privacy-preserving inference. However, existing schemes face challenges with transformer-based models due to massive matrix multiplications and complex non-linear functions. To address these challenges, we propose PrivTI, an efficient end-to-end privacy-preserving inference framework for transformer-based models. Specifically, we design an HE-based single-instruction multiple-data (SIMD) encoding scheme to perform matrix multiplication efficiently. To handle non-linear functions, we develop MPC-friendly protocols for core functions like GELU, Softmax, and LayerNorm, and further extend support to SiLU and RMSNorm. Moreover, PrivTI supports conversions between HE and MPC, enabling end-to-end privacy-preserving inference. Experimental results demonstrate that PrivTI achieves speedups 1.9-45.2× in matrix multiplication and 1.1-18.9× in non-linear function evaluation compared to baseline schemes IRON, BOLT, and NEXUS under the Local Area Network (LAN) setting. Additionally, PrivTI reduces communication overhead by 1.1-4898.7×. Moreover, end-to-end benchmarks show 3.3-9.8× improvements in computation speed and 1.6-6.2× reductions in communication overhead. © 2026 IEEE.
Recommended Citation
Luo, Mingshun; He, Haolei; Yang, Wenti; Li, Meng; Wu, Longfei; Zhang, Zijian; and Guan, Zhitao, "PrivTI: Efficient End-to-End Privacy-Preserving Inference for Transformer-based Models in MLaaS" (2026). College of Health, Science, and Technology. 1204.
https://digitalcommons.uncfsu.edu/college_health_science_technology/1204