IIUM Repository

Hybrid watermarking and adaptive misinformation for protection against AI model extraction in edge-deployed cyber-physical security

Razali, Fatimah Azzahrah and Habaebi, Mohamed Hadi and Al-Hussaini, Mohammed Abdulla Salim (2026) Hybrid watermarking and adaptive misinformation for protection against AI model extraction in edge-deployed cyber-physical security. Network, 6 (3). pp. 1-20. E-ISSN 2673-8732

[img] PDF - Published Version
Restricted to Repository staff only

Download (1MB) | Request a copy

Abstract

Artificial intelligence (AI) models are increasingly deployed in edge-deployed cyber-physical security systems for tasks encompassing monitoring, threat classification, and automated decision-making. While these models offer robust performance, their deployment through open or semi-open Machine Learning as a Service (MLaaS) interfaces exposes them to severe security threats, prominently model extraction attacks. In such attacks, an adversary systematically queries a target API to replicate the victim model’s behavior. This study proposes a novel hybrid defense framework combining Adaptive Misinformation (AM) and Trigger-Based Watermarking (WM) to protect AI models against black-box extraction. Utilizing a LeNet architecture, the victim model was trained on the MNIST dataset, while a simulated attack utilized 50,000 EMNIST samples to train a clone model. The framework employs Maximum Softmax Probability (MSP) for out-of-distribution (OOD) detection to identify suspicious queries and strategically inject misleading responses, alongside a fine-tuned embedded watermark for ownership verification. Experimental evaluations using 10-fold cross-validation reveal that the baseline extraction attack yielded a clone model accuracy of 96.32%. Upon implementing the AM + WM framework, clone model accuracy degraded significantly to 53.67%, while the victim model maintained an accuracy of 98.97%. Furthermore, the protected model achieved a 100% Trigger Match Rate (TMR), ensuring reliable intellectual property verification. The proposed framework provides a prototype validation for lightweight edge architectures to balance security, model utility, and ownership protection in cyber-physical deployments

Item Type: Article (Journal)
Uncontrolled Keywords: model extraction attacks, adaptive misinformation, trigger-based watermarking, cyber-physical systems, machine learning security
Subjects: T Technology > TK Electrical engineering. Electronics Nuclear engineering > TK5101 Telecommunication. Including telegraphy, radio, radar, television
T Technology > TK Electrical engineering. Electronics Nuclear engineering > TK7800 Electronics. Computer engineering. Computer hardware. Photoelectronic devices > TK7885 Computer engineering
Kulliyyahs/Centres/Divisions/Institutes (Can select more than one option. Press CONTROL button): Kulliyyah of Engineering > Department of Electrical and Computer Engineering
Depositing User: Dr. Mohamed Hadi Habaebi
Date Deposited: 17 Sep 2026 11:22
Last Update: 17 Sep 2026 11:22
Queue Number: 2026-09-Q5124
URI: http://irep.iium.edu.my/id/eprint/131287
Indexed In: Google Scholar

Actions (login required)

View Item View Item