Real-time saliency-guided deep watermarking on Kria KV260: a Vitis AI accelerated proxy architecture


Gedik M. İ., COŞKUN A.

Frontiers in Computer Science, cilt.8, ss.1-13, 2026 (ESCI, Scopus)

  • Yayın Türü: Makale / Tam Makale
  • Cilt numarası: 8
  • Basım Tarihi: 2026
  • Doi Numarası: 10.3389/fcomp.2026.1886719
  • Dergi Adı: Frontiers in Computer Science
  • Derginin Tarandığı İndeksler: Emerging Sources Citation Index (ESCI), Scopus, INSPEC, Directory of Open Access Journals
  • Sayfa Sayıları: ss.1-13
  • Anahtar Kelimeler: deep watermarking, DPU acceleration, FCN-ResNet50, Kria KV260, MobileNetV2, noise-assisted dithering, PL-PS DNA
  • Gazi Üniversitesi Adresli: Evet

Özet

The rapid growth of Industrial Internet of Things (IIoT) ecosystems and autonomous surveillance networks has necessitated the shift of digital content security from central servers to the data-generating edge. In industrial security scenarios, operators must continuously monitor live video streams and optionally capture high-resolution, verifiable evidence snapshots. However, high computational costs and hardware-based precision losses prevent the real-time execution of deep learning-based watermarking models on resource-limited embedded devices. This study proposes a hardware-aware and semantically oriented real-time watermarking architecture running on the Xilinx Kria KV260 FPGA platform. The fundamental innovation of the proposed system is the asynchronous “Proxy Frame” software architecture, which allows heavy Convolutional Neural Networks (CNNs) to run in the background, isolated from the live video stream. Thus, highly secure watermarking can be performed at 1080p resolutions without compromising the fluidity of the 30 FPS live preview offered to the operator. Furthermore, a Noise-Assisted Dithering technique, inspired by stochastic resonance, was used to mitigate the “Signal Fading” problem arising from watermark signal loss during 8-bit integer (INT8) quantization on the Deep Learning Processing Unit (DPU). By injecting controlled Gaussian noise into the quantized inference pipeline, the detectability of subthreshold weak watermark signals was increased, reducing the hardware Bit Error Rate (BER) from 18.75% to 13.06%. MobileNetV2 and ResNet50-FCN based saliency models achieved an average PSNR visual quality of 39.86 dB by concealing the payload in perceptually insignificant regions. Under clean hardware conditions, the system achieved a baseline BER of 2.6% (with MobileNetV2). In attack tests, the BER remained below 15% for MobileNetV2 under most standard degradations, with the stated exceptions of severe JPEG compression and deliberate geometric cropping. This end-to-end hardware and software solution demonstrates that theoretical deep learning models can be integrated into industrial smart cameras without experiencing performance bottlenecks, particularly for static infrastructure monitoring.