A System-Level Analysis of Cross-Domain Generalization Failure in Audio Deepfake Detection


OYUCU S., ARSLAN B., SAĞIROĞLU Ş.

IEEE Access, 2026 (SCI-Expanded, Scopus)

  • Yayın Türü: Makale / Tam Makale
  • Basım Tarihi: 2026
  • Doi Numarası: 10.1109/access.2026.3716180
  • Dergi Adı: IEEE Access
  • Derginin Tarandığı İndeksler: Science Citation Index Expanded (SCI-EXPANDED), Scopus, Compendex, INSPEC, Directory of Open Access Journals
  • Anahtar Kelimeler: ASVspoof, audio deepfake detection, computational efficiency, Conformer, cross-domain evaluation, deployment-oriented assessment, domain shift, generalization, HeAR, HuBERT, In-the-Wild, MFCC, robustness, TCM-Conformer, transfer learning, Wav2Vec
  • Gazi Üniversitesi Adresli: Evet

Özet

Audio deepfake detectors often report high accuracy on individual benchmarks, yet their reliability under domain shift remains largely untested. This study presents a system-level analysis of cross-domain generalization failure, evaluating five hybrid architectures (CNN–LSTM, TCN, TCN–LSTM, Conformer, and TCM-Conformer) across four feature extractors (MFCC, Wav2Vec 2.0, HuBERT, and HeAR) and four distinct acoustic domains: ASVspoof 2019 (LA), ASVspoof 2021 (DF), the DEEP-VOICE dataset (RVC-based voice conversion), and the In-the-Wild corpus representing real-world deployment conditions. Across 320 cross-domain train–test configurations, severe and consistent degradation is observed: while models achieve up to 99.98% accuracy under in-domain evaluation, performance collapses to near-chance or below-chance levels under domain shift. The most extreme failure occurs when MFCC-based models trained on In-the-Wild are evaluated on protocol-based benchmarks, with accuracy collapsing to as low as 5.55% (TCN–LSTM architecture), representing a complete functional inversion of the detection logic. Substantial degradation is also observed within benchmark families:Wav2Vec-based models trained on ASVspoof 2021 and evaluated on ASVspoof 2019 lose up to 18.5 percentage points despite the two datasets sharing the same protocol structure. Critically, the observed degradation is architecture-invariant in direction: all five evaluated architectures exhibit consistent collapse patterns across all four feature extractors, confirming that neither increased model complexity nor attention-based design is sufficient to address domain gaps; however, the severity of collapse differs markedly across feature extractors, with MFCC exhibiting the most extreme degradation under real-world training conditions and HeAR retaining substantially higher cross-domain accuracy under the same transfer conditions. Computational efficiency analysis further confirms that the evaluated architectures are suitable for practical deployment, with mean inference latencies ranging from 0.001 to 0.052 ms per sample and throughput ranging from approximately 27,000 to 361,000 samples per second depending on the feature extractor, with MFCC-based configurations exceeding 100,000 samples per second and self-supervised and health-acoustic extractors achieving approximately 27,000 to 43,000 samples per second. These findings demonstrate that high in-domain accuracy does not translate to deployment reliability and that single-dataset evaluation provides little evidence of real-world robustness without explicit cross-domain validation and adaptation mechanisms.