Secure speech processing and communication
Abstract
Keynote given for the Fall School 2025 in Bilbao, Spain.
Full text
Secure Speech Processing and Communication Antonio M. Peinado Research Group on Signal Processing, Multimedia Transmission and Speech/Audio Technologies (SigMAT) Antonio M. Peinado Secure Speech Processing 1 / 46
Contents 1Introduction: why secure speech processing? 2Speech spoofing and deepfake detection ▶Metrics: which are our goals? ▶Corpora: which are our raw materials? ▶Systems: which are our tools? 3Speech watermarking: a proactive solution ▶Classical/DSP-based techniques ▶DNN-based techniques ▶Watermarking and deepfake detection Antonio M. Peinado Secure Speech Processing 2 / 46
1. INTRO: WHY SECURE SPEECH PROCESSING? Spread of speech technology applications involves threats and misuses [Europol22,Mubarak23]. Legal acknowledgement of these threats: ▶EU GPDR: ”... right to the protection of personal data ... This should apply in particular to the processing of personal data in the audiovisual field ...”. ▶EU AI Act: ”AI systems presenting limited risk, such as systems that interact with humans (i.e. chatbots), emotion recognition systems, biometric categorisation systems, and AI systems that generate or manipulate image, audio or video content (i.e. deepfakes), would be subject to a limited set of transparency obligations. Threads/attacks considered here: false speech (digitally created or manipulated) for impersonation. Antonio M. Peinado Secure Speech Processing 3 / 46
Possible attack/countermeasures scenarios: ▶Impersonation in Biometric Systems Physical Access (PA) att. Logical Access Access (LA) att. Mimicking Pre-recorded speech Replay TTS/VC TTS/VC playback Countermeasures: anti-spoofing (presentacion attack detection, PAD) ▶Impersonation for social misuses New deep synthesis (TTS or VC) or other DSP technologies may be used to generate or manipulate speech signals. Goals: desinformation, blackmail or denigration. Countermeasures: deepfake detection Examples of common attacks: Replay: Genuine Audio 1 Audio 2 TTS/VC: Genuine Audio 1 Audio 2 Antonio M. Peinado Secure Speech Processing 4 / 46
Countermeasures (CMs) Passive (classical) approaches: ▶Standalone spoofing and deepfake detection. Binary (genuine or spoofed/fake speech) detection. Biometrics: CM facilitates the ASV decision. ▶Spoofing-Aware Speaker Verification (SASV). Single detector approach (possibility of CM+ASV joint optimization). Trials can be classified into 3 classes: Class C1C2C3 System / Trial Genuine target Genuine non-target Spoof target ASV Positive Negative - CM Positive - Negative ASV + CM Positive Negative Negative Proactive solutions for TTS/VC: watermarking. ▶An imperceptible watermark (indicating a synthetic origin) is embedded in the speech signal. ▶It requires the collaboration of the TTS/VC provider. ▶The watermark can be erased or tampered. Antonio M. Peinado Secure Speech Processing 5 / 46
Research boosted by Challenges Challenge Standalone Spoofing & Deepfake Detection SASV PA LA/Deepfake Det. (TTS/VC) Replay TTS/VC Playback TTS/VC ASVspoof 2015 x ASVspoof 2017 x ASVspoof 2019 x x ASVspoof 2021 x x SASV 2022 x ADD 2022 x ADD 2023 x ASVspoof 5 (2024) x x ASVspoof Automatic Speaker Verification Spoofing and Countermeasures Challenge [ASVspoof]. SASV Spoofing-Aware Speaker Verification Challenge [SASV22]. ADD Audio Deepfake Detection Challenge [ADD]. Antonio M. Peinado Secure Speech Processing 6 / 46
2.1. SPOOF & DEEPFAKE DETECTION: METRICS Detection: binary classification problem (genuine/spoof). Given an input utterance x, the spoof detector must provide an score s=f(x)to decide the class: y(x) = 1 if s≥τ(accepted) 0 if s< τ (rejected) Performance is afected by the chosen threshold! Biometrics applications: CM and ASV are integrated. ▶Tandem: decision-level integration, CM and ASV independent. ▶Single detector: CM and ASV are fully integrated. Antonio M. Peinado Secure Speech Processing 7 / 46
Basic metrics for binary detection Error rates in binary classifiers: ▶False acceptance rate (FAR): FAR(τ) = Zτ −∞ p(s|genuine)ds ≈No. spoof detected as genuine No. spoof A low value indicate secure system ▶False rejection (miss) rate (FRR): FRR(τ) = Z∞ τ p(s|spoof)ds ≈No. genuine detected as spoof No. genuine A low value indicate user-friendly system Problem: both goals are opposite. A FAR/FRR trade-off must be achieved! Antonio M. Peinado Secure Speech Processing 8 / 46
Detection error tradeoff (DET) curve: FAR(τ)vs FRR(τ). FRR EER system 1 system 2 FAR ▶Each point represents an operating point (τvalue). ▶Allows to provide a FAR value given FRR (or viceversa). ▶Equal Error Rate (EER): point where FAR =FRR. Score distribution and EER: Antonio M. Peinado Secure Speech Processing 9 / 46
SASV (single detector) evaluation: a-DCF Tandems can be built with many different configurations. Agnostic DCF (a-DCF): DCF independient of the tandem architecture. Requirement: the whole CM/ASV system must employ a single score and a single threshold τsasv for decisions. Definition of a-DCF (same error rates as t-EER): a-DCF(τsasv ) = πtar CmissPmiss(τsasv ) + πnonCfa,nonPfa,non(τsasv ) +πspoof Cfa,spf Pfa,spf (τsasv ) Normalization w.r.t. a system which accepts (τsasv → −∞) or rejects (τsasv →+∞) every trial: a-DCFnorm(τsasv ) = a-DCF(τsasv )/a-DCFdef a-DCFdef = m´ın{πtar Cmiss, πnonCfa,non +πspoof Cfa,spf } Minimum normalized a-DCF: a-DCFmin norm = m´ın τsasv a-DCFnorm(τsasv ) ASVspoof 5 Parameters: same as t-DCF. Antonio M. Peinado Secure Speech Processing 16 / 46
2.2. SPOOF & DEEPFAKE DETECTION: CORPORA ASVspoof 2015 Goal: standalone spoofing detection. VC and TTS attacks. Based on VCTK: hemi-anechoic from imni-directional microphone at 96 kHz (downsampled at 16 kHz). Composition: ▶No speaker overlap across subsets. ▶Genuine speech: VCTK utts. without modification. ▶Spoof speech: same utts. modified for VC, same transcripts for TTS. Spoofing techniques: all of them based on classical DSP. ▶Training: 3 VCs and 2 TTS (HMM-based). ▶Development: same as training. ▶Evaluation: same (known attacks) plus other (unkown attacks) 4 classical DSP-based VC and 1 TTS (MaryTTS). Antonio M. Peinado Secure Speech Processing 17 / 46
ASVspoof 2017 corpus Goal: standalone spoofing detection. Replay attacks. Based on RedDots: crowdsourced audios from Android devices. Composition: ▶No speaker overlap across subsets. ▶Genuine speech: RedDots utts. without modification. ▶Spoof speech: same utts. replayed. Antonio M. Peinado Secure Speech Processing 18 / 46
Replay configuration (RC): {loudspeaker, acoustics, microphone}. ▶26 acoustic environments: ⋆Outer and inner recordings ⋆Inner: different types of rooms (different reverberation level). ⋆Different acoustic noise conditions. ⋆Anechoic and perfect/wired replays: included as very challenging conditions. ▶26 Loudspeakers: different qualities, including audio interfaces for wired environments. ▶25 ASV microphones: different qualities, including audio interfaces for wired environments. Each RC involves different recording sessions. Possible RCs: 26 ×26 ×25!! Only 61 distinct RCs are considered. Distribution: ▶Training: 3 RCs ▶Development: 10 RCs ▶Evaluation: 57 RCs (including those from Dev). Antonio M. Peinado Secure Speech Processing 19 / 46
ASVspoof 2019 corpus Goal: standalone spoofing detection with ASV-constrained evaluation. ▶Two scenarios: TTS/VC (LA) and (simulated) Replay attacks (PA). ▶Unified LA/PA data generation. Based on 107 speakers of VCTK (20 train + 20 dev + 67 eval). Generation of LA partition: Antonio M. Peinado Secure Speech Processing 20 / 46
Generation of simulated PA partition: ▶Microphones: ideal. ▶27 Acoustic environments simulated (S,R,Ds): 3 room size (S), 3 T60 (R), 3 distance to ASV microphone (Ds). ▶9 Attack types (Da,Q): Loudspeakers: 41 devices (mod.), 3 quality ranges (Q). Distance speaker to attacker micro: 3 intervals (Da). Antonio M. Peinado Secure Speech Processing 21 / 46
ASVspoof 2021 corpus Three scenarios: ▶Standalone anti-spoofing (ASV-constrained evaluation). Two scenarios: TTS/VC (LA) and Replay attacks (PA). ▶Deepfake detection task (DF): TTS/VC attacks not related to ASV. No new training data: training with ASVspoof 2019 training partitions. Only evaluation data.Focus on generalization to new conditions. ▶Progress data provided. ▶LA/PA: data from VCTK (same speakers as in 2019). ▶DF: data from ASVspoof 2019 and VCCs challenges 2018 and 2020. LA scenario: ASVspoof 2019 evaluation data but including new transmission channel conditions. ▶Condition C1: same ASVspoof 2019 condition. ▶Transmission: may include transcoding and packet loss. Antonio M. Peinado Secure Speech Processing 22 / 46
PA scenario: real recordings on multiple combinations of rooms/noise, microphones, loudspeakers and distances to microphones. Not totally real data!: no real talkers (substituted by a high quality loudspeaker). Antonio M. Peinado Secure Speech Processing 23 / 46
DF scenario: ▶Goal: ”detection of deepfakes in compressed audio used in television and media hosted on news websites and social media platforms, etc.” ▶Data from ASVspoof 2019 LA, and VCC 2022 and VCC 2018 challenges. ▶More than 100 different attack types. ▶Data with 9 compression conditions: ⋆Condition C1: same ASVspoof 2019 condition. ⋆C8/C9: cases when the audio content change from one web platform to another. Antonio M. Peinado Secure Speech Processing 24 / 46
Other corpora SASV 2022 (joint ASV/CN)) [SASV22]: VoxCeleb2 for ASV training, ASVspoof 2019 for CM training and SASV evaluation. ADD 2022/23 challenges [ADD22]: TTS/VC data from mandarin corpora AISHELL-1, -3 and -4. Large number of synthesis methods. Noisy data. Fully and partially faked audio. In-The-Wild [ITW22]: 37,9 hours of audio (average 4,3 sec/utt) from celebrities, includes 17,2 hours of fake audios (from publicly available audio/video sources, same speakers). Intended for generalization evaluation. ASVspoof 5 (2024): ▶Source: Multilingual Librispeech (English). 32 TTS/VC attacks. ▶Intended for both standalone spoof/deepfake detection and SASV. ▶Large # of speakers (≈2000!) to allow development of SASV systems. ▶Includes adversarial attacks (refined for surrogate models). SpoofCeleb [SC25]: developed from VoxCeleb1, very large database (1251 spk, 2.5M utt) for spoof/deepfake detection and SASV in the wild. 23 TTS attacks. Antonio M. Peinado Secure Speech Processing 25 / 46
RawNet2 [Tak23]: beyond hand-crafted features ⇒end-to-end solution. ▶GRU for frame aggregation. ▶Residual blocks for increased discrimination. ▶A TRAINABLE filterbank. ▶EER Results: Challenge LA (Win) DF (Win) ASVspoof 2019 3.50 (0.22) – ASVspoof 2021 (base) 9.50 (1.32) 22.38 (15.64) ASVspoof 5 (base1) – 36.04 (8.61) ASVspoof 5 (base2) – 29.12 ” ASVspoof 5: track 1 (standalone CM), closed condition. ASVspoof 21/5 base/base1: RawNet2. ASVspoof 5 base2: AASIST (includes RawNet2 encoder). What happened in ASVspoof 2021 PA (replay)? The winner was ... Vocoder: WORLD Classifier: PCA+GMM EER Result (PA): 24.26 % !! Antonio M. Peinado Secure Speech Processing 32 / 46
SYSTEMS: Incorporation of self-supervised learning What is an SSL speech model? ⇒similar to LLMs, it learns to predicts next samples given a signal context [Baevski20]. ▶(Pre)Trained on many hours of audio. ▶CNN as feature encoder. ▶Multiple transformer layers (each representating a different abstraction levels). ▶Quantization is introduced to deal with the non-discrete nature of speech. Why an SSL model? ▶Problem of systems until 2021 challenge: lack of diverse data ⇒ lack of generalization capabilities. ▶Hypoyhesis: use of SSL models trained with a huge amount of data (even for other tasks) may help to reduce overfitting. ▶A good representation (usually fine-tuned) may help to ligthen classifiers. Fine-tuning typically required (to see spoofed speech). Antonio M. Peinado Secure Speech Processing 33 / 46
Direct approach: use SSL to provide a (sort of deep) spectrogram to feed a (downstream) classifier: AASIST [Tak23] Conformer [Rosell´ o23] Fully connected layer (FC) Batch Normalization (BN) & SeLU Feed Forward Module (FFN) + Multi-Head Self Attention Module (MHSA) Convolutional Module (Conv) Feed Forward Module + + + Layernorm Linear Layer Conformer Encoder Classification Token Processed Classification Token · L Conformer Blocks Class W2V 2.0 Model XL s O' O X0 ·1/2 ·1/2 xclass xL 0Encoder Outputs Antonio M. Peinado Secure Speech Processing 34 / 46
Alternative use of SSL: combining features from different layers of the SSL model (frozen) [Martin24]. ▶Adapter: linear cmbination of layer embeddings. ▶Frame proc.: dimensionality reduction. ▶Time pooling: provides a single utterance embedding. ▶Score: cosine similarity. Results with SSL models (Wav2Vec 2.0): Technique Winner AASIST Conformer Layer Comb. ASVspoof 2021 LA 1.32 1.00 0.97 8.66 ASVspoof 2021 DF 15.64 3.69 2.58 1.16 Antonio M. Peinado Secure Speech Processing 35 / 46
SYSTEMS: SASV systems 1st approach: standalone CM and ASV. Example: ASVspoof-5 baseline 3 (B03): ▶Models: AASIST (CM) + ECAPA-TDNN (ASV). ▶Score fusion: ⋆Applies calibration to transform scores into LLR(x) = log P(H1|x) P(H0|x). ⋆Non-linear combination CM and ASV scores into a single one. 2nd approach: joint CM+ASV system. Example: ASVspoof-5 baseline 4 (B04): ▶SASV encoder: MFA-Conformer. ▶Three training stages: 1Only target/nontarget bonafide data is presented. 2Batch enlarged with vocoded (spoofed) speech. 3Fine-tuning with in-domain data. Antonio M. Peinado Secure Speech Processing 36 / 46
ASvspoof 5 (track 2 - SASV) results: min a-DCF (t-EER) Closed condition Open condition REF 0.6869 – B03 (1st approach) 0.6803 (28.78 %) – B04 (2nd approach) 0.5741 – Winner 1 (T45) 0.2814 0.0756 Winner 2 0.2954 (9.58 %) 0.1156 (4.32 %) Comments: ▶REF: standalone baseline ASV + random CM scores. ▶B03: CM subsystem does not help. ▶Winner: standalone CM + ASV systems Antonio M. Peinado Secure Speech Processing 37 / 46
3. WATERMARKING Initially intended for copyright protection, copy protection, etc ... Application to deepfake detection: ▶Proactive solution: requires the cooperation of service providers. ▶Regulated in the EU AI Act (art. 133): ”... it is appropriate to require providers of those systems to embed technical solutions that enable marking in a machine readable format ... that the output has been generated or manipulated by an AI system and not a human.” Embedding of complementary information into multimedia content. Objectives: ▶Imperceptibility: low signal distortion. ▶Watermark detection accuracy: low bit error rate (BER). ▶Robustness against processing and attacks (usually, removal). Common attacks: LP/BP/HP filtering, echo, additive noises, sample suppression, amplitude alterations, desynchronization, etc. Antonio M. Peinado Secure Speech Processing 38 / 46
Classical/DSP-based techniques WM Tech. Main features Resilience against attacks Bit rate (bps) Echo [Lin15] - Time domain technique using echo hiding: watermarks embedded as echoes (perceived as weak resonance, not disturbing to human hearing). - Detection easily performed in the auto-cepstrum domain. - Repetition coding and variable frame length. - Not robust to pitch modifications or echo attacks. Natural echoes may cause false positives. 3.43 DCT [Hu15] - Frame-based DCT domain method. Message embedded via quantization index modulation (over DCT vectors). - Imperceptibility ensured by auditory masking constraints (distortion kept below threshold). - Auditory masking improves resistance to attacks. - A synchronization procedure mitigates desynchronization attacks. - Robust to most common DSP attacks and audio compression. 24 Patchwork [Nat17] - Multi-layer watermarking based on the patchwork method in the DCT domain. - DCT reordering reduces perceptual quality degradation. - Error buffer and frequency selection enhance robustness. - Resistant to several DSP attacks, compression, and desynchronization. 10 FSVC [Zhao21] - SVD-based method on the DCT domain using ratios between singular values of consecutive segments. - Operates on mid-frequency bands only, minimizing perceptual impact, assisted by the embedding procedure. - Error buffer supports robustness to common DSP attacks (re-sampling, echo, requantization, and AWGN). - Robust to DSP, compression, and desynchronization. 8 Norm [Saad19] - Combines DWT and DCT to embed 1 bit via the L2-norm of frame vectors using Haar wavelets (suitable for human auditory system). - Sensitive to silent frames due to norm-based embedding. - Spreading via sub-sampling decomposition improves robustness against attacks (excluding re-sampling). - Note: no reported results for filtering attacks. 7 Antonio M. Peinado Secure Speech Processing 39 / 46
DNN-based techniques FIrst approach [Pav22,Pei24]: Embedder Attacks Detector Watermarked Speech Speech Signal Extracted Watermark Watermark ▶Embedders typically adopt an encoder/decoder form (U-Net [Pav22]): CONV2D CONV2D DOWN UP DOWN DOWN DOWN UP UP DOWN CONV2D UP UP 256x2x16 128x4x32 64x8x64 32x16x128 16x32x256 8x64x512 256x2x16 Message 512x2x16 8x64x512 16x32x256 128x4x32 64x8x64 32x16x128 2x64x512 STFT 2x64x512 STFT ▶The watermark is inserted in the encoder/decoder interface. ▶Detectors: much lighter architecture (e.g., CNN). ▶Attacks: integrated as non-trainable layers. ▶Training: losses for imperceptibility (e.g., MMSE, MAE, PMSQE, etc) and accuracy (e.g., BCE). ▶Performance: 256 bits/s, PESQ ≈4,2; SNR ≈35 dB; BER ≈0 % (no attack)). Antonio M. Peinado Secure Speech Processing 40 / 46
Adapting watermarking to deepfake detection AudioSeal (Meta) [SRom24]. ▶AI-generated speech can be inserted in genuine speech ⇒AudioSeal allows localized watermark detection (up to the sample level!). ⋆Zero-bit watermarking: watermark present just means synthetic audio. ⋆Multi-bit watermarking: several bits (allows detection of the specific audio synthesis generator). ▶Embedder: encoder/decoder based on EnCodec (Meta). ▶Detector: (EnCodec encoder) + (convolutional + linear layer; one per bit). ▶Perceptual losses: MAE, multiscale Mel-spectrogram loss, and TF-loudness loss (to encourage watermark masking). ▶BCE for watermark detection. ▶Very robust against attacks (including sampling rate changes). Attacks presented during training. ▶Performance (zero-bit): PESQ ≈4,5; SNR ≈26 dB; BER ≈0 % (no attack). Antonio M. Peinado Secure Speech Processing 41 / 46