SpeechVerifier: Robust Acoustic Fingerprint against Tampering Attacks via Watermarking

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Yao, Lingfeng, Huang, Chenpei, Wang, Shengyao, Xue, Junpei, Guo, Hanqing, Liu, Jiang, Chen, Xun, Pan, Miao
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912408252973056
author Yao, Lingfeng
Huang, Chenpei
Wang, Shengyao
Xue, Junpei
Guo, Hanqing
Liu, Jiang
Chen, Xun
Pan, Miao
author_facet Yao, Lingfeng
Huang, Chenpei
Wang, Shengyao
Xue, Junpei
Guo, Hanqing
Liu, Jiang
Chen, Xun
Pan, Miao
contents With the surge of social media, maliciously tampered public speeches, especially those from influential figures, have seriously affected social stability and public trust. Existing speech tampering detection methods remain insufficient: they either rely on external reference data or fail to be both sensitive to attacks and robust to benign operations, such as compression and resampling. To tackle these challenges, we introduce SpeechVerifer to proactively verify speech integrity using only the published speech itself, i.e., without requiring any external references. Inspired by audio fingerprinting and watermarking, SpeechVerifier can (i) effectively detect tampering attacks, (ii) be robust to benign operations and (iii) verify the integrity only based on published speeches. Briefly, SpeechVerifier utilizes multiscale feature extraction to capture speech features across different temporal resolutions. Then, it employs contrastive learning to generate fingerprints that can detect modifications at varying granularities. These fingerprints are designed to be robust to benign operations, but exhibit significant changes when malicious tampering occurs. To enable speech verification in a self-contained manner, the generated fingerprints are then embedded into the speech signal by segment-wise watermarking. Without external references, SpeechVerifier can retrieve the fingerprint from the published audio and check it with the embedded watermark to verify the integrity of the speech. Extensive experimental results demonstrate that the proposed SpeechVerifier is effective in detecting tampering attacks and robust to benign operations.
format Preprint
id arxiv_https___arxiv_org_abs_2505_23821
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle SpeechVerifier: Robust Acoustic Fingerprint against Tampering Attacks via Watermarking
Yao, Lingfeng
Huang, Chenpei
Wang, Shengyao
Xue, Junpei
Guo, Hanqing
Liu, Jiang
Chen, Xun
Pan, Miao
Cryptography and Security
Sound
Audio and Speech Processing
With the surge of social media, maliciously tampered public speeches, especially those from influential figures, have seriously affected social stability and public trust. Existing speech tampering detection methods remain insufficient: they either rely on external reference data or fail to be both sensitive to attacks and robust to benign operations, such as compression and resampling. To tackle these challenges, we introduce SpeechVerifer to proactively verify speech integrity using only the published speech itself, i.e., without requiring any external references. Inspired by audio fingerprinting and watermarking, SpeechVerifier can (i) effectively detect tampering attacks, (ii) be robust to benign operations and (iii) verify the integrity only based on published speeches. Briefly, SpeechVerifier utilizes multiscale feature extraction to capture speech features across different temporal resolutions. Then, it employs contrastive learning to generate fingerprints that can detect modifications at varying granularities. These fingerprints are designed to be robust to benign operations, but exhibit significant changes when malicious tampering occurs. To enable speech verification in a self-contained manner, the generated fingerprints are then embedded into the speech signal by segment-wise watermarking. Without external references, SpeechVerifier can retrieve the fingerprint from the published audio and check it with the embedded watermark to verify the integrity of the speech. Extensive experimental results demonstrate that the proposed SpeechVerifier is effective in detecting tampering attacks and robust to benign operations.
title SpeechVerifier: Robust Acoustic Fingerprint against Tampering Attacks via Watermarking
topic Cryptography and Security
Sound
Audio and Speech Processing
url https://arxiv.org/abs/2505.23821