SpeechForensics: Audio-Visual Speech Representation Learning for Face Forgery Detection
Fuente:
arXiv
Saved in:
| Main Authors: | Liang, Yachao, Yu, Min, Li, Gang, Jiang, Jianguo, Li, Boquan, Yu, Feng, Zhang, Ning, Meng, Xiang, Huang, Weiqing |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
UniForensics: Face Forgery Detection via General Facial Representation
by: Fang, Ziyuan, et al.
Published: (2024)
by: Fang, Ziyuan, et al.
Published: (2024)
Leveraging Mamba with Full-Face Vision for Audio-Visual Speech Enhancement
by: Chao, Rong, et al.
Published: (2025)
by: Chao, Rong, et al.
Published: (2025)
LLM-Guided Reinforcement Learning for Audio-Visual Speech Enhancement
by: Chen, Chih-Ning, et al.
Published: (2026)
by: Chen, Chih-Ning, et al.
Published: (2026)
Forensics Adapter: Unleashing CLIP for Generalizable Face Forgery Detection
by: Cui, Xinjie, et al.
Published: (2024)
by: Cui, Xinjie, et al.
Published: (2024)
Distilled Transformers with Locally Enhanced Global Representations for Face Forgery Detection
by: Zhang, Yaning, et al.
Published: (2024)
by: Zhang, Yaning, et al.
Published: (2024)
Learning Natural Consistency Representation for Face Forgery Video Detection
by: Zhang, Daichi, et al.
Published: (2024)
by: Zhang, Daichi, et al.
Published: (2024)
Learning to Discover Forgery Cues for Face Forgery Detection
by: Tian, Jiahe, et al.
Published: (2024)
by: Tian, Jiahe, et al.
Published: (2024)
AV2AV: Direct Audio-Visual Speech to Audio-Visual Speech Translation with Unified Audio-Visual Speech Representation
by: Choi, Jeongsoo, et al.
Published: (2023)
by: Choi, Jeongsoo, et al.
Published: (2023)
Identity-Aware Vision-Language Model for Explainable Face Forgery Detection
by: Xu, Junhao, et al.
Published: (2025)
by: Xu, Junhao, et al.
Published: (2025)
XLAVS-R: Cross-Lingual Audio-Visual Speech Representation Learning for Noise-Robust Speech Perception
by: Han, HyoJung, et al.
Published: (2024)
by: Han, HyoJung, et al.
Published: (2024)
Conversational Speech Recognition by Learning Audio-textual Cross-modal Contextual Representation
by: Wei, Kun, et al.
Published: (2023)
by: Wei, Kun, et al.
Published: (2023)
Poisoned Forgery Face: Towards Backdoor Attacks on Face Forgery Detection
by: Liang, Jiawei, et al.
Published: (2024)
by: Liang, Jiawei, et al.
Published: (2024)
Revisiting Face Forgery Detection: From Facial Representation to Forgery Detection
by: Guo, Zonghui, et al.
Published: (2024)
by: Guo, Zonghui, et al.
Published: (2024)
Benchmarking Joint Face Spoofing and Forgery Detection with Visual and Physiological Cues
by: Yu, Zitong, et al.
Published: (2022)
by: Yu, Zitong, et al.
Published: (2022)
Federated Face Forgery Detection Learning with Personalized Representation
by: Liu, Decheng, et al.
Published: (2024)
by: Liu, Decheng, et al.
Published: (2024)
DomainForensics: Exposing Face Forgery across Domains via Bi-directional Adaptation
by: Lv, Qingxuan, et al.
Published: (2023)
by: Lv, Qingxuan, et al.
Published: (2023)
Audio-Visual Speech Representation Expert for Enhanced Talking Face Video Generation and Evaluation
by: Yaman, Dogucan, et al.
Published: (2024)
by: Yaman, Dogucan, et al.
Published: (2024)
GLIMMER: Incorporating Graph and Lexical Features in Unsupervised Multi-Document Summarization
by: Liu, Ran, et al.
Published: (2024)
by: Liu, Ran, et al.
Published: (2024)
AV-data2vec: Self-supervised Learning of Audio-Visual Speech Representations with Contextualized Target Representations
by: Lian, Jiachen, et al.
Published: (2023)
by: Lian, Jiachen, et al.
Published: (2023)
ForgeryVCR: Visual-Centric Reasoning via Efficient Forensic Tools in MLLMs for Image Forgery Detection and Localization
by: Wang, Youqi, et al.
Published: (2026)
by: Wang, Youqi, et al.
Published: (2026)
Agent4FaceForgery: Multi-Agent LLM Framework for Realistic Face Forgery Detection
by: Lai, Yingxin, et al.
Published: (2025)
by: Lai, Yingxin, et al.
Published: (2025)
Deep Learning Technology for Face Forgery Detection: A Survey
by: Ma, Lixia, et al.
Published: (2024)
by: Ma, Lixia, et al.
Published: (2024)
STAR: Speech-to-Audio Generation via Representation Learning
by: Xie, Zeyu, et al.
Published: (2025)
by: Xie, Zeyu, et al.
Published: (2025)
Code-in-the-Loop Forensics: Agentic Tool Use for Image Forgery Detection
by: Zhang, Fanrui, et al.
Published: (2025)
by: Zhang, Fanrui, et al.
Published: (2025)
CorrDetail: Visual Detail Enhanced Self-Correction for Face Forgery Detection
by: Zhou, Binjia, et al.
Published: (2025)
by: Zhou, Binjia, et al.
Published: (2025)
Real-Time Audio-Visual Speech Enhancement Using Pre-trained Visual Representations
by: Ma, T. Aleksandra, et al.
Published: (2025)
by: Ma, T. Aleksandra, et al.
Published: (2025)
Generalizable Face Forgery Detection via Separable Prompt Learning
by: Yang, Enrui, et al.
Published: (2026)
by: Yang, Enrui, et al.
Published: (2026)
Towards General Visual-Linguistic Face Forgery Detection
by: Sun, Ke, et al.
Published: (2023)
by: Sun, Ke, et al.
Published: (2023)
Contrastive Desensitization Learning for Cross Domain Face Forgery Detection
by: Qiu, Lingyu, et al.
Published: (2025)
by: Qiu, Lingyu, et al.
Published: (2025)
Human-Inspired Computing for Robust and Efficient Audio-Visual Speech Recognition
by: Liu, Qianhui, et al.
Published: (2024)
by: Liu, Qianhui, et al.
Published: (2024)
Zero-AVSR: Zero-Shot Audio-Visual Speech Recognition with LLMs by Learning Language-Agnostic Speech Representations
by: Yeo, Jeong Hun, et al.
Published: (2025)
by: Yeo, Jeong Hun, et al.
Published: (2025)
LAKAN: Landmark-assisted Adaptive Kolmogorov-Arnold Network for Face Forgery Detection
by: Jiang, Jiayao, et al.
Published: (2025)
by: Jiang, Jiayao, et al.
Published: (2025)
Multi-Task Corrupted Prediction for Learning Robust Audio-Visual Speech Representation
by: Kim, Sungnyun, et al.
Published: (2025)
by: Kim, Sungnyun, et al.
Published: (2025)
Contrastive Learning With Audio Discrimination For Customizable Keyword Spotting In Continuous Speech
by: Xi, Yu, et al.
Published: (2024)
by: Xi, Yu, et al.
Published: (2024)
ELEGANCE: Efficient LLM Guidance for Audio-Visual Target Speech Extraction
by: Wu, Wenxuan, et al.
Published: (2025)
by: Wu, Wenxuan, et al.
Published: (2025)
Speech-to-See: End-to-End Speech-Driven Open-Set Object Detection
by: Lu, Wenhuan, et al.
Published: (2025)
by: Lu, Wenhuan, et al.
Published: (2025)
Detecting Adversarial Data using Perturbation Forgery
by: Wang, Qian, et al.
Published: (2024)
by: Wang, Qian, et al.
Published: (2024)
DCIM-AVSR : Efficient Audio-Visual Speech Recognition via Dual Conformer Interaction Module
by: Wang, Xinyu, et al.
Published: (2024)
by: Wang, Xinyu, et al.
Published: (2024)
Robust Audio-Visual Speech Enhancement: Correcting Misassignments in Complex Environments with Advanced Post-Processing
by: Ren, Wenze, et al.
Published: (2024)
by: Ren, Wenze, et al.
Published: (2024)
RepCodec: A Speech Representation Codec for Speech Tokenization
by: Huang, Zhichao, et al.
Published: (2023)
by: Huang, Zhichao, et al.
Published: (2023)
Similar Items
-
UniForensics: Face Forgery Detection via General Facial Representation
by: Fang, Ziyuan, et al.
Published: (2024) -
Leveraging Mamba with Full-Face Vision for Audio-Visual Speech Enhancement
by: Chao, Rong, et al.
Published: (2025) -
LLM-Guided Reinforcement Learning for Audio-Visual Speech Enhancement
by: Chen, Chih-Ning, et al.
Published: (2026) -
Forensics Adapter: Unleashing CLIP for Generalizable Face Forgery Detection
by: Cui, Xinjie, et al.
Published: (2024) -
Distilled Transformers with Locally Enhanced Global Representations for Face Forgery Detection
by: Zhang, Yaning, et al.
Published: (2024)