Do You See What I Say? Generalizable Deepfake Detection based on Visual Speech Recognition
Fuente:
arXiv
Saved in:
| Main Authors: | Bora, Maheswar, Dhamija, Tashvik, Reddy, Shukesh, Chopin, Baptiste, Balaji, Pranav, Das, Abhijit, Dantcheva, Antitza |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
AI killed the video star. Audio-driven diffusion model for expressive talking head generation
by: Chopin, Baptiste, et al.
Published: (2025)
by: Chopin, Baptiste, et al.
Published: (2025)
Dimitra: Audio-driven Diffusion model for Expressive Talking Head Generation
by: Chopin, Baptiste, et al.
Published: (2025)
by: Chopin, Baptiste, et al.
Published: (2025)
DenVisCoM: Dense Vision Correspondence Mamba for Efficient and Real-time Optical Flow and Stereo Estimation
by: Anand, Tushar, et al.
Published: (2026)
by: Anand, Tushar, et al.
Published: (2026)
Beyond Real versus Fake Towards Intent-Aware Video Analysis
by: Atreya, Saurabh, et al.
Published: (2025)
by: Atreya, Saurabh, et al.
Published: (2025)
THEval. Evaluation Framework for Talking Head Video Generation
by: Quignon, Nabyl, et al.
Published: (2025)
by: Quignon, Nabyl, et al.
Published: (2025)
Fusion-SSAT: Unleashing the Potential of Self-supervised Auxiliary Task by Feature Fusion for Generalized Deepfake Detection
by: Reddy, Shukesh, et al.
Published: (2026)
by: Reddy, Shukesh, et al.
Published: (2026)
Now You See Me, Now You Don't: A Unified Framework for Expression Consistent Anonymization in Talking Head Videos
by: Egin, Anil, et al.
Published: (2026)
by: Egin, Anil, et al.
Published: (2026)
A Backbone Benchmarking Study on Self-supervised Learning as a Auxiliary Task with Texture-based Local Descriptors for Face Analysis
by: Reddy, Shukesh, et al.
Published: (2026)
by: Reddy, Shukesh, et al.
Published: (2026)
Investigating the Viability of Employing Multi-modal Large Language Models in the Context of Audio Deepfake Detection
by: Chuchra, Akanksha, et al.
Published: (2026)
by: Chuchra, Akanksha, et al.
Published: (2026)
Self-supervised Auxiliary Learning for Texture and Model-based Hybrid Robust and Fair Featuring in Face Analysis
by: Reddy, Shukesh, et al.
Published: (2024)
by: Reddy, Shukesh, et al.
Published: (2024)
AM Flow: Adapters for Temporal Processing in Action Recognition
by: Agrawal, Tanay, et al.
Published: (2024)
by: Agrawal, Tanay, et al.
Published: (2024)
Beyond the Visible: A Survey on Cross-spectral Face Recognition
by: Anghelone, David, et al.
Published: (2022)
by: Anghelone, David, et al.
Published: (2022)
KDC-MAE: Knowledge Distilled Contrastive Mask Auto-Encoder
by: Bora, Maheswar, et al.
Published: (2024)
by: Bora, Maheswar, et al.
Published: (2024)
Enhancing 3D-Air Signature by Pen Tip Tail Trajectory Awareness: Dataset and Featuring by Novel Spatio-temporal CNN
by: Atreya, Saurabh, et al.
Published: (2024)
by: Atreya, Saurabh, et al.
Published: (2024)
Seeing What You Say: Expressive Image Generation from Speech
by: Lee, Jiyoung, et al.
Published: (2025)
by: Lee, Jiyoung, et al.
Published: (2025)
ViM-Disparity: Bridging the Gap of Speed, Accuracy and Memory for Disparity Map Generation
by: Bora, Maheswar, et al.
Published: (2024)
by: Bora, Maheswar, et al.
Published: (2024)
HFNeRF: Learning Human Biomechanic Features with Neural Radiance Fields
by: Dey, Arnab, et al.
Published: (2024)
by: Dey, Arnab, et al.
Published: (2024)
GHNeRF: Learning Generalizable Human Features with Efficient Neural Radiance Fields
by: Dey, Arnab, et al.
Published: (2024)
by: Dey, Arnab, et al.
Published: (2024)
Do You See What I See? A Qualitative Study Eliciting High-Level Visualization Comprehension
by: Quadri, Ghulam Jilani, et al.
Published: (2024)
by: Quadri, Ghulam Jilani, et al.
Published: (2024)
Do What I Say! Voice Recognition Makes Major Advances.
by: Ruley, C. Dorsey
Published: (1994)
by: Ruley, C. Dorsey
Published: (1994)
Do You See What I See?: Observed Race and the Ascription of American Identity
by: Raul S. Casarez
Published: (2025)
by: Raul S. Casarez
Published: (2025)
Teacher: Can You See What I'm Saying? A Research Experience with Deaf Learners
by: Olga Lucía Ávila Caica
Published: (2011)
by: Olga Lucía Ávila Caica
Published: (2011)
Visual-Aware Speech Recognition for Noisy Scenarios
by: Balaji, Lakshmipathi, et al.
Published: (2025)
by: Balaji, Lakshmipathi, et al.
Published: (2025)
Do You See What I See? Leader–Follower Congruence in Authentic Leadership and Employee Participation
by: Hsing‐Er Lin, et al.
Published: (2026)
by: Hsing‐Er Lin, et al.
Published: (2026)
What You Read Isn't What You Hear: Linguistic Sensitivity in Deepfake Speech Detection
by: Nguyen, Binh, et al.
Published: (2025)
by: Nguyen, Binh, et al.
Published: (2025)
What Exactly is a Deepfake?
by: Liu, Yizhi, et al.
Published: (2025)
by: Liu, Yizhi, et al.
Published: (2025)
“If You See Something, Say Something,” for Patient Advocacy
by: M. Imelda Wright
Published: (2025)
by: M. Imelda Wright
Published: (2025)
Do You Hear What I See? Assessing Accessibility of Digital Commons and CONTENTdm
by: Walker, Wendy, et al.
Published: (2015)
by: Walker, Wendy, et al.
Published: (2015)
LIA-X: Interpretable Latent Portrait Animator
by: Wang, Yaohui, et al.
Published: (2025)
by: Wang, Yaohui, et al.
Published: (2025)
What You Say Predicts How You Do: A Multilevel Analysis of Corporate Environmental Performance
by: C. José García, et al.
Published: (2026)
by: C. José García, et al.
Published: (2026)
Vision Language Models See What You Want but not What You See
by: Gao, Qingying, et al.
Published: (2024)
by: Gao, Qingying, et al.
Published: (2024)
Saying What We Will Do, and Doing What We Say: Implementing a Customer Service Plan.
by: Wehmeyer, Susan, et al.
Published: (1996)
by: Wehmeyer, Susan, et al.
Published: (1996)
Browsing without Third-Party Cookies: What Do You See?
by: Lin, Maxwell, et al.
Published: (2024)
by: Lin, Maxwell, et al.
Published: (2024)
Do You See What I Am Pointing At? Gesture-Based Egocentric Video Question Answering
by: Choi, Yura, et al.
Published: (2026)
by: Choi, Yura, et al.
Published: (2026)
LEO: Generative Latent Image Animator for Human Video Synthesis
by: Wang, Yaohui, et al.
Published: (2023)
by: Wang, Yaohui, et al.
Published: (2023)
LAC: Latent Action Composition for Skeleton-based Action Segmentation
by: Yang, Di, et al.
Published: (2023)
by: Yang, Di, et al.
Published: (2023)
What You Think is What You See: Driving Exploration in VLM Agents via Visual-Linguistic Curiosity
by: Li, Haoxi, et al.
Published: (2026)
by: Li, Haoxi, et al.
Published: (2026)
Union Truth: Do as I Say, Not as I Do
Published: (2025)
Published: (2025)
You Know What I'm Saying: Jailbreak Attack via Implicit Reference
by: Wu, Tianyu, et al.
Published: (2024)
by: Wu, Tianyu, et al.
Published: (2024)
Euphemisms: Tell Me What You Do and I'll Tell You What You Are!
by: Perez, Cenel Augusto
Published: (2025)
by: Perez, Cenel Augusto
Published: (2025)
Similar Items
-
AI killed the video star. Audio-driven diffusion model for expressive talking head generation
by: Chopin, Baptiste, et al.
Published: (2025) -
Dimitra: Audio-driven Diffusion model for Expressive Talking Head Generation
by: Chopin, Baptiste, et al.
Published: (2025) -
DenVisCoM: Dense Vision Correspondence Mamba for Efficient and Real-time Optical Flow and Stereo Estimation
by: Anand, Tushar, et al.
Published: (2026) -
Beyond Real versus Fake Towards Intent-Aware Video Analysis
by: Atreya, Saurabh, et al.
Published: (2025) -
THEval. Evaluation Framework for Talking Head Video Generation
by: Quignon, Nabyl, et al.
Published: (2025)