KLASSify to Verify: Audio-Visual Deepfake Detection Using SSL-based Audio and Handcrafted Visual Features
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Kukanov, Ivan, Ng, Jun Wah |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
AVFF: Audio-Visual Feature Fusion for Video Deepfake Detection
von: Oorloff, Trevine, et al.
Veröffentlicht: (2024)
von: Oorloff, Trevine, et al.
Veröffentlicht: (2024)
Detecting Audio-Visual Deepfakes with Fine-Grained Inconsistencies
von: Astrid, Marcella, et al.
Veröffentlicht: (2024)
von: Astrid, Marcella, et al.
Veröffentlicht: (2024)
Pindrop it! Audio and Visual Deepfake Countermeasures for Robust Detection and Fine Grained-Localization
von: Klein, Nicholas, et al.
Veröffentlicht: (2025)
von: Klein, Nicholas, et al.
Veröffentlicht: (2025)
Audio-Visual Deepfake Detection With Local Temporal Inconsistencies
von: Astrid, Marcella, et al.
Veröffentlicht: (2025)
von: Astrid, Marcella, et al.
Veröffentlicht: (2025)
Straight Through Gumbel Softmax Estimator based Bimodal Neural Architecture Search for Audio-Visual Deepfake Detection
von: PN, Aravinda Reddy, et al.
Veröffentlicht: (2024)
von: PN, Aravinda Reddy, et al.
Veröffentlicht: (2024)
Bootstrapping Audio-Visual Segmentation by Strengthening Audio Cues
von: Chen, Tianxiang, et al.
Veröffentlicht: (2024)
von: Chen, Tianxiang, et al.
Veröffentlicht: (2024)
Localizing Audio-Visual Deepfakes via Hierarchical Boundary Modeling
von: Chen, Xuanjun, et al.
Veröffentlicht: (2025)
von: Chen, Xuanjun, et al.
Veröffentlicht: (2025)
Whisper-Flamingo: Integrating Visual Features into Whisper for Audio-Visual Speech Recognition and Translation
von: Rouditchenko, Andrew, et al.
Veröffentlicht: (2024)
von: Rouditchenko, Andrew, et al.
Veröffentlicht: (2024)
Contextual Cross-Modal Attention for Audio-Visual Deepfake Detection and Localization
von: Katamneni, Vinaya Sree, et al.
Veröffentlicht: (2024)
von: Katamneni, Vinaya Sree, et al.
Veröffentlicht: (2024)
Learning Spatial Features from Audio-Visual Correspondence in Egocentric Videos
von: Majumder, Sagnik, et al.
Veröffentlicht: (2023)
von: Majumder, Sagnik, et al.
Veröffentlicht: (2023)
CCStereo: Audio-Visual Contextual and Contrastive Learning for Binaural Audio Generation
von: Chen, Yuanhong, et al.
Veröffentlicht: (2025)
von: Chen, Yuanhong, et al.
Veröffentlicht: (2025)
Dynamic Derivation and Elimination: Audio Visual Segmentation with Enhanced Audio Semantics
von: Liu, Chen, et al.
Veröffentlicht: (2025)
von: Liu, Chen, et al.
Veröffentlicht: (2025)
DDAVS: Disentangled Audio Semantics and Delayed Bidirectional Alignment for Audio-Visual Segmentation
von: Tian, Jingqi, et al.
Veröffentlicht: (2025)
von: Tian, Jingqi, et al.
Veröffentlicht: (2025)
Lightweight Wasserstein Audio-Visual Model for Unified Speech Enhancement and Separation
von: Park, Jisoo, et al.
Veröffentlicht: (2025)
von: Park, Jisoo, et al.
Veröffentlicht: (2025)
Benchmarking Cross-Domain Audio-Visual Deception Detection
von: Guo, Xiaobao, et al.
Veröffentlicht: (2024)
von: Guo, Xiaobao, et al.
Veröffentlicht: (2024)
Exploring Green AI for Audio Deepfake Detection
von: Saha, Subhajit, et al.
Veröffentlicht: (2024)
von: Saha, Subhajit, et al.
Veröffentlicht: (2024)
Classifying Shelf Life Quality of Pineapples by Combining Audio and Visual Features
von: Jiang, Yi-Lu, et al.
Veröffentlicht: (2025)
von: Jiang, Yi-Lu, et al.
Veröffentlicht: (2025)
AV2AV: Direct Audio-Visual Speech to Audio-Visual Speech Translation with Unified Audio-Visual Speech Representation
von: Choi, Jeongsoo, et al.
Veröffentlicht: (2023)
von: Choi, Jeongsoo, et al.
Veröffentlicht: (2023)
Audio-Visual Person Verification based on Recursive Fusion of Joint Cross-Attention
von: Praveen, R. Gnana, et al.
Veröffentlicht: (2024)
von: Praveen, R. Gnana, et al.
Veröffentlicht: (2024)
3D Audio-Visual Segmentation
von: Sokolov, Artem, et al.
Veröffentlicht: (2024)
von: Sokolov, Artem, et al.
Veröffentlicht: (2024)
Language-Guided Contrastive Audio-Visual Masked Autoencoder with Automatically Generated Audio-Visual-Text Triplets from Videos
von: Ishikawa, Yuchi, et al.
Veröffentlicht: (2025)
von: Ishikawa, Yuchi, et al.
Veröffentlicht: (2025)
Statistics-aware Audio-visual Deepfake Detector
von: Astrid, Marcella, et al.
Veröffentlicht: (2024)
von: Astrid, Marcella, et al.
Veröffentlicht: (2024)
Seeing Speech and Sound: Distinguishing and Locating Audios in Visual Scenes
von: Ryu, Hyeonggon, et al.
Veröffentlicht: (2025)
von: Ryu, Hyeonggon, et al.
Veröffentlicht: (2025)
UniSync: A Unified Framework for Audio-Visual Synchronization
von: Feng, Tao, et al.
Veröffentlicht: (2025)
von: Feng, Tao, et al.
Veröffentlicht: (2025)
Look, Listen and Recognise: Character-Aware Audio-Visual Subtitling
von: Korbar, Bruno, et al.
Veröffentlicht: (2024)
von: Korbar, Bruno, et al.
Veröffentlicht: (2024)
Audio-Visual Talker Localization in Video for Spatial Sound Reproduction
von: Berghi, Davide, et al.
Veröffentlicht: (2024)
von: Berghi, Davide, et al.
Veröffentlicht: (2024)
SAM Audio: Segment Anything in Audio
von: Shi, Bowen, et al.
Veröffentlicht: (2025)
von: Shi, Bowen, et al.
Veröffentlicht: (2025)
Seeing Soundscapes: Audio-Visual Generation and Separation from Soundscapes Using Audio-Visual Separator
von: Kang, Minjae, et al.
Veröffentlicht: (2025)
von: Kang, Minjae, et al.
Veröffentlicht: (2025)
UWAV: Uncertainty-weighted Weakly-supervised Audio-Visual Video Parsing
von: Lai, Yung-Hsuan, et al.
Veröffentlicht: (2025)
von: Lai, Yung-Hsuan, et al.
Veröffentlicht: (2025)
MoME: Mixture of Matryoshka Experts for Audio-Visual Speech Recognition
von: Cappellazzo, Umberto, et al.
Veröffentlicht: (2025)
von: Cappellazzo, Umberto, et al.
Veröffentlicht: (2025)
EchoInk-R1: Exploring Audio-Visual Reasoning in Multimodal LLMs via Reinforcement Learning
von: Xing, Zhenghao, et al.
Veröffentlicht: (2025)
von: Xing, Zhenghao, et al.
Veröffentlicht: (2025)
AlignVSR: Audio-Visual Cross-Modal Alignment for Visual Speech Recognition
von: Liu, Zehua, et al.
Veröffentlicht: (2024)
von: Liu, Zehua, et al.
Veröffentlicht: (2024)
Mixture of Low-Rank Adapter Experts in Generalizable Audio Deepfake Detection
von: Laakkonen, Janne, et al.
Veröffentlicht: (2025)
von: Laakkonen, Janne, et al.
Veröffentlicht: (2025)
Can You Hear, Localize, and Segment Continually? An Exemplar-Free Continual Learning Benchmark for Audio-Visual Segmentation
von: Raghavan, Siddeshwar, et al.
Veröffentlicht: (2026)
von: Raghavan, Siddeshwar, et al.
Veröffentlicht: (2026)
Mitigating Audiovisual Mismatch in Visual-Guide Audio Captioning
von: Xu, Le, et al.
Veröffentlicht: (2025)
von: Xu, Le, et al.
Veröffentlicht: (2025)
mWhisper-Flamingo for Multilingual Audio-Visual Noise-Robust Speech Recognition
von: Rouditchenko, Andrew, et al.
Veröffentlicht: (2025)
von: Rouditchenko, Andrew, et al.
Veröffentlicht: (2025)
SSAVSV: Towards Unified Model for Self-Supervised Audio-Visual Speaker Verification
von: Rajasekhar, Gnana Praveen, et al.
Veröffentlicht: (2025)
von: Rajasekhar, Gnana Praveen, et al.
Veröffentlicht: (2025)
Mitigating Attention Sinks and Massive Activations in Audio-Visual Speech Recognition with LLMs
von: Anand, et al.
Veröffentlicht: (2025)
von: Anand, et al.
Veröffentlicht: (2025)
PEAVS: Perceptual Evaluation of Audio-Visual Synchrony Grounded in Viewers' Opinion Scores
von: Goncalves, Lucas, et al.
Veröffentlicht: (2024)
von: Goncalves, Lucas, et al.
Veröffentlicht: (2024)
Listening without Looking: Modality Bias in Audio-Visual Captioning
von: Ishikawa, Yuchi, et al.
Veröffentlicht: (2025)
von: Ishikawa, Yuchi, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
AVFF: Audio-Visual Feature Fusion for Video Deepfake Detection
von: Oorloff, Trevine, et al.
Veröffentlicht: (2024) -
Detecting Audio-Visual Deepfakes with Fine-Grained Inconsistencies
von: Astrid, Marcella, et al.
Veröffentlicht: (2024) -
Pindrop it! Audio and Visual Deepfake Countermeasures for Robust Detection and Fine Grained-Localization
von: Klein, Nicholas, et al.
Veröffentlicht: (2025) -
Audio-Visual Deepfake Detection With Local Temporal Inconsistencies
von: Astrid, Marcella, et al.
Veröffentlicht: (2025) -
Straight Through Gumbel Softmax Estimator based Bimodal Neural Architecture Search for Audio-Visual Deepfake Detection
von: PN, Aravinda Reddy, et al.
Veröffentlicht: (2024)