Enhancing Self-Supervised Talking Head Forgery Detection via a Training-Free Dual-System Framework
Fuente:
arXiv
Salvato in:
| Autori principali: | Liu, Ke, Wei, Jiwei, Zhou, Shuchang, Xiao, Yutong, Chai, Ruikun, Qin, Yitong, Zhou, Yuyang, Yang, Yang |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
From Talking to Singing: A New Challenge for Audio-Visual Deepfake Detection
di: Liu, Ke, et al.
Pubblicazione: (2026)
di: Liu, Ke, et al.
Pubblicazione: (2026)
Identity-Aware Vision-Language Model for Explainable Face Forgery Detection
di: Xu, Junhao, et al.
Pubblicazione: (2025)
di: Xu, Junhao, et al.
Pubblicazione: (2025)
Band-Attention Modulated RetNet for Face Forgery Detection
di: Zhang, Zhida, et al.
Pubblicazione: (2024)
di: Zhang, Zhida, et al.
Pubblicazione: (2024)
Identity-Driven Multimedia Forgery Detection via Reference Assistance
di: Xu, Junhao, et al.
Pubblicazione: (2024)
di: Xu, Junhao, et al.
Pubblicazione: (2024)
EmotionTalk: An Interactive Chinese Multimodal Emotion Dataset With Rich Annotations
di: Sun, Haoqin, et al.
Pubblicazione: (2025)
di: Sun, Haoqin, et al.
Pubblicazione: (2025)
Optimization-Free Universal Watermark Forgery with Regenerative Diffusion Models
di: Zhu, Chaoyi, et al.
Pubblicazione: (2025)
di: Zhu, Chaoyi, et al.
Pubblicazione: (2025)
UniTalking: A Unified Audio-Video Framework for Talking Portrait Generation
di: Li, Hebeizi, et al.
Pubblicazione: (2026)
di: Li, Hebeizi, et al.
Pubblicazione: (2026)
CEM-Net: Cross-Emotion Memory Network for Emotional Talking Face Generation
di: Wu, Kangyi, et al.
Pubblicazione: (2025)
di: Wu, Kangyi, et al.
Pubblicazione: (2025)
MG-Former: A Transformer-Based Framework for Music-Driven 3D Conducting Gesture Generation
di: Qiu, Ke, et al.
Pubblicazione: (2026)
di: Qiu, Ke, et al.
Pubblicazione: (2026)
Pianist Transformer: Towards Expressive Piano Performance Rendering via Scalable Self-Supervised Pre-Training
di: You, Hong-Jie, et al.
Pubblicazione: (2025)
di: You, Hong-Jie, et al.
Pubblicazione: (2025)
DreamHead: Learning Spatial-Temporal Correspondence via Hierarchical Diffusion for Audio-driven Talking Head Synthesis
di: Hong, Fa-Ting, et al.
Pubblicazione: (2024)
di: Hong, Fa-Ting, et al.
Pubblicazione: (2024)
FD2Talk: Towards Generalized Talking Head Generation with Facial Decoupled Diffusion Model
di: Yao, Ziyu, et al.
Pubblicazione: (2024)
di: Yao, Ziyu, et al.
Pubblicazione: (2024)
Structure-Aware Residual-Center Representation for Self-Supervised Open-Set 3D Cross-Modal Retrieval
di: Xu, Yang, et al.
Pubblicazione: (2024)
di: Xu, Yang, et al.
Pubblicazione: (2024)
Panonut360: A Head and Eye Tracking Dataset for Panoramic Video
di: Xu, Yutong, et al.
Pubblicazione: (2024)
di: Xu, Yutong, et al.
Pubblicazione: (2024)
Inclusion 2024 Global Multimedia Deepfake Detection Challenge: Towards Multi-dimensional Face Forgery Detection
di: Zhang, Yi, et al.
Pubblicazione: (2024)
di: Zhang, Yi, et al.
Pubblicazione: (2024)
Test-Time Self-Adaptive Conditioning for Stable Audio-Driven Talking-Head Generation
di: Zhang, Zhicheng, et al.
Pubblicazione: (2026)
di: Zhang, Zhicheng, et al.
Pubblicazione: (2026)
DDNet: A Dual-Stream Graph Learning and Disentanglement Framework for Temporal Forgery Localization
di: Zhao, Boyang, et al.
Pubblicazione: (2026)
di: Zhao, Boyang, et al.
Pubblicazione: (2026)
CustomDancer: Customized Dance Recommendation by Text-Dance Retrieval
di: Qin, Yawen, et al.
Pubblicazione: (2026)
di: Qin, Yawen, et al.
Pubblicazione: (2026)
An Inverse Partial Optimal Transport Framework for Music-guided Movie Trailer Generation
di: Wang, Yutong, et al.
Pubblicazione: (2024)
di: Wang, Yutong, et al.
Pubblicazione: (2024)
Copy-Move Forgery Detection and Question Answering for Remote Sensing Image
di: Zhang, Ze, et al.
Pubblicazione: (2024)
di: Zhang, Ze, et al.
Pubblicazione: (2024)
Look, Listen and Segment: Towards Weakly Supervised Audio-visual Semantic Segmentation
di: Li, Chengzhi, et al.
Pubblicazione: (2026)
di: Li, Chengzhi, et al.
Pubblicazione: (2026)
DAT: Dual-Aware Adaptive Transmission for Efficient Multimodal LLM Inference in Edge-Cloud Systems
di: Guo, Qi, et al.
Pubblicazione: (2026)
di: Guo, Qi, et al.
Pubblicazione: (2026)
SIDQL: An Efficient Keyframe Extraction and Motion Reconstruction Framework in Motion Capture
di: Zhang, Xuling, et al.
Pubblicazione: (2024)
di: Zhang, Xuling, et al.
Pubblicazione: (2024)
Challenging Dataset and Multi-modal Gated Mixture of Experts Model for Remote Sensing Copy-Move Forgery Understanding
di: Zhang, Ze, et al.
Pubblicazione: (2025)
di: Zhang, Ze, et al.
Pubblicazione: (2025)
MoLEx: Mixture of LoRA Experts in Speech Self-Supervised Models for Audio Deepfake Detection
di: Pan, Zihan, et al.
Pubblicazione: (2025)
di: Pan, Zihan, et al.
Pubblicazione: (2025)
RDTF: Resource-efficient Dual-mask Training Framework for Multi-frame Animated Sticker Generation
di: Yuan, Zhiqiang, et al.
Pubblicazione: (2025)
di: Yuan, Zhiqiang, et al.
Pubblicazione: (2025)
TalkVerse: Democratizing Minute-Long Audio-Driven Video Generation
di: Wang, Zhenzhi, et al.
Pubblicazione: (2025)
di: Wang, Zhenzhi, et al.
Pubblicazione: (2025)
PoseTalk: Text-and-Audio-based Pose Control and Motion Refinement for One-Shot Talking Head Generation
di: Ling, Jun, et al.
Pubblicazione: (2024)
di: Ling, Jun, et al.
Pubblicazione: (2024)
Generalized Face Forgery Detection via Adaptive Learning for Pre-trained Vision Transformer
di: Luo, Anwei, et al.
Pubblicazione: (2023)
di: Luo, Anwei, et al.
Pubblicazione: (2023)
Listening Deepfake Detection: A New Perspective Beyond Speaking-Centric Forgery Analysis
di: Liu, Miao, et al.
Pubblicazione: (2026)
di: Liu, Miao, et al.
Pubblicazione: (2026)
GaussianTalker: Speaker-specific Talking Head Synthesis via 3D Gaussian Splatting
di: Yu, Hongyun, et al.
Pubblicazione: (2024)
di: Yu, Hongyun, et al.
Pubblicazione: (2024)
InconVAD: A Two-Stage Dual-Tower Framework for Multimodal Emotion Inconsistency Detection
di: Li, Zongyi, et al.
Pubblicazione: (2025)
di: Li, Zongyi, et al.
Pubblicazione: (2025)
MarkIt: Training-Free Visual Markers for Precise Video Temporal Grounding
di: Fang, Pengcheng, et al.
Pubblicazione: (2026)
di: Fang, Pengcheng, et al.
Pubblicazione: (2026)
PAME: Self-Supervised Masked Autoencoder for No-Reference Point Cloud Quality Assessment
di: Shan, Ziyu, et al.
Pubblicazione: (2024)
di: Shan, Ziyu, et al.
Pubblicazione: (2024)
Generalizable Deepfake Detection Based on Forgery-aware Layer Masking and Multi-artifact Subspace Decomposition
di: Zhang, Xiang, et al.
Pubblicazione: (2026)
di: Zhang, Xiang, et al.
Pubblicazione: (2026)
Think2Sing: Orchestrating Structured Motion Subtitles for Singing-Driven 3D Head Animation
di: Huang, Zikai, et al.
Pubblicazione: (2025)
di: Huang, Zikai, et al.
Pubblicazione: (2025)
MRATTS: An MR-Based Acupoint Therapy Training System with Real-Time Acupoint Detection and Evaluation Standards
di: Liu, Jiacheng, et al.
Pubblicazione: (2026)
di: Liu, Jiacheng, et al.
Pubblicazione: (2026)
Multimodal Interaction Modeling via Self-Supervised Multi-Task Learning for Review Helpfulness Prediction
di: Gong, HongLin, et al.
Pubblicazione: (2024)
di: Gong, HongLin, et al.
Pubblicazione: (2024)
Towards Automatic Soccer Commentary Generation with Knowledge-Enhanced Visual Reasoning
di: Jin, Zeyu, et al.
Pubblicazione: (2026)
di: Jin, Zeyu, et al.
Pubblicazione: (2026)
Dual Knowledge-Enhanced Two-Stage Reasoner for Multimodal Dialog Systems
di: Chen, Xiaolin, et al.
Pubblicazione: (2025)
di: Chen, Xiaolin, et al.
Pubblicazione: (2025)
Documenti analoghi
-
From Talking to Singing: A New Challenge for Audio-Visual Deepfake Detection
di: Liu, Ke, et al.
Pubblicazione: (2026) -
Identity-Aware Vision-Language Model for Explainable Face Forgery Detection
di: Xu, Junhao, et al.
Pubblicazione: (2025) -
Band-Attention Modulated RetNet for Face Forgery Detection
di: Zhang, Zhida, et al.
Pubblicazione: (2024) -
Identity-Driven Multimedia Forgery Detection via Reference Assistance
di: Xu, Junhao, et al.
Pubblicazione: (2024) -
EmotionTalk: An Interactive Chinese Multimodal Emotion Dataset With Rich Annotations
di: Sun, Haoqin, et al.
Pubblicazione: (2025)