Towards multi-modal forgery representation learning for AI-generated video detection and localization
Fuente:
arXiv
Saved in:
| Main Authors: | Le, Dat, Nguyen, Khoa, Wang, Xin, Hu, Shu |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Learning to mask: Towards generalized face forgery detection
by: Fei, Jianwei, et al.
Published: (2022)
by: Fei, Jianwei, et al.
Published: (2022)
PhotoHolmes: a Python library for forgery detection in digital images
by: O'Flaherty, Julián, et al.
Published: (2024)
by: O'Flaherty, Julián, et al.
Published: (2024)
Deep video representation learning: a survey
by: Ravanbakhsh, Elham, et al.
Published: (2024)
by: Ravanbakhsh, Elham, et al.
Published: (2024)
Multi-view Action Recognition via Directed Gromov-Wasserstein Discrepancy
by: Nguyen, Hoang-Quan, et al.
Published: (2024)
by: Nguyen, Hoang-Quan, et al.
Published: (2024)
Cross-view Action Recognition Understanding From Exocentric to Egocentric Perspective
by: Truong, Thanh-Dat, et al.
Published: (2023)
by: Truong, Thanh-Dat, et al.
Published: (2023)
A multi-center analysis of deep learning methods for video polyp detection and segmentation
by: Ghatwary, Noha, et al.
Published: (2026)
by: Ghatwary, Noha, et al.
Published: (2026)
Insect-Foundation: A Foundation Model and Large-scale 1M Dataset for Visual Insect Understanding
by: Nguyen, Hoang-Quan, et al.
Published: (2023)
by: Nguyen, Hoang-Quan, et al.
Published: (2023)
Insect-Foundation: A Foundation Model and Large Multimodal Dataset for Vision-Language Insect Understanding
by: Truong, Thanh-Dat, et al.
Published: (2025)
by: Truong, Thanh-Dat, et al.
Published: (2025)
ED-SAM: An Efficient Diffusion Sampling Approach to Domain Generalization in Vision-Language Foundation Models
by: Truong, Thanh-Dat, et al.
Published: (2024)
by: Truong, Thanh-Dat, et al.
Published: (2024)
VirDA: Reusing Backbone for Unsupervised Domain Adaptation with Visual Reprogramming
by: Nguyen, Duy, et al.
Published: (2025)
by: Nguyen, Duy, et al.
Published: (2025)
Understanding normalization in contrastive representation learning and out-of-distribution detection
by: Le-Gia, Tai, et al.
Published: (2023)
by: Le-Gia, Tai, et al.
Published: (2023)
Multi-modal, multi-scale representation learning for satellite imagery analysis just needs a good ALiBi
by: Kage, Patrick, et al.
Published: (2026)
by: Kage, Patrick, et al.
Published: (2026)
Timealign: A multi-modal object detection method for time misalignment fusing in autonomous driving
by: Song, Zhihang, et al.
Published: (2024)
by: Song, Zhihang, et al.
Published: (2024)
BRAIN: Bias-Mitigation Continual Learning Approach to Vision-Brain Understanding
by: Nguyen, Xuan-Bac, et al.
Published: (2025)
by: Nguyen, Xuan-Bac, et al.
Published: (2025)
VideoGen-of-Thought: Step-by-step generating multi-shot video with minimal manual intervention
by: Zheng, Mingzhe, et al.
Published: (2025)
by: Zheng, Mingzhe, et al.
Published: (2025)
MANGO: Multimodal Attention-based Normalizing Flow Approach to Fusion Learning
by: Truong, Thanh-Dat, et al.
Published: (2025)
by: Truong, Thanh-Dat, et al.
Published: (2025)
Domain Generalization through Spatial Relation Induction over Visual Primitives
by: Nguyen, Dat, et al.
Published: (2026)
by: Nguyen, Dat, et al.
Published: (2026)
A comparison of extended object tracking with multi-modal sensors in indoor environment
by: Shuai, Jiangtao, et al.
Published: (2024)
by: Shuai, Jiangtao, et al.
Published: (2024)
Venomancer: Towards Imperceptible and Target-on-Demand Backdoor Attacks in Federated Learning
by: Nguyen, Son, et al.
Published: (2024)
by: Nguyen, Son, et al.
Published: (2024)
FALCONEye: Finding Answers and Localizing Content in ONE-hour-long videos with multi-modal LLMs
by: Plou, Carlos, et al.
Published: (2025)
by: Plou, Carlos, et al.
Published: (2025)
Amodal Instance Segmentation with Diffusion Shape Prior Estimation
by: Tran, Minh, et al.
Published: (2024)
by: Tran, Minh, et al.
Published: (2024)
Towards Robust and Fair Vision Learning in Open-World Environments
by: Truong, Thanh-Dat
Published: (2024)
by: Truong, Thanh-Dat
Published: (2024)
Phantom: Subject-consistent video generation via cross-modal alignment
by: Liu, Lijie, et al.
Published: (2025)
by: Liu, Lijie, et al.
Published: (2025)
Synthetic images aid the recognition of human-made art forgeries
by: Ostmeyer, Johann, et al.
Published: (2023)
by: Ostmeyer, Johann, et al.
Published: (2023)
MV-MR: multi-views and multi-representations for self-supervised learning and knowledge distillation
by: Kinakh, Vitaliy, et al.
Published: (2023)
by: Kinakh, Vitaliy, et al.
Published: (2023)
BIMA: Bijective Maximum Likelihood Learning Approach to Hallucination Prediction and Mitigation in Large Vision-Language Models
by: Tran, Huu-Thien, et al.
Published: (2025)
by: Tran, Huu-Thien, et al.
Published: (2025)
Adaptive thresholding pattern for fingerprint forgery detection
by: Farzadpour, Zahra, et al.
Published: (2025)
by: Farzadpour, Zahra, et al.
Published: (2025)
FALCON: Fairness Learning via Contrastive Attention Approach to Continual Semantic Scene Understanding
by: Truong, Thanh-Dat, et al.
Published: (2023)
by: Truong, Thanh-Dat, et al.
Published: (2023)
OmViD: Omni-supervised active learning for video action detection
by: Rana, Aayush, et al.
Published: (2025)
by: Rana, Aayush, et al.
Published: (2025)
SHREC 2025: Retrieval of Optimal Objects for Multi-modal Enhanced Language and Spatial Assistance (ROOMELSA)
by: Nguyen, Trong-Thuan, et al.
Published: (2025)
by: Nguyen, Trong-Thuan, et al.
Published: (2025)
A multi-modal vision-language model for generalizable annotation-free pathology localization
by: Yang, Hao, et al.
Published: (2024)
by: Yang, Hao, et al.
Published: (2024)
Cross-modal ultra-scale learning with tri-modalities of renal biopsy images for glomerular multi-disease auxiliary diagnosis
by: Long, Kaixing, et al.
Published: (2025)
by: Long, Kaixing, et al.
Published: (2025)
A re-calibration method for object detection with multi-modal alignment bias in autonomous driving
by: Song, Zhihang, et al.
Published: (2024)
by: Song, Zhihang, et al.
Published: (2024)
HIG: Hierarchical Interlacement Graph Approach to Scene Graph Generation in Video Understanding
by: Nguyen, Trong-Thuan, et al.
Published: (2023)
by: Nguyen, Trong-Thuan, et al.
Published: (2023)
ShapeFormer: Shape Prior Visible-to-Amodal Transformer-based Amodal Instance Segmentation
by: Tran, Minh, et al.
Published: (2024)
by: Tran, Minh, et al.
Published: (2024)
AI killed the video star. Audio-driven diffusion model for expressive talking head generation
by: Chopin, Baptiste, et al.
Published: (2025)
by: Chopin, Baptiste, et al.
Published: (2025)
VideoGen-of-Thought: Step-by-step generating multi-shot video with minimal manual intervention
by: Zheng, Mingzhe, et al.
Published: (2024)
by: Zheng, Mingzhe, et al.
Published: (2024)
CONDA: Continual Unsupervised Domain Adaptation Learning in Visual Perception for Self-Driving Cars
by: Truong, Thanh-Dat, et al.
Published: (2022)
by: Truong, Thanh-Dat, et al.
Published: (2022)
AI-Face: A Million-Scale Demographically Annotated AI-Generated Face Dataset and Fairness Benchmark
by: Lin, Li, et al.
Published: (2024)
by: Lin, Li, et al.
Published: (2024)
FrameDiT: Diffusion Transformer with Matrix Attention for Efficient Video Generation
by: Le, Minh Khoa, et al.
Published: (2026)
by: Le, Minh Khoa, et al.
Published: (2026)
Similar Items
-
Learning to mask: Towards generalized face forgery detection
by: Fei, Jianwei, et al.
Published: (2022) -
PhotoHolmes: a Python library for forgery detection in digital images
by: O'Flaherty, Julián, et al.
Published: (2024) -
Deep video representation learning: a survey
by: Ravanbakhsh, Elham, et al.
Published: (2024) -
Multi-view Action Recognition via Directed Gromov-Wasserstein Discrepancy
by: Nguyen, Hoang-Quan, et al.
Published: (2024) -
Cross-view Action Recognition Understanding From Exocentric to Egocentric Perspective
by: Truong, Thanh-Dat, et al.
Published: (2023)