Beyond Surface Artifacts: Capturing Shared Latent Forgery Knowledge Across Modalities
Fuente:
arXiv
Saved in:
| Main Authors: | Dou, Jingtong, Shi, Chuancheng, Wang, Jian, Shen, Fei, Wang, Zhiyong, Chua, Tat-Seng |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
DNA: Uncovering Universal Latent Forgery Knowledge
by: Dou, Jingtong, et al.
Published: (2026)
by: Dou, Jingtong, et al.
Published: (2026)
Latent Anomaly Knowledge Excavation: Unveiling Sparse Sensitive Neurons in Vision-Language Models
by: Li, Shaotian, et al.
Published: (2026)
by: Li, Shaotian, et al.
Published: (2026)
HarmoniAD: Harmonizing Local Structures and Global Semantics for Anomaly Detection
by: Zhang, Naiqi, et al.
Published: (2026)
by: Zhang, Naiqi, et al.
Published: (2026)
Where Culture Fades: Revealing the Cultural Gap in Text-to-Image Generation
by: Shi, Chuancheng, et al.
Published: (2025)
by: Shi, Chuancheng, et al.
Published: (2025)
Suppressing Forgery-Specific Shortcuts for Generalizable Deepfake Detection
by: Wang, Yihui, et al.
Published: (2026)
by: Wang, Yihui, et al.
Published: (2026)
TraceRouter: Robust Safety for Large Foundation Models via Path-Level Intervention
by: Shi, Chuancheng, et al.
Published: (2026)
by: Shi, Chuancheng, et al.
Published: (2026)
Universal Scene Graph Generation
by: Wu, Shengqiong, et al.
Published: (2025)
by: Wu, Shengqiong, et al.
Published: (2025)
Who Transfers Safety? Identifying and Targeting Cross-Lingual Shared Safety Neurons
by: Zhang, Xianhui, et al.
Published: (2026)
by: Zhang, Xianhui, et al.
Published: (2026)
Turing Patterns for Multimedia: Reaction-Diffusion Multi-Modal Fusion for Language-Guided Video Moment Retrieval
by: Fang, Xiang, et al.
Published: (2026)
by: Fang, Xiang, et al.
Published: (2026)
Compose Your Aesthetics: Empowering Text-to-Image Models with the Principles of Art
by: Jin, Zhe, et al.
Published: (2025)
by: Jin, Zhe, et al.
Published: (2025)
Disentangling Masked Autoencoders for Unsupervised Domain Generalization
by: Zhang, An, et al.
Published: (2024)
by: Zhang, An, et al.
Published: (2024)
CLUE: Leveraging Low-Rank Adaptation to Capture Latent Uncovered Evidence for Image Forgery Localization
by: Wang, Youqi, et al.
Published: (2025)
by: Wang, Youqi, et al.
Published: (2025)
Precise Shield: Explaining and Aligning VLLM Safety via Neuron-Level Guidance
by: Shi, Enyi, et al.
Published: (2026)
by: Shi, Enyi, et al.
Published: (2026)
Generative Cross-Modal Retrieval: Memorizing Images in Multimodal Language Models for Retrieval and Beyond
by: Li, Yongqi, et al.
Published: (2024)
by: Li, Yongqi, et al.
Published: (2024)
Long-Term TalkingFace Generation via Motion-Prior Conditional Diffusion Model
by: Shen, Fei, et al.
Published: (2025)
by: Shen, Fei, et al.
Published: (2025)
OrthoEraser: Coupled-Neuron Orthogonal Projection for Concept Erasure
by: Shi, Chuancheng, et al.
Published: (2026)
by: Shi, Chuancheng, et al.
Published: (2026)
Towards Modality Generalization: A Benchmark and Prospective Analysis
by: Liu, Xiaohao, et al.
Published: (2024)
by: Liu, Xiaohao, et al.
Published: (2024)
Counterfactual Explanations for Face Forgery Detection via Adversarial Removal of Artifacts
by: Li, Yang, et al.
Published: (2024)
by: Li, Yang, et al.
Published: (2024)
Lingua-SafetyBench: A Benchmark for Safety Evaluation of Multilingual Vision-Language Models
by: Shi, Enyi, et al.
Published: (2026)
by: Shi, Enyi, et al.
Published: (2026)
Combating Multimodal LLM Hallucination via Bottom-Up Holistic Reasoning
by: Wu, Shengqiong, et al.
Published: (2024)
by: Wu, Shengqiong, et al.
Published: (2024)
Understanding Long Videos via LLM-Powered Entity Relation Graphs
by: Chu, Meng, et al.
Published: (2025)
by: Chu, Meng, et al.
Published: (2025)
AnchorFlow: Training-Free 3D Editing via Latent Anchor-Aligned Flows
by: Zhou, Zhenglin, et al.
Published: (2025)
by: Zhou, Zhenglin, et al.
Published: (2025)
Zero-1-to-A: Zero-Shot One Image to Animatable Head Avatars Using Video Diffusion
by: Zhou, Zhenglin, et al.
Published: (2025)
by: Zhou, Zhenglin, et al.
Published: (2025)
Dysen-VDM: Empowering Dynamics-aware Text-to-Video Diffusion with LLMs
by: Fei, Hao, et al.
Published: (2023)
by: Fei, Hao, et al.
Published: (2023)
Towards Natural Language-Guided Drones: GeoText-1652 Benchmark with Spatial Relation Matching
by: Chu, Meng, et al.
Published: (2023)
by: Chu, Meng, et al.
Published: (2023)
Extending Visual Dynamics for Video-to-Music Generation
by: Liu, Xiaohao, et al.
Published: (2025)
by: Liu, Xiaohao, et al.
Published: (2025)
Vitron: A Unified Pixel-level Vision LLM for Understanding, Generating, Segmenting, Editing
by: Fei, Hao, et al.
Published: (2024)
by: Fei, Hao, et al.
Published: (2024)
Multiple-environment Self-adaptive Network for Aerial-view Geo-localization
by: Wang, Tingyu, et al.
Published: (2022)
by: Wang, Tingyu, et al.
Published: (2022)
Thinking with Blueprints: Assisting Vision-Language Models in Spatial Reasoning via Structured Object Representation
by: Ma, Weijian, et al.
Published: (2026)
by: Ma, Weijian, et al.
Published: (2026)
Active Zero: Self-Evolving Vision-Language Models through Active Environment Exploration
by: He, Jinghan, et al.
Published: (2026)
by: He, Jinghan, et al.
Published: (2026)
ExpLLM: Towards Chain of Thought for Facial Expression Recognition
by: Lan, Xing, et al.
Published: (2024)
by: Lan, Xing, et al.
Published: (2024)
Composed Image Retrieval with Text Feedback via Multi-grained Uncertainty Regularization
by: Chen, Yiyang, et al.
Published: (2022)
by: Chen, Yiyang, et al.
Published: (2022)
3D Magic Mirror: Clothing Reconstruction from a Single Image via a Causal Perspective
by: Zheng, Zhedong, et al.
Published: (2022)
by: Zheng, Zhedong, et al.
Published: (2022)
Calibrated Multimodal Representation Learning with Missing Modalities
by: Liu, Xiaohao, et al.
Published: (2025)
by: Liu, Xiaohao, et al.
Published: (2025)
Towards Semantic Equivalence of Tokenization in Multimodal LLM
by: Wu, Shengqiong, et al.
Published: (2024)
by: Wu, Shengqiong, et al.
Published: (2024)
Uncertainty-Driven Expert Control: Enhancing the Reliability of Medical Vision-Language Models
by: Liang, Xiao, et al.
Published: (2025)
by: Liang, Xiao, et al.
Published: (2025)
Show Me What and Where has Changed? Question Answering and Grounding for Remote Sensing Change Detection
by: Li, Ke, et al.
Published: (2024)
by: Li, Ke, et al.
Published: (2024)
Enhancing Video-Language Representations with Structural Spatio-Temporal Alignment
by: Fei, Hao, et al.
Published: (2024)
by: Fei, Hao, et al.
Published: (2024)
Can I Trust Your Answer? Visually Grounded Video Question Answering
by: Xiao, Junbin, et al.
Published: (2023)
by: Xiao, Junbin, et al.
Published: (2023)
Agent4FaceForgery: Multi-Agent LLM Framework for Realistic Face Forgery Detection
by: Lai, Yingxin, et al.
Published: (2025)
by: Lai, Yingxin, et al.
Published: (2025)
Similar Items
-
DNA: Uncovering Universal Latent Forgery Knowledge
by: Dou, Jingtong, et al.
Published: (2026) -
Latent Anomaly Knowledge Excavation: Unveiling Sparse Sensitive Neurons in Vision-Language Models
by: Li, Shaotian, et al.
Published: (2026) -
HarmoniAD: Harmonizing Local Structures and Global Semantics for Anomaly Detection
by: Zhang, Naiqi, et al.
Published: (2026) -
Where Culture Fades: Revealing the Cultural Gap in Text-to-Image Generation
by: Shi, Chuancheng, et al.
Published: (2025) -
Suppressing Forgery-Specific Shortcuts for Generalizable Deepfake Detection
by: Wang, Yihui, et al.
Published: (2026)