Cultivating Forensic Reasoning for Generalizable Multimodal Manipulation Detection
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Zhang, Yuchen, Wang, Yaxiong, Han, Kecheng, Wu, Yujiao, Wu, Lianwei, Zhu, Li, Zheng, Zhedong |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
The Coherence Trap: When MLLM-Crafted Narratives Exploit Manipulated Visual Contexts
von: Zhang, Yuchen, et al.
Veröffentlicht: (2025)
von: Zhang, Yuchen, et al.
Veröffentlicht: (2025)
Generating Attribution Reports for Manipulated Facial Images: A Dataset and Baseline
von: Lian, Jingchun, et al.
Veröffentlicht: (2024)
von: Lian, Jingchun, et al.
Veröffentlicht: (2024)
REVEAL: Reference-Grounded Reasoning for Multimodal Manipulation Detection
von: Zhou, Jun, et al.
Veröffentlicht: (2026)
von: Zhou, Jun, et al.
Veröffentlicht: (2026)
EntityCLIP: Entity-Centric Image-Text Matching via Multimodal Attentive Contrastive Learning
von: Wang, Yaxiong, et al.
Veröffentlicht: (2024)
von: Wang, Yaxiong, et al.
Veröffentlicht: (2024)
Towards Unified Multimodal Misinformation Detection in Social Media: A Benchmark Dataset and Baseline
von: Li, Haiyang, et al.
Veröffentlicht: (2025)
von: Li, Haiyang, et al.
Veröffentlicht: (2025)
Can Video Diffusion Models Predict Past Frames? Bidirectional Cycle Consistency for Reversible Interpolation
von: Liu, Lingyu, et al.
Veröffentlicht: (2026)
von: Liu, Lingyu, et al.
Veröffentlicht: (2026)
Beyond Walking: A Large-Scale Image-Text Benchmark for Text-based Person Anomaly Search
von: Yang, Shuyu, et al.
Veröffentlicht: (2024)
von: Yang, Shuyu, et al.
Veröffentlicht: (2024)
Every Painting Awakened: A Training-free Framework for Painting-to-Animation Generation
von: Liu, Lingyu, et al.
Veröffentlicht: (2025)
von: Liu, Lingyu, et al.
Veröffentlicht: (2025)
Minimizing the Pretraining Gap: Domain-aligned Text-Based Person Retrieval
von: Yang, Shuyu, et al.
Veröffentlicht: (2025)
von: Yang, Shuyu, et al.
Veröffentlicht: (2025)
Towards Scalable Video Anomaly Retrieval: A Synthetic Video-Text Benchmark
von: Yang, Shuyu, et al.
Veröffentlicht: (2025)
von: Yang, Shuyu, et al.
Veröffentlicht: (2025)
OmniVL-Guard Pro: A Tool-Augmented Agent for Omnibus Vision-Language Forensics
von: Shen, Jinjie, et al.
Veröffentlicht: (2026)
von: Shen, Jinjie, et al.
Veröffentlicht: (2026)
Look, Compare and Draw: Differential Query Transformer for Automatic Oil Painting
von: Liu, Lingyu, et al.
Veröffentlicht: (2026)
von: Liu, Lingyu, et al.
Veröffentlicht: (2026)
SketchThinker-R1: Towards Efficient Sketch-Style Reasoning in Large Multimodal Models
von: Zhang, Ruiyang, et al.
Veröffentlicht: (2026)
von: Zhang, Ruiyang, et al.
Veröffentlicht: (2026)
Pretrain-then-Adapt: Uncertainty-Aware Test-Time Adaptation for Text-based Person Search
von: Zhang, Jiahao, et al.
Veröffentlicht: (2026)
von: Zhang, Jiahao, et al.
Veröffentlicht: (2026)
AINet+: Advancing Superpixel Segmentation via Cascaded Association Implantation
von: Wang, Yaxiong, et al.
Veröffentlicht: (2021)
von: Wang, Yaxiong, et al.
Veröffentlicht: (2021)
Beyond Artificial Misalignment: Detecting and Grounding Semantic-Coordinated Multimodal Manipulations
von: Shen, Jinjie, et al.
Veröffentlicht: (2025)
von: Shen, Jinjie, et al.
Veröffentlicht: (2025)
Dual Relation Alignment for Composed Image Retrieval
von: Jiang, Xintong, et al.
Veröffentlicht: (2023)
von: Jiang, Xintong, et al.
Veröffentlicht: (2023)
VL-Uncertainty: Detecting Hallucination in Large Vision-Language Model via Uncertainty Estimation
von: Zhang, Ruiyang, et al.
Veröffentlicht: (2024)
von: Zhang, Ruiyang, et al.
Veröffentlicht: (2024)
GeoSense: Internalizing Geometric Necessity Perception for Multimodal Reasoning
von: Liu, Ruiheng, et al.
Veröffentlicht: (2026)
von: Liu, Ruiheng, et al.
Veröffentlicht: (2026)
Scale Up Composed Image Retrieval Learning via Modification Text Generation
von: Zhou, Yinan, et al.
Veröffentlicht: (2025)
von: Zhou, Yinan, et al.
Veröffentlicht: (2025)
Learning Consistent Taxonomic Classification through Hierarchical Reasoning
von: Li, Zhenghong, et al.
Veröffentlicht: (2026)
von: Li, Zhenghong, et al.
Veröffentlicht: (2026)
Afford-X: Generalizable and Slim Affordance Reasoning for Task-oriented Manipulation
von: Zhu, Xiaomeng, et al.
Veröffentlicht: (2025)
von: Zhu, Xiaomeng, et al.
Veröffentlicht: (2025)
TIGeR: Text-Instructed Generation and Refinement for Template-Free Hand-Object Interaction
von: Huang, Yiyao, et al.
Veröffentlicht: (2025)
von: Huang, Yiyao, et al.
Veröffentlicht: (2025)
Deepfake Forensics Adapter: A Dual-Stream Network for Generalizable Deepfake Detection
von: Liao, Jianfeng, et al.
Veröffentlicht: (2026)
von: Liao, Jianfeng, et al.
Veröffentlicht: (2026)
Harnessing Uncertainty-aware Bounding Boxes for Unsupervised 3D Object Detection
von: Zhang, Ruiyang, et al.
Veröffentlicht: (2024)
von: Zhang, Ruiyang, et al.
Veröffentlicht: (2024)
Beyond Known Clusters: Probe New Prototypes for Efficient Generalized Class Discovery
von: Wang, Ye, et al.
Veröffentlicht: (2024)
von: Wang, Ye, et al.
Veröffentlicht: (2024)
Approaching Outside: Scaling Unsupervised 3D Object Detection from 2D Scene
von: Zhang, Ruiyang, et al.
Veröffentlicht: (2024)
von: Zhang, Ruiyang, et al.
Veröffentlicht: (2024)
Forensics Adapter: Unleashing CLIP for Generalizable Face Forgery Detection
von: Cui, Xinjie, et al.
Veröffentlicht: (2024)
von: Cui, Xinjie, et al.
Veröffentlicht: (2024)
VSearcher: Long-Horizon Multimodal Search Agent via Reinforcement Learning
von: Zhang, Ruiyang, et al.
Veröffentlicht: (2026)
von: Zhang, Ruiyang, et al.
Veröffentlicht: (2026)
Explainable Forensics of Manipulated Segments in Untrimmed Long Videos
von: Feng, Yue, et al.
Veröffentlicht: (2026)
von: Feng, Yue, et al.
Veröffentlicht: (2026)
Celeb-DF++: A Large-scale Challenging Video DeepFake Benchmark for Generalizable Forensics
von: Li, Yuezun, et al.
Veröffentlicht: (2025)
von: Li, Yuezun, et al.
Veröffentlicht: (2025)
Collaborative Group: Composed Image Retrieval via Consensus Learning from Noisy Annotations
von: Zhang, Xu, et al.
Veröffentlicht: (2023)
von: Zhang, Xu, et al.
Veröffentlicht: (2023)
Can Vision-Language Models Think from the Sky? Unifying UAV Reasoning and Generation
von: Sun, Jintao, et al.
Veröffentlicht: (2026)
von: Sun, Jintao, et al.
Veröffentlicht: (2026)
FauForensics: Boosting Audio-Visual Deepfake Detection with Facial Action Units
von: Wang, Jian, et al.
Veröffentlicht: (2025)
von: Wang, Jian, et al.
Veröffentlicht: (2025)
Contextual AD Narration with Interleaved Multimodal Sequence
von: Wang, Hanlin, et al.
Veröffentlicht: (2024)
von: Wang, Hanlin, et al.
Veröffentlicht: (2024)
REVEAL: Reasoning-Enhanced Forensic Evidence Analysis for Explainable AI-Generated Image Detection
von: Cao, Huangsen, et al.
Veröffentlicht: (2025)
von: Cao, Huangsen, et al.
Veröffentlicht: (2025)
Uncertainty-o: One Model-agnostic Framework for Unveiling Uncertainty in Large Multimodal Models
von: Zhang, Ruiyang, et al.
Veröffentlicht: (2025)
von: Zhang, Ruiyang, et al.
Veröffentlicht: (2025)
Seeing What Matters: Generalizable AI-generated Video Detection with Forensic-Oriented Augmentation
von: Corvi, Riccardo, et al.
Veröffentlicht: (2025)
von: Corvi, Riccardo, et al.
Veröffentlicht: (2025)
Weather-R1: Logically Consistent Reinforcement Fine-Tuning for Multimodal Reasoning in Meteorology
von: Wu, Kaiyu, et al.
Veröffentlicht: (2026)
von: Wu, Kaiyu, et al.
Veröffentlicht: (2026)
From Instruction to Event: Sound-Triggered Mobile Manipulation
von: Ju, Hao, et al.
Veröffentlicht: (2026)
von: Ju, Hao, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
The Coherence Trap: When MLLM-Crafted Narratives Exploit Manipulated Visual Contexts
von: Zhang, Yuchen, et al.
Veröffentlicht: (2025) -
Generating Attribution Reports for Manipulated Facial Images: A Dataset and Baseline
von: Lian, Jingchun, et al.
Veröffentlicht: (2024) -
REVEAL: Reference-Grounded Reasoning for Multimodal Manipulation Detection
von: Zhou, Jun, et al.
Veröffentlicht: (2026) -
EntityCLIP: Entity-Centric Image-Text Matching via Multimodal Attentive Contrastive Learning
von: Wang, Yaxiong, et al.
Veröffentlicht: (2024) -
Towards Unified Multimodal Misinformation Detection in Social Media: A Benchmark Dataset and Baseline
von: Li, Haiyang, et al.
Veröffentlicht: (2025)