Towards Unified Multimodal Misinformation Detection in Social Media: A Benchmark Dataset and Baseline
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Li, Haiyang, Wang, Yaxiong, Tang, Shengeng, Wu, Lianwei, Cheng, Lechao, Zhong, Zhun |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
OmniVL-Guard: Towards Unified Vision-Language Forgery Detection and Grounding via Balanced RL
von: Shen, Jinjie, et al.
Veröffentlicht: (2026)
von: Shen, Jinjie, et al.
Veröffentlicht: (2026)
Beyond Artificial Misalignment: Detecting and Grounding Semantic-Coordinated Multimodal Manipulations
von: Shen, Jinjie, et al.
Veröffentlicht: (2025)
von: Shen, Jinjie, et al.
Veröffentlicht: (2025)
Towards Fine-Grained Emotion Understanding via Skeleton-Based Micro-Gesture Recognition
von: Xu, Hao, et al.
Veröffentlicht: (2025)
von: Xu, Hao, et al.
Veröffentlicht: (2025)
EntityCLIP: Entity-Centric Image-Text Matching via Multimodal Attentive Contrastive Learning
von: Wang, Yaxiong, et al.
Veröffentlicht: (2024)
von: Wang, Yaxiong, et al.
Veröffentlicht: (2024)
OmniVL-Guard Pro: A Tool-Augmented Agent for Omnibus Vision-Language Forensics
von: Shen, Jinjie, et al.
Veröffentlicht: (2026)
von: Shen, Jinjie, et al.
Veröffentlicht: (2026)
Generating Attribution Reports for Manipulated Facial Images: A Dataset and Baseline
von: Lian, Jingchun, et al.
Veröffentlicht: (2024)
von: Lian, Jingchun, et al.
Veröffentlicht: (2024)
Knowledge Swapping via Learning and Unlearning
von: Xing, Mingyu, et al.
Veröffentlicht: (2025)
von: Xing, Mingyu, et al.
Veröffentlicht: (2025)
TDEdit: A Unified Diffusion Framework for Text-Drag Guided Image Manipulation
von: Wang, Qihang, et al.
Veröffentlicht: (2025)
von: Wang, Qihang, et al.
Veröffentlicht: (2025)
Towards Micro-Action Recognition with Limited Annotations: An Asynchronous Pseudo Labeling and Training Approach
von: Zhang, Yan, et al.
Veröffentlicht: (2025)
von: Zhang, Yan, et al.
Veröffentlicht: (2025)
Towards Ancient Plant Seed Classification: A Benchmark Dataset and Baseline Model
von: Xing, Rui, et al.
Veröffentlicht: (2025)
von: Xing, Rui, et al.
Veröffentlicht: (2025)
Text2Lip: Progressive Lip-Synced Talking Face Generation from Text via Viseme-Guided Rendering
von: Wang, Xu, et al.
Veröffentlicht: (2025)
von: Wang, Xu, et al.
Veröffentlicht: (2025)
VMID: A Multimodal Fusion LLM Framework for Detecting and Identifying Misinformation of Short Videos
von: Zhong, Weihao, et al.
Veröffentlicht: (2024)
von: Zhong, Weihao, et al.
Veröffentlicht: (2024)
CanonSLR: Canonical-View Guided Multi-View Continuous Sign Language Recognition
von: Wang, Xu, et al.
Veröffentlicht: (2026)
von: Wang, Xu, et al.
Veröffentlicht: (2026)
Motion is the Choreographer: Learning Latent Pose Dynamics for Seamless Sign Language Generation
von: He, Jiayi, et al.
Veröffentlicht: (2025)
von: He, Jiayi, et al.
Veröffentlicht: (2025)
Text-Driven Diffusion Model for Sign Language Production
von: He, Jiayi, et al.
Veröffentlicht: (2025)
von: He, Jiayi, et al.
Veröffentlicht: (2025)
ASAP: Advancing Semantic Alignment Promotes Multi-Modal Manipulation Detecting and Grounding
von: Zhang, Zhenxing, et al.
Veröffentlicht: (2024)
von: Zhang, Zhenxing, et al.
Veröffentlicht: (2024)
SSAM: Self-Supervised Association Modeling for Test-Time Adaption
von: Wang, Yaxiong, et al.
Veröffentlicht: (2025)
von: Wang, Yaxiong, et al.
Veröffentlicht: (2025)
A Large-scale Universal Evaluation Benchmark For Face Forgery Detection
von: Bei, Yijun, et al.
Veröffentlicht: (2024)
von: Bei, Yijun, et al.
Veröffentlicht: (2024)
SIDA: Social Media Image Deepfake Detection, Localization and Explanation with Large Multimodal Model
von: Huang, Zhenglin, et al.
Veröffentlicht: (2024)
von: Huang, Zhenglin, et al.
Veröffentlicht: (2024)
PointCloud-Text Matching: Benchmark Datasets and a Baseline
von: Feng, Yanglin, et al.
Veröffentlicht: (2024)
von: Feng, Yanglin, et al.
Veröffentlicht: (2024)
Modality Alignment Meets Federated Broadcasting
von: Ma, Yuting, et al.
Veröffentlicht: (2024)
von: Ma, Yuting, et al.
Veröffentlicht: (2024)
EventSTR: A Benchmark Dataset and Baselines for Event Stream based Scene Text Recognition
von: Wang, Xiao, et al.
Veröffentlicht: (2025)
von: Wang, Xiao, et al.
Veröffentlicht: (2025)
UNIAA: A Unified Multi-modal Image Aesthetic Assessment Baseline and Benchmark
von: Zhou, Zhaokun, et al.
Veröffentlicht: (2024)
von: Zhou, Zhaokun, et al.
Veröffentlicht: (2024)
Complex Mathematical Expression Recognition: Benchmark, Large-Scale Dataset and Strong Baseline
von: Bai, Weikang, et al.
Veröffentlicht: (2025)
von: Bai, Weikang, et al.
Veröffentlicht: (2025)
FedHPL: Efficient Heterogeneous Federated Learning with Prompt Tuning and Logit Distillation
von: Ma, Yuting, et al.
Veröffentlicht: (2024)
von: Ma, Yuting, et al.
Veröffentlicht: (2024)
A Simple Aerial Detection Baseline of Multimodal Language Models
von: Li, Qingyun, et al.
Veröffentlicht: (2025)
von: Li, Qingyun, et al.
Veröffentlicht: (2025)
Fact-R1: Towards Explainable Video Misinformation Detection with Deep Reasoning
von: Zhang, Fanrui, et al.
Veröffentlicht: (2025)
von: Zhang, Fanrui, et al.
Veröffentlicht: (2025)
Dataset Distillers Are Good Label Denoisers In the Wild
von: Cheng, Lechao, et al.
Veröffentlicht: (2024)
von: Cheng, Lechao, et al.
Veröffentlicht: (2024)
MMR-AD: A Large-Scale Multimodal Dataset for Benchmarking General Anomaly Detection with Multimodal Large Language Models
von: Yao, Xincheng, et al.
Veröffentlicht: (2026)
von: Yao, Xincheng, et al.
Veröffentlicht: (2026)
Cultivating Forensic Reasoning for Generalizable Multimodal Manipulation Detection
von: Zhang, Yuchen, et al.
Veröffentlicht: (2026)
von: Zhang, Yuchen, et al.
Veröffentlicht: (2026)
Image-Text Out-Of-Context Detection Using Synthetic Multimodal Misinformation
von: Shalabi, Fatma, et al.
Veröffentlicht: (2024)
von: Shalabi, Fatma, et al.
Veröffentlicht: (2024)
Open-World 3D Scene Graph Generation for Retrieval-Augmented Reasoning
von: Yu, Fei, et al.
Veröffentlicht: (2025)
von: Yu, Fei, et al.
Veröffentlicht: (2025)
Shaping a Stabilized Video by Mitigating Unintended Changes for Concept-Augmented Video Editing
von: Guo, Mingce, et al.
Veröffentlicht: (2024)
von: Guo, Mingce, et al.
Veröffentlicht: (2024)
VIVID-10M: A Dataset and Baseline for Versatile and Interactive Video Local Editing
von: Hu, Jiahao, et al.
Veröffentlicht: (2024)
von: Hu, Jiahao, et al.
Veröffentlicht: (2024)
MCITlib: Multimodal Continual Instruction Tuning Library and Benchmark
von: Guo, Haiyang, et al.
Veröffentlicht: (2025)
von: Guo, Haiyang, et al.
Veröffentlicht: (2025)
RAW: Robust Avatar Watermarking -- Benchmarking and Baseline
von: Parry, Jack, et al.
Veröffentlicht: (2026)
von: Parry, Jack, et al.
Veröffentlicht: (2026)
A Multimodal Benchmark Dataset and Model for Crop Disease Diagnosis
von: Liu, Xiang, et al.
Veröffentlicht: (2025)
von: Liu, Xiang, et al.
Veröffentlicht: (2025)
ChangeDiff: A Multi-Temporal Change Detection Data Generator with Flexible Text Prompts via Diffusion Model
von: Zang, Qi, et al.
Veröffentlicht: (2024)
von: Zang, Qi, et al.
Veröffentlicht: (2024)
Omni-Fake: Benchmarking Unified Multimodal Social Media Deepfake Detection
von: Li, Tianxiao, et al.
Veröffentlicht: (2026)
von: Li, Tianxiao, et al.
Veröffentlicht: (2026)
Envision: Benchmarking Unified Understanding & Generation for Causal World Process Insights
von: Tian, Juanxi, et al.
Veröffentlicht: (2025)
von: Tian, Juanxi, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
OmniVL-Guard: Towards Unified Vision-Language Forgery Detection and Grounding via Balanced RL
von: Shen, Jinjie, et al.
Veröffentlicht: (2026) -
Beyond Artificial Misalignment: Detecting and Grounding Semantic-Coordinated Multimodal Manipulations
von: Shen, Jinjie, et al.
Veröffentlicht: (2025) -
Towards Fine-Grained Emotion Understanding via Skeleton-Based Micro-Gesture Recognition
von: Xu, Hao, et al.
Veröffentlicht: (2025) -
EntityCLIP: Entity-Centric Image-Text Matching via Multimodal Attentive Contrastive Learning
von: Wang, Yaxiong, et al.
Veröffentlicht: (2024) -
OmniVL-Guard Pro: A Tool-Augmented Agent for Omnibus Vision-Language Forensics
von: Shen, Jinjie, et al.
Veröffentlicht: (2026)