Spot the Fake: Large Multimodal Model-Based Synthetic Image Detection with Artifact Explanation
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Wen, Siwei, Ye, Junyan, Feng, Peilin, Kang, Hengrui, Wen, Zichen, Chen, Yize, Wu, Jiang, Wu, Wenjun, He, Conghui, Li, Weijia |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
LEGION: Learning to Ground and Explain for Synthetic Image Detection
von: Kang, Hengrui, et al.
Veröffentlicht: (2025)
von: Kang, Hengrui, et al.
Veröffentlicht: (2025)
OmniAID: Decoupling Semantic and Artifacts for Universal AI-Generated Image Detection in the Wild
von: Guo, Yuncheng, et al.
Veröffentlicht: (2025)
von: Guo, Yuncheng, et al.
Veröffentlicht: (2025)
FakeVLM-R1: Internalizing Physical Laws via CoT for Synthetic Image Detection
von: Zhu, Leqi, et al.
Veröffentlicht: (2026)
von: Zhu, Leqi, et al.
Veröffentlicht: (2026)
LOKI: A Comprehensive Synthetic Data Detection Benchmark using Large Multimodal Models
von: Ye, Junyan, et al.
Veröffentlicht: (2024)
von: Ye, Junyan, et al.
Veröffentlicht: (2024)
Token Pruning in Multimodal Large Language Models: Are We Solving the Right Problem?
von: Wen, Zichen, et al.
Veröffentlicht: (2025)
von: Wen, Zichen, et al.
Veröffentlicht: (2025)
OmniDocLayout: Towards Diverse Document Layout Generation via Coarse-to-Fine LLM Learning
von: Kang, Hengrui, et al.
Veröffentlicht: (2025)
von: Kang, Hengrui, et al.
Veröffentlicht: (2025)
Respecting Self-Uncertainty in On-Policy Self-Distillation for Efficient LLM Reasoning
von: Ke, Junlong, et al.
Veröffentlicht: (2026)
von: Ke, Junlong, et al.
Veröffentlicht: (2026)
DocLayout-YOLO: Enhancing Document Layout Analysis through Diverse Synthetic Data and Global-to-Local Adaptive Perception
von: Zhao, Zhiyuan, et al.
Veröffentlicht: (2024)
von: Zhao, Zhiyuan, et al.
Veröffentlicht: (2024)
Can ChatGPT Detect DeepFakes? A Study of Using Multimodal Large Language Models for Media Forensics
von: Jia, Shan, et al.
Veröffentlicht: (2024)
von: Jia, Shan, et al.
Veröffentlicht: (2024)
Stop Looking for Important Tokens in Multimodal Language Models: Duplication Matters More
von: Wen, Zichen, et al.
Veröffentlicht: (2025)
von: Wen, Zichen, et al.
Veröffentlicht: (2025)
Echo-4o: Harnessing the Power of GPT-4o Synthetic Images for Improved Image Generation
von: Ye, Junyan, et al.
Veröffentlicht: (2025)
von: Ye, Junyan, et al.
Veröffentlicht: (2025)
Leveraging BEV Paradigm for Ground-to-Aerial Image Synthesis
von: Ye, Junyan, et al.
Veröffentlicht: (2024)
von: Ye, Junyan, et al.
Veröffentlicht: (2024)
UrBench: A Comprehensive Benchmark for Evaluating Large Multimodal Models in Multi-View Urban Scenarios
von: Zhou, Baichuan, et al.
Veröffentlicht: (2024)
von: Zhou, Baichuan, et al.
Veröffentlicht: (2024)
Prune2Drive: A Plug-and-Play Framework for Accelerating Vision-Language Models in Autonomous Driving
von: Xiong, Minhao, et al.
Veröffentlicht: (2025)
von: Xiong, Minhao, et al.
Veröffentlicht: (2025)
Bench-CoE: a Framework for Collaboration of Experts from Benchmark
von: Wang, Yuanshuai, et al.
Veröffentlicht: (2024)
von: Wang, Yuanshuai, et al.
Veröffentlicht: (2024)
A Self-Learning Multimodal Approach for Fake News Detection
von: Chen, Hao, et al.
Veröffentlicht: (2024)
von: Chen, Hao, et al.
Veröffentlicht: (2024)
Detect Fake with Fake: Leveraging Synthetic Data-driven Representation for Synthetic Image Detection
von: Otake, Hina, et al.
Veröffentlicht: (2024)
von: Otake, Hina, et al.
Veröffentlicht: (2024)
IPCV: Information-Preserving Compression for MLLM Visual Encoders
von: Chen, Yuan, et al.
Veröffentlicht: (2025)
von: Chen, Yuan, et al.
Veröffentlicht: (2025)
ThinkFake: Reasoning in Multimodal Large Language Models for AI-Generated Image Detection
von: Huang, Tai-Ming, et al.
Veröffentlicht: (2025)
von: Huang, Tai-Ming, et al.
Veröffentlicht: (2025)
Earth-Agent: Unlocking the Full Landscape of Earth Observation with Agents
von: Feng, Peilin, et al.
Veröffentlicht: (2025)
von: Feng, Peilin, et al.
Veröffentlicht: (2025)
Parrot Captions Teach CLIP to Spot Text
von: Lin, Yiqi, et al.
Veröffentlicht: (2023)
von: Lin, Yiqi, et al.
Veröffentlicht: (2023)
TinyLLaVA-Video-R1: Towards Smaller LMMs for Video Reasoning
von: Zhang, Xingjian, et al.
Veröffentlicht: (2025)
von: Zhang, Xingjian, et al.
Veröffentlicht: (2025)
EDGE-GRPO: Entropy-Driven GRPO with Guided Error Correction for Advantage Diversity
von: Zhang, Xingjian, et al.
Veröffentlicht: (2025)
von: Zhang, Xingjian, et al.
Veröffentlicht: (2025)
SG-BEV: Satellite-Guided BEV Fusion for Cross-View Semantic Segmentation
von: Ye, Junyan, et al.
Veröffentlicht: (2024)
von: Ye, Junyan, et al.
Veröffentlicht: (2024)
Cross-view image geo-localization with Panorama-BEV Co-Retrieval Network
von: Ye, Junyan, et al.
Veröffentlicht: (2024)
von: Ye, Junyan, et al.
Veröffentlicht: (2024)
RealGen: Photorealistic Text-to-Image Generation via Detector-Guided Rewards
von: Ye, Junyan, et al.
Veröffentlicht: (2025)
von: Ye, Junyan, et al.
Veröffentlicht: (2025)
FakeGPT: Fake News Generation, Explanation and Detection of Large Language Models
von: Huang, Yue, et al.
Veröffentlicht: (2023)
von: Huang, Yue, et al.
Veröffentlicht: (2023)
GPT-ImgEval: A Comprehensive Benchmark for Diagnosing GPT4o in Image Generation
von: Yan, Zhiyuan, et al.
Veröffentlicht: (2025)
von: Yan, Zhiyuan, et al.
Veröffentlicht: (2025)
FakeBench: Probing Explainable Fake Image Detection via Large Multimodal Models
von: Li, Yixuan, et al.
Veröffentlicht: (2024)
von: Li, Yixuan, et al.
Veröffentlicht: (2024)
Large Visual-Language Models Are Also Good Classifiers: A Study of In-Context Multimodal Fake News Detection
von: Jiang, Ye, et al.
Veröffentlicht: (2024)
von: Jiang, Ye, et al.
Veröffentlicht: (2024)
BLINK-Twice: You see, but do you observe? A Reasoning Benchmark on Visual Perception
von: Ye, Junyan, et al.
Veröffentlicht: (2025)
von: Ye, Junyan, et al.
Veröffentlicht: (2025)
SIDA: Social Media Image Deepfake Detection, Localization and Explanation with Large Multimodal Model
von: Huang, Zhenglin, et al.
Veröffentlicht: (2024)
von: Huang, Zhenglin, et al.
Veröffentlicht: (2024)
CrossViewDiff: A Cross-View Diffusion Model for Satellite-to-Street View Synthesis
von: Li, Weijia, et al.
Veröffentlicht: (2024)
von: Li, Weijia, et al.
Veröffentlicht: (2024)
Omni-Fake: Benchmarking Unified Multimodal Social Media Deepfake Detection
von: Li, Tianxiao, et al.
Veröffentlicht: (2026)
von: Li, Tianxiao, et al.
Veröffentlicht: (2026)
Efficient Multi-modal Large Language Models via Progressive Consistency Distillation
von: Wen, Zichen, et al.
Veröffentlicht: (2025)
von: Wen, Zichen, et al.
Veröffentlicht: (2025)
GenClaw: Code-Driven Agentic Image Generation
von: Ye, Junyan, et al.
Veröffentlicht: (2026)
von: Ye, Junyan, et al.
Veröffentlicht: (2026)
Where am I? Cross-View Geo-localization with Natural Language Descriptions
von: Ye, Junyan, et al.
Veröffentlicht: (2024)
von: Ye, Junyan, et al.
Veröffentlicht: (2024)
Exploring Modality Disruption in Multimodal Fake News Detection
von: Liu, Moyang, et al.
Veröffentlicht: (2025)
von: Liu, Moyang, et al.
Veröffentlicht: (2025)
EventMemAgent: Hierarchical Event-Centric Memory for Online Video Understanding with Adaptive Tool Use
von: Wen, Siwei, et al.
Veröffentlicht: (2026)
von: Wen, Siwei, et al.
Veröffentlicht: (2026)
Scene4U: Hierarchical Layered 3D Scene Reconstruction from Single Panoramic Image for Your Immerse Exploration
von: Huang, Zilong, et al.
Veröffentlicht: (2025)
von: Huang, Zilong, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
LEGION: Learning to Ground and Explain for Synthetic Image Detection
von: Kang, Hengrui, et al.
Veröffentlicht: (2025) -
OmniAID: Decoupling Semantic and Artifacts for Universal AI-Generated Image Detection in the Wild
von: Guo, Yuncheng, et al.
Veröffentlicht: (2025) -
FakeVLM-R1: Internalizing Physical Laws via CoT for Synthetic Image Detection
von: Zhu, Leqi, et al.
Veröffentlicht: (2026) -
LOKI: A Comprehensive Synthetic Data Detection Benchmark using Large Multimodal Models
von: Ye, Junyan, et al.
Veröffentlicht: (2024) -
Token Pruning in Multimodal Large Language Models: Are We Solving the Right Problem?
von: Wen, Zichen, et al.
Veröffentlicht: (2025)