Leveraging Chat-Based Large Vision Language Models for Multimodal Out-Of-Context Detection
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Shalabi, Fatma, Felouat, Hichem, Nguyen, Huy H., Echizen, Isao |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Image-Text Out-Of-Context Detection Using Synthetic Multimodal Misinformation
von: Shalabi, Fatma, et al.
Veröffentlicht: (2024)
von: Shalabi, Fatma, et al.
Veröffentlicht: (2024)
GFT-GCN: Privacy-Preserving 3D Face Mesh Recognition with Spectral Diffusion
von: Felouat, Hichem, et al.
Veröffentlicht: (2025)
von: Felouat, Hichem, et al.
Veröffentlicht: (2025)
Multimodal Adversarial Defense for Vision-Language Models by Leveraging One-To-Many Relationships
von: Waseda, Futa, et al.
Veröffentlicht: (2024)
von: Waseda, Futa, et al.
Veröffentlicht: (2024)
Exploring Self-Supervised Vision Transformers for Deepfake Detection: A Comparative Analysis
von: Nguyen, Huy H., et al.
Veröffentlicht: (2024)
von: Nguyen, Huy H., et al.
Veröffentlicht: (2024)
Tell-Tale Watermarks for Explanatory Reasoning in Synthetic Media Forensics
von: Chang, Ching-Chun, et al.
Veröffentlicht: (2025)
von: Chang, Ching-Chun, et al.
Veröffentlicht: (2025)
UlcerGPT: A Multimodal Approach Leveraging Large Language and Vision Models for Diabetic Foot Ulcer Image Transcription
von: Basiri, Reza, et al.
Veröffentlicht: (2024)
von: Basiri, Reza, et al.
Veröffentlicht: (2024)
On the Out-Of-Distribution Generalization of Multimodal Large Language Models
von: Zhang, Xingxuan, et al.
Veröffentlicht: (2024)
von: Zhang, Xingxuan, et al.
Veröffentlicht: (2024)
Quality Text, Robust Vision: The Role of Language in Enhancing Visual Robustness of Vision-Language Models
von: Waseda, Futa, et al.
Veröffentlicht: (2025)
von: Waseda, Futa, et al.
Veröffentlicht: (2025)
Beyond Standard Benchmarks: A Systematic Audit of Vision-Language Model's Robustness to Natural Semantic Variation Across Diverse Tasks
von: Chengyu, Jia, et al.
Veröffentlicht: (2026)
von: Chengyu, Jia, et al.
Veröffentlicht: (2026)
VisionGraph: Leveraging Large Multimodal Models for Graph Theory Problems in Visual Context
von: Li, Yunxin, et al.
Veröffentlicht: (2024)
von: Li, Yunxin, et al.
Veröffentlicht: (2024)
SNIFFER: Multimodal Large Language Model for Explainable Out-of-Context Misinformation Detection
von: Qi, Peng, et al.
Veröffentlicht: (2024)
von: Qi, Peng, et al.
Veröffentlicht: (2024)
Phantasia: Context-Adaptive Backdoors in Vision Language Models
von: Tran, Nam Duong, et al.
Veröffentlicht: (2026)
von: Tran, Nam Duong, et al.
Veröffentlicht: (2026)
Leveraging Vision-Language Models to Detect Attention in Educational Videos
von: Becquet, Gabriel, et al.
Veröffentlicht: (2026)
von: Becquet, Gabriel, et al.
Veröffentlicht: (2026)
Defending Against Physical Adversarial Patch Attacks on Infrared Human Detection
von: Strack, Lukas, et al.
Veröffentlicht: (2023)
von: Strack, Lukas, et al.
Veröffentlicht: (2023)
Exploring Active Data Selection Strategies for Continuous Training in Deepfake Detection
von: Furuhashi, Yoshihiko, et al.
Veröffentlicht: (2025)
von: Furuhashi, Yoshihiko, et al.
Veröffentlicht: (2025)
AgriChat: A Multimodal Large Language Model for Agriculture Image Understanding
von: Boudiaf, Abderrahmene, et al.
Veröffentlicht: (2026)
von: Boudiaf, Abderrahmene, et al.
Veröffentlicht: (2026)
Vision Language Model-based Caption Evaluation Method Leveraging Visual Context Extraction
von: Maeda, Koki, et al.
Veröffentlicht: (2024)
von: Maeda, Koki, et al.
Veröffentlicht: (2024)
WAVER: Writing-style Agnostic Text-Video Retrieval via Distilling Vision-Language Models Through Open-Vocabulary Knowledge
von: Le, Huy, et al.
Veröffentlicht: (2023)
von: Le, Huy, et al.
Veröffentlicht: (2023)
OODBench: Out-of-Distribution Benchmark for Large Vision-Language Models
von: Lin, Ling, et al.
Veröffentlicht: (2026)
von: Lin, Ling, et al.
Veröffentlicht: (2026)
Check Field Detection Agent (CFD-Agent) using Multimodal Large Language and Vision Language Models
von: Halder, Sourav, et al.
Veröffentlicht: (2025)
von: Halder, Sourav, et al.
Veröffentlicht: (2025)
Large Vision-Language Models as Emotion Recognizers in Context Awareness
von: Lei, Yuxuan, et al.
Veröffentlicht: (2024)
von: Lei, Yuxuan, et al.
Veröffentlicht: (2024)
Hierarchical Vision-Language Learning for Medical Out-of-Distribution Detection
von: Lai, Runhe, et al.
Veröffentlicht: (2025)
von: Lai, Runhe, et al.
Veröffentlicht: (2025)
Multimodal Contextualized Support for Enhancing Video Retrieval System
von: Nguyen-Le, Quoc-Bao, et al.
Veröffentlicht: (2024)
von: Nguyen-Le, Quoc-Bao, et al.
Veröffentlicht: (2024)
CrashChat: A Multimodal Large Language Model for Multitask Traffic Crash Video Analysis
von: Liang, Kaidi, et al.
Veröffentlicht: (2025)
von: Liang, Kaidi, et al.
Veröffentlicht: (2025)
STER-VLM: Spatio-Temporal With Enhanced Reference Vision-Language Models
von: Nguyen-Nhu, Tinh-Anh, et al.
Veröffentlicht: (2025)
von: Nguyen-Nhu, Tinh-Anh, et al.
Veröffentlicht: (2025)
Fine-Tuning Text-To-Image Diffusion Models for Class-Wise Spurious Feature Generation
von: MaungMaung, AprilPyone, et al.
Veröffentlicht: (2024)
von: MaungMaung, AprilPyone, et al.
Veröffentlicht: (2024)
From Captions to Rewards (CAREVL): Leveraging Large Language Model Experts for Enhanced Reward Modeling in Large Vision-Language Models
von: Dai, Muzhi, et al.
Veröffentlicht: (2025)
von: Dai, Muzhi, et al.
Veröffentlicht: (2025)
Leveraging ChatGPT's Multimodal Vision Capabilities to Rank Satellite Images by Poverty Level: Advancing Tools for Social Science Research
von: Sarmadi, Hamid, et al.
Veröffentlicht: (2025)
von: Sarmadi, Hamid, et al.
Veröffentlicht: (2025)
Advancing Multimodal In-Context Learning in Large Vision-Language Models with Task-aware Demonstrations
von: Li, Yanshu
Veröffentlicht: (2025)
von: Li, Yanshu
Veröffentlicht: (2025)
The Paradigm Shift: A Comprehensive Survey on Large Vision Language Models for Multimodal Fake News Detection
von: Ai, Wei, et al.
Veröffentlicht: (2026)
von: Ai, Wei, et al.
Veröffentlicht: (2026)
Object Detection with Multimodal Large Vision-Language Models: An In-depth Review
von: Sapkota, Ranjan, et al.
Veröffentlicht: (2025)
von: Sapkota, Ranjan, et al.
Veröffentlicht: (2025)
LAST: Leveraging Tools as Hints to Enhance Spatial Reasoning for Multimodal Large Language Models
von: Tian, Shi-Yu, et al.
Veröffentlicht: (2026)
von: Tian, Shi-Yu, et al.
Veröffentlicht: (2026)
Contextual Object Detection with Multimodal Large Language Models
von: Zang, Yuhang, et al.
Veröffentlicht: (2023)
von: Zang, Yuhang, et al.
Veröffentlicht: (2023)
LookupForensics: A Large-Scale Multi-Task Dataset for Multi-Phase Image-Based Fact Verification
von: Cui, Shuhan, et al.
Veröffentlicht: (2024)
von: Cui, Shuhan, et al.
Veröffentlicht: (2024)
White-box Multimodal Jailbreaks Against Large Vision-Language Models
von: Wang, Ruofan, et al.
Veröffentlicht: (2024)
von: Wang, Ruofan, et al.
Veröffentlicht: (2024)
Investigating Redundancy in Multimodal Large Language Models with Multiple Vision Encoders
von: Wang, Yizhou, et al.
Veröffentlicht: (2025)
von: Wang, Yizhou, et al.
Veröffentlicht: (2025)
A Controllable 3D Deepfake Generation Framework with Gaussian Splatting
von: Liu, Wending, et al.
Veröffentlicht: (2025)
von: Liu, Wending, et al.
Veröffentlicht: (2025)
Visual-Noise Guided In-Context Distillation for Multimodal Large Language Model Unlearning
von: Chen, Junkai, et al.
Veröffentlicht: (2026)
von: Chen, Junkai, et al.
Veröffentlicht: (2026)
Respecting Modality Gap in Post-hoc Out-of-distribution Detection with Pre-trained Vision-Language Models
von: Hu, Yuanwei, et al.
Veröffentlicht: (2026)
von: Hu, Yuanwei, et al.
Veröffentlicht: (2026)
Self-Calibrated Tuning of Vision-Language Models for Out-of-Distribution Detection
von: Yu, Geng, et al.
Veröffentlicht: (2024)
von: Yu, Geng, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Image-Text Out-Of-Context Detection Using Synthetic Multimodal Misinformation
von: Shalabi, Fatma, et al.
Veröffentlicht: (2024) -
GFT-GCN: Privacy-Preserving 3D Face Mesh Recognition with Spectral Diffusion
von: Felouat, Hichem, et al.
Veröffentlicht: (2025) -
Multimodal Adversarial Defense for Vision-Language Models by Leveraging One-To-Many Relationships
von: Waseda, Futa, et al.
Veröffentlicht: (2024) -
Exploring Self-Supervised Vision Transformers for Deepfake Detection: A Comparative Analysis
von: Nguyen, Huy H., et al.
Veröffentlicht: (2024) -
Tell-Tale Watermarks for Explanatory Reasoning in Synthetic Media Forensics
von: Chang, Ching-Chun, et al.
Veröffentlicht: (2025)