Burn After Reading: Do Multimodal Large Language Models Truly Capture Order of Events in Image Sequences?
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Song, Yingjin, Du, Yupei, Paperno, Denis, Gatt, Albert |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Context-aware Visual Storytelling with Visual Prefix Tuning and Contrastive Learning
von: Song, Yingjin, et al.
Veröffentlicht: (2024)
von: Song, Yingjin, et al.
Veröffentlicht: (2024)
From Image Captioning to Visual Storytelling
von: Passadakis, Admitos, et al.
Veröffentlicht: (2025)
von: Passadakis, Admitos, et al.
Veröffentlicht: (2025)
Don't Learn, Ground: A Case for Natural Language Inference with Visual Grounding
von: Ignatev, Daniil, et al.
Veröffentlicht: (2025)
von: Ignatev, Daniil, et al.
Veröffentlicht: (2025)
Common Objects Out of Context (COOCo): Investigating Multimodal Context and Semantic Scene Violations in Referential Communication
von: Merlo, Filippo, et al.
Veröffentlicht: (2025)
von: Merlo, Filippo, et al.
Veröffentlicht: (2025)
Do Vision-Language Models Truly Perform Vision Reasoning? A Rigorous Study of the Modality Gap
von: Xu, Yige, et al.
Veröffentlicht: (2026)
von: Xu, Yige, et al.
Veröffentlicht: (2026)
Disentangling the Roles of Representation and Selection in Data Pruning
von: Du, Yupei, et al.
Veröffentlicht: (2025)
von: Du, Yupei, et al.
Veröffentlicht: (2025)
Do Multimodal Large Language Models Understand Welding?
von: Khvatskii, Grigorii, et al.
Veröffentlicht: (2025)
von: Khvatskii, Grigorii, et al.
Veröffentlicht: (2025)
LogicOCR: Do Your Large Multimodal Models Excel at Logical Reasoning on Text-Rich Images?
von: Ye, Maoyuan, et al.
Veröffentlicht: (2025)
von: Ye, Maoyuan, et al.
Veröffentlicht: (2025)
FTFT: Efficient and Robust Fine-Tuning by Transferring Training Dynamics
von: Du, Yupei, et al.
Veröffentlicht: (2023)
von: Du, Yupei, et al.
Veröffentlicht: (2023)
WSI-LLaVA: A Multimodal Large Language Model for Whole Slide Image
von: Liang, Yuci, et al.
Veröffentlicht: (2024)
von: Liang, Yuci, et al.
Veröffentlicht: (2024)
The Impact of Image Resolution on Biomedical Multimodal Large Language Models
von: Chen, Liangyu, et al.
Veröffentlicht: (2025)
von: Chen, Liangyu, et al.
Veröffentlicht: (2025)
Do Vision Encoders Truly Explain Object Hallucination?: Mitigating Object Hallucination via Simple Fine-Grained CLIPScore
von: Oh, Hongseok, et al.
Veröffentlicht: (2025)
von: Oh, Hongseok, et al.
Veröffentlicht: (2025)
Kosmos-G: Generating Images in Context with Multimodal Large Language Models
von: Pan, Xichen, et al.
Veröffentlicht: (2023)
von: Pan, Xichen, et al.
Veröffentlicht: (2023)
Do Egocentric Video-Language Models Truly Understand Hand-Object Interactions?
von: Xu, Boshen, et al.
Veröffentlicht: (2024)
von: Xu, Boshen, et al.
Veröffentlicht: (2024)
Do Composed Image Retrieval Benchmarks Require Multimodal Composition?
von: Attimonelli, Matteo, et al.
Veröffentlicht: (2026)
von: Attimonelli, Matteo, et al.
Veröffentlicht: (2026)
ROAP: A Reading-Order and Attention-Prior Pipeline for Optimizing Layout Transformers in Key Information Extraction
von: Xie, Tingwei, et al.
Veröffentlicht: (2026)
von: Xie, Tingwei, et al.
Veröffentlicht: (2026)
Mementos: A Comprehensive Benchmark for Multimodal Large Language Model Reasoning over Image Sequences
von: Wang, Xiyao, et al.
Veröffentlicht: (2024)
von: Wang, Xiyao, et al.
Veröffentlicht: (2024)
AesBench: An Expert Benchmark for Multimodal Large Language Models on Image Aesthetics Perception
von: Huang, Yipo, et al.
Veröffentlicht: (2024)
von: Huang, Yipo, et al.
Veröffentlicht: (2024)
From Perception to Reasoning: Deep Thinking Empowers Multimodal Large Language Models
von: Zhu, Wenxin, et al.
Veröffentlicht: (2025)
von: Zhu, Wenxin, et al.
Veröffentlicht: (2025)
Understanding Multimodal Procedural Knowledge by Sequencing Multimodal Instructional Manuals
von: Wu, Te-Lin, et al.
Veröffentlicht: (2021)
von: Wu, Te-Lin, et al.
Veröffentlicht: (2021)
Images are Achilles' Heel of Alignment: Exploiting Visual Vulnerabilities for Jailbreaking Multimodal Large Language Models
von: Li, Yifan, et al.
Veröffentlicht: (2024)
von: Li, Yifan, et al.
Veröffentlicht: (2024)
RestoreAgent: Autonomous Image Restoration Agent via Multimodal Large Language Models
von: Chen, Haoyu, et al.
Veröffentlicht: (2024)
von: Chen, Haoyu, et al.
Veröffentlicht: (2024)
Correlates of Image Memorability in Vision Encoders: Activations, Attention Entropy, Patch Uniformity and Autoencoder Losses
von: Takmaz, Ece, et al.
Veröffentlicht: (2025)
von: Takmaz, Ece, et al.
Veröffentlicht: (2025)
II-Bench: An Image Implication Understanding Benchmark for Multimodal Large Language Models
von: Liu, Ziqiang, et al.
Veröffentlicht: (2024)
von: Liu, Ziqiang, et al.
Veröffentlicht: (2024)
Indexing Multimodal Language Models for Large-scale Image Retrieval
von: Tharwat, Bahey, et al.
Veröffentlicht: (2026)
von: Tharwat, Bahey, et al.
Veröffentlicht: (2026)
ChatEXAONEPath: An Expert-level Multimodal Large Language Model for Histopathology Using Whole Slide Images
von: Kim, Sangwook, et al.
Veröffentlicht: (2025)
von: Kim, Sangwook, et al.
Veröffentlicht: (2025)
Model Composition for Multimodal Large Language Models
von: Chen, Chi, et al.
Veröffentlicht: (2024)
von: Chen, Chi, et al.
Veröffentlicht: (2024)
ICM-Assistant: Instruction-tuning Multimodal Large Language Models for Rule-based Explainable Image Content Moderation
von: Wu, Mengyang, et al.
Veröffentlicht: (2024)
von: Wu, Mengyang, et al.
Veröffentlicht: (2024)
ImageChain: Advancing Sequential Image-to-Text Reasoning in Multimodal Large Language Models
von: Villegas, Danae Sánchez, et al.
Veröffentlicht: (2025)
von: Villegas, Danae Sánchez, et al.
Veröffentlicht: (2025)
Looking Beyond Text: Reducing Language bias in Large Vision-Language Models via Multimodal Dual-Attention and Soft-Image Guidance
von: Zhao, Haozhe, et al.
Veröffentlicht: (2024)
von: Zhao, Haozhe, et al.
Veröffentlicht: (2024)
AdaMMS: Model Merging for Heterogeneous Multimodal Large Language Models with Unsupervised Coefficient Optimization
von: Du, Yiyang, et al.
Veröffentlicht: (2025)
von: Du, Yiyang, et al.
Veröffentlicht: (2025)
Grounding Partially-Defined Events in Multimodal Data
von: Sanders, Kate, et al.
Veröffentlicht: (2024)
von: Sanders, Kate, et al.
Veröffentlicht: (2024)
Mitigating Behavioral Hallucination in Multimodal Large Language Models for Sequential Images
von: You, Liangliang, et al.
Veröffentlicht: (2025)
von: You, Liangliang, et al.
Veröffentlicht: (2025)
Argus Inspection: Do Multimodal Large Language Models Possess the Eye of Panoptes?
von: Yao, Yang, et al.
Veröffentlicht: (2025)
von: Yao, Yang, et al.
Veröffentlicht: (2025)
The Use of Multimodal Large Language Models to Detect Objects from Thermal Images: Transportation Applications
von: Ashqar, Huthaifa I., et al.
Veröffentlicht: (2024)
von: Ashqar, Huthaifa I., et al.
Veröffentlicht: (2024)
Mixed-R1: Unified Reward Perspective For Reasoning Capability in Multimodal Large Language Models
von: Xu, Shilin, et al.
Veröffentlicht: (2025)
von: Xu, Shilin, et al.
Veröffentlicht: (2025)
ScImage: How Good Are Multimodal Large Language Models at Scientific Text-to-Image Generation?
von: Zhang, Leixin, et al.
Veröffentlicht: (2024)
von: Zhang, Leixin, et al.
Veröffentlicht: (2024)
Plug-and-Play Grounding of Reasoning in Multimodal Large Language Models
von: Chen, Jiaxing, et al.
Veröffentlicht: (2024)
von: Chen, Jiaxing, et al.
Veröffentlicht: (2024)
Can Large Vision-Language Models Understand Multimodal Sarcasm?
von: Wang, Xinyu, et al.
Veröffentlicht: (2025)
von: Wang, Xinyu, et al.
Veröffentlicht: (2025)
LLAVADI: What Matters For Multimodal Large Language Models Distillation
von: Xu, Shilin, et al.
Veröffentlicht: (2024)
von: Xu, Shilin, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Context-aware Visual Storytelling with Visual Prefix Tuning and Contrastive Learning
von: Song, Yingjin, et al.
Veröffentlicht: (2024) -
From Image Captioning to Visual Storytelling
von: Passadakis, Admitos, et al.
Veröffentlicht: (2025) -
Don't Learn, Ground: A Case for Natural Language Inference with Visual Grounding
von: Ignatev, Daniil, et al.
Veröffentlicht: (2025) -
Common Objects Out of Context (COOCo): Investigating Multimodal Context and Semantic Scene Violations in Referential Communication
von: Merlo, Filippo, et al.
Veröffentlicht: (2025) -
Do Vision-Language Models Truly Perform Vision Reasoning? A Rigorous Study of the Modality Gap
von: Xu, Yige, et al.
Veröffentlicht: (2026)