Common Objects Out of Context (COOCo): Investigating Multimodal Context and Semantic Scene Violations in Referential Communication
Fuente:
arXiv
Saved in:
| Main Authors: | Merlo, Filippo, Takmaz, Ece, Chen, Wenkai, Gatt, Albert |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Correlates of Image Memorability in Vision Encoders: Activations, Attention Entropy, Patch Uniformity and Autoencoder Losses
by: Takmaz, Ece, et al.
Published: (2025)
by: Takmaz, Ece, et al.
Published: (2025)
Model Merging to Maintain Language-Only Performance in Developmentally Plausible Multimodal Models
by: Takmaz, Ece, et al.
Published: (2025)
by: Takmaz, Ece, et al.
Published: (2025)
Decoding Emotions in Abstract Art: Cognitive Plausibility of CLIP in Recognizing Color-Emotion Associations
by: Widhoelzl, Hanna-Sophia, et al.
Published: (2024)
by: Widhoelzl, Hanna-Sophia, et al.
Published: (2024)
Context-aware Visual Storytelling with Visual Prefix Tuning and Contrastive Learning
by: Song, Yingjin, et al.
Published: (2024)
by: Song, Yingjin, et al.
Published: (2024)
Describing Images $\textit{Fast and Slow}$: Quantifying and Predicting the Variation in Human Signals during Visuo-Linguistic Processes
by: Takmaz, Ece, et al.
Published: (2024)
by: Takmaz, Ece, et al.
Published: (2024)
Burn After Reading: Do Multimodal Large Language Models Truly Capture Order of Events in Image Sequences?
by: Song, Yingjin, et al.
Published: (2025)
by: Song, Yingjin, et al.
Published: (2025)
From Image Captioning to Visual Storytelling
by: Passadakis, Admitos, et al.
Published: (2025)
by: Passadakis, Admitos, et al.
Published: (2025)
Common Inpainted Objects In-N-Out of Context
by: Yang, Tianze, et al.
Published: (2025)
by: Yang, Tianze, et al.
Published: (2025)
Image-Text Out-Of-Context Detection Using Synthetic Multimodal Misinformation
by: Shalabi, Fatma, et al.
Published: (2024)
by: Shalabi, Fatma, et al.
Published: (2024)
Make LVLMs Focus: Context-Aware Attention Modulation for Better Multimodal In-Context Learning
by: Li, Yanshu, et al.
Published: (2025)
by: Li, Yanshu, et al.
Published: (2025)
Summarize the Past to Predict the Future: Natural Language Descriptions of Context Boost Multimodal Object Interaction Anticipation
by: Pasca, Razvan-George, et al.
Published: (2023)
by: Pasca, Razvan-George, et al.
Published: (2023)
Mitigating GenAI-powered Evidence Pollution for Out-of-Context Multimodal Misinformation Detection
by: Yan, Zehong, et al.
Published: (2025)
by: Yan, Zehong, et al.
Published: (2025)
Support or Refute: Analyzing the Stance of Evidence to Detect Out-of-Context Mis- and Disinformation
by: Yuan, Xin, et al.
Published: (2023)
by: Yuan, Xin, et al.
Published: (2023)
Kosmos-G: Generating Images in Context with Multimodal Large Language Models
by: Pan, Xichen, et al.
Published: (2023)
by: Pan, Xichen, et al.
Published: (2023)
Grounding Language in Multi-Perspective Referential Communication
by: Tang, Zineng, et al.
Published: (2024)
by: Tang, Zineng, et al.
Published: (2024)
SNIFFER: Multimodal Large Language Model for Explainable Out-of-Context Misinformation Detection
by: Qi, Peng, et al.
Published: (2024)
by: Qi, Peng, et al.
Published: (2024)
Context-Aware Multimodal Pretraining
by: Roth, Karsten, et al.
Published: (2024)
by: Roth, Karsten, et al.
Published: (2024)
CODIS: Benchmarking Context-Dependent Visual Comprehension for Multimodal Large Language Models
by: Luo, Fuwen, et al.
Published: (2024)
by: Luo, Fuwen, et al.
Published: (2024)
CommVQA: Situating Visual Question Answering in Communicative Contexts
by: Naik, Nandita Shankar, et al.
Published: (2024)
by: Naik, Nandita Shankar, et al.
Published: (2024)
MathScape: Benchmarking Multimodal Large Language Models in Real-World Mathematical Contexts
by: Liang, Hao, et al.
Published: (2024)
by: Liang, Hao, et al.
Published: (2024)
Disentangling Instance and Scene Contexts for 3D Semantic Scene Completion
by: Liu, Enyu, et al.
Published: (2025)
by: Liu, Enyu, et al.
Published: (2025)
Scaling Beyond Context: A Survey of Multimodal Retrieval-Augmented Generation for Document Understanding
by: Gao, Sensen, et al.
Published: (2025)
by: Gao, Sensen, et al.
Published: (2025)
Referencing Where to Focus: Improving VisualGrounding with Referential Query
by: Wang, Yabing, et al.
Published: (2024)
by: Wang, Yabing, et al.
Published: (2024)
Learning from Fine-Grained Visual Discrepancies: Mitigating Multimodal Hallucinations via In-Context Visual Contrastive Optimization
by: Deng, Haolin, et al.
Published: (2026)
by: Deng, Haolin, et al.
Published: (2026)
Culture-Aware Humorous Captioning: Multimodal Humor Generation across Cultural Contexts
by: Xu, Run, et al.
Published: (2026)
by: Xu, Run, et al.
Published: (2026)
A Multimodal In-Context Tuning Approach for E-Commerce Product Description Generation
by: Li, Yunxin, et al.
Published: (2024)
by: Li, Yunxin, et al.
Published: (2024)
VimRAG: Navigating Massive Visual Context in Retrieval-Augmented Generation via Multimodal Memory Graph
by: Wang, Qiuchen, et al.
Published: (2026)
by: Wang, Qiuchen, et al.
Published: (2026)
SK-VQA: Synthetic Knowledge Generation at Scale for Training Context-Augmented Multimodal LLMs
by: Su, Xin, et al.
Published: (2024)
by: Su, Xin, et al.
Published: (2024)
FutureOmni: Evaluating Future Forecasting from Omni-Modal Context for Multimodal LLMs
by: Chen, Qian, et al.
Published: (2026)
by: Chen, Qian, et al.
Published: (2026)
LOOK-M: Look-Once Optimization in KV Cache for Efficient Multimodal Long-Context Inference
by: Wan, Zhongwei, et al.
Published: (2024)
by: Wan, Zhongwei, et al.
Published: (2024)
Detecting Offensive Memes with Social Biases in Singapore Context Using Multimodal Large Language Models
by: Yuxuan, Cao, et al.
Published: (2025)
by: Yuxuan, Cao, et al.
Published: (2025)
Improving Classification of Occluded Objects through Scene Context
by: King, Courtney M., et al.
Published: (2025)
by: King, Courtney M., et al.
Published: (2025)
OmChat: A Recipe to Train Multimodal Language Models with Strong Long Context and Video Understanding
by: Zhao, Tiancheng, et al.
Published: (2024)
by: Zhao, Tiancheng, et al.
Published: (2024)
According to Me: Long-Term Personalized Referential Memory QA
by: Mei, Jingbiao, et al.
Published: (2026)
by: Mei, Jingbiao, et al.
Published: (2026)
Placing Objects in Context via Inpainting for Out-of-distribution Segmentation
by: de Jorge, Pau, et al.
Published: (2024)
by: de Jorge, Pau, et al.
Published: (2024)
Contextual Image Attack: How Visual Context Exposes Multimodal Safety Vulnerabilities
by: Xiong, Yuan, et al.
Published: (2025)
by: Xiong, Yuan, et al.
Published: (2025)
Video-guided Machine Translation with Global Video Context
by: Chen, Jian, et al.
Published: (2026)
by: Chen, Jian, et al.
Published: (2026)
With Great Context Comes Great Prediction Power: Classifying Objects via Geo-Semantic Scene Graphs
by: Constantinescu, Ciprian, et al.
Published: (2025)
by: Constantinescu, Ciprian, et al.
Published: (2025)
SRPO: Self-Referential Policy Optimization for Vision-Language-Action Models
by: Fei, Senyu, et al.
Published: (2025)
by: Fei, Senyu, et al.
Published: (2025)
ZALM3: Zero-Shot Enhancement of Vision-Language Alignment via In-Context Information in Multi-Turn Multimodal Medical Dialogue
by: Li, Zhangpu, et al.
Published: (2024)
by: Li, Zhangpu, et al.
Published: (2024)
Similar Items
-
Correlates of Image Memorability in Vision Encoders: Activations, Attention Entropy, Patch Uniformity and Autoencoder Losses
by: Takmaz, Ece, et al.
Published: (2025) -
Model Merging to Maintain Language-Only Performance in Developmentally Plausible Multimodal Models
by: Takmaz, Ece, et al.
Published: (2025) -
Decoding Emotions in Abstract Art: Cognitive Plausibility of CLIP in Recognizing Color-Emotion Associations
by: Widhoelzl, Hanna-Sophia, et al.
Published: (2024) -
Context-aware Visual Storytelling with Visual Prefix Tuning and Contrastive Learning
by: Song, Yingjin, et al.
Published: (2024) -
Describing Images $\textit{Fast and Slow}$: Quantifying and Predicting the Variation in Human Signals during Visuo-Linguistic Processes
by: Takmaz, Ece, et al.
Published: (2024)