What If the TV Was Off? Examining Counterfactual Reasoning Abilities of Multi-modal Language Models
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Zhang, Letian, Zhai, Xiaotong, Zhao, Zhongkai, Zong, Yongshuo, Wen, Xin, Zhao, Bingchen |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2023
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Benchmarking Multi-Image Understanding in Vision and Language Models: Perception, Knowledge, Reasoning, and Multi-Hop Reasoning
von: Zhao, Bingchen, et al.
Veröffentlicht: (2024)
von: Zhao, Bingchen, et al.
Veröffentlicht: (2024)
What Makes CLIP More Robust to Long-Tailed Pre-Training Data? A Controlled Study for Transferable Insights
von: Wen, Xin, et al.
Veröffentlicht: (2024)
von: Wen, Xin, et al.
Veröffentlicht: (2024)
What if...?: Thinking Counterfactual Keywords Helps to Mitigate Hallucination in Large Multi-modal Models
von: Kim, Junho, et al.
Veröffentlicht: (2024)
von: Kim, Junho, et al.
Veröffentlicht: (2024)
Customizing Visual-Language Foundation Models for Multi-modal Anomaly Detection and Reasoning
von: Xu, Xiaohao, et al.
Veröffentlicht: (2024)
von: Xu, Xiaohao, et al.
Veröffentlicht: (2024)
Structured and Abstractive Reasoning on Multi-modal Relational Knowledge Images
von: Zhang, Yichi, et al.
Veröffentlicht: (2025)
von: Zhang, Yichi, et al.
Veröffentlicht: (2025)
Actial: Activate Spatial Reasoning Ability of Multimodal Large Language Models
von: Zhan, Xiaoyu, et al.
Veröffentlicht: (2025)
von: Zhan, Xiaoyu, et al.
Veröffentlicht: (2025)
PROGRESSLM: Towards Progress Reasoning in Vision-Language Models
von: Zhang, Jianshu, et al.
Veröffentlicht: (2026)
von: Zhang, Jianshu, et al.
Veröffentlicht: (2026)
Advancing Large Multi-modal Models with Explicit Chain-of-Reasoning and Visual Question Generation
von: Uehara, Kohei, et al.
Veröffentlicht: (2024)
von: Uehara, Kohei, et al.
Veröffentlicht: (2024)
FSMR: A Feature Swapping Multi-modal Reasoning Approach with Joint Textual and Visual Clues
von: Li, Shuang, et al.
Veröffentlicht: (2024)
von: Li, Shuang, et al.
Veröffentlicht: (2024)
Empowering Segmentation Ability to Multi-modal Large Language Models
von: Yang, Yuqi, et al.
Veröffentlicht: (2024)
von: Yang, Yuqi, et al.
Veröffentlicht: (2024)
Enhancing the Reasoning Ability of Multimodal Large Language Models via Mixed Preference Optimization
von: Wang, Weiyun, et al.
Veröffentlicht: (2024)
von: Wang, Weiyun, et al.
Veröffentlicht: (2024)
Reasoning Can Hurt the Inductive Abilities of Large Language Models
von: Jin, Haibo, et al.
Veröffentlicht: (2025)
von: Jin, Haibo, et al.
Veröffentlicht: (2025)
GroundingGPT:Language Enhanced Multi-modal Grounding Model
von: Li, Zhaowei, et al.
Veröffentlicht: (2024)
von: Li, Zhaowei, et al.
Veröffentlicht: (2024)
AlignGPT: Multi-modal Large Language Models with Adaptive Alignment Capability
von: Zhao, Fei, et al.
Veröffentlicht: (2024)
von: Zhao, Fei, et al.
Veröffentlicht: (2024)
Large Multi-modal Models Can Interpret Features in Large Multi-modal Models
von: Zhang, Kaichen, et al.
Veröffentlicht: (2024)
von: Zhang, Kaichen, et al.
Veröffentlicht: (2024)
MMBench: Is Your Multi-modal Model an All-around Player?
von: Liu, Yuan, et al.
Veröffentlicht: (2023)
von: Liu, Yuan, et al.
Veröffentlicht: (2023)
What Makes for Good Visual Instructions? Synthesizing Complex Visual Reasoning Instructions for Visual Instruction Tuning
von: Du, Yifan, et al.
Veröffentlicht: (2023)
von: Du, Yifan, et al.
Veröffentlicht: (2023)
Expanding the Boundaries of Vision Prior Knowledge in Multi-modal Large Language Models
von: Liang, Qiao, et al.
Veröffentlicht: (2025)
von: Liang, Qiao, et al.
Veröffentlicht: (2025)
II-MMR: Identifying and Improving Multi-modal Multi-hop Reasoning in Visual Question Answering
von: Kil, Jihyung, et al.
Veröffentlicht: (2024)
von: Kil, Jihyung, et al.
Veröffentlicht: (2024)
Hybrid-Level Instruction Injection for Video Token Compression in Multi-modal Large Language Models
von: Liu, Zhihang, et al.
Veröffentlicht: (2025)
von: Liu, Zhihang, et al.
Veröffentlicht: (2025)
Assessing News Thumbnail Representativeness: Counterfactual text can enhance the cross-modal matching ability
von: Yoon, Yejun, et al.
Veröffentlicht: (2024)
von: Yoon, Yejun, et al.
Veröffentlicht: (2024)
Chartographer: Counterfactual Chart Generation for Evaluating Vision-Language Models
von: Jiang, Yifan, et al.
Veröffentlicht: (2026)
von: Jiang, Yifan, et al.
Veröffentlicht: (2026)
Behind the Magic, MERLIM: Multi-modal Evaluation Benchmark for Large Image-Language Models
von: Villa, Andrés, et al.
Veröffentlicht: (2023)
von: Villa, Andrés, et al.
Veröffentlicht: (2023)
UrbanLLaVA: A Multi-modal Large Language Model for Urban Intelligence with Spatial Reasoning and Understanding
von: Feng, Jie, et al.
Veröffentlicht: (2025)
von: Feng, Jie, et al.
Veröffentlicht: (2025)
Centurio: On Drivers of Multilingual Ability of Large Vision-Language Model
von: Geigle, Gregor, et al.
Veröffentlicht: (2025)
von: Geigle, Gregor, et al.
Veröffentlicht: (2025)
What If We Recaption Billions of Web Images with LLaMA-3?
von: Li, Xianhang, et al.
Veröffentlicht: (2024)
von: Li, Xianhang, et al.
Veröffentlicht: (2024)
Images are Achilles' Heel of Alignment: Exploiting Visual Vulnerabilities for Jailbreaking Multimodal Large Language Models
von: Li, Yifan, et al.
Veröffentlicht: (2024)
von: Li, Yifan, et al.
Veröffentlicht: (2024)
Can Vision Language Models Learn from Visual Demonstrations of Ambiguous Spatial Reasoning?
von: Zhao, Bowen, et al.
Veröffentlicht: (2024)
von: Zhao, Bowen, et al.
Veröffentlicht: (2024)
Measuring Social Bias in Vision-Language Models with Face-Only Counterfactuals from Real Photos
von: Chen, Haodong, et al.
Veröffentlicht: (2026)
von: Chen, Haodong, et al.
Veröffentlicht: (2026)
RoboSVG: A Unified Framework for Interactive SVG Generation with Multi-modal Guidance
von: Wang, Jiuniu, et al.
Veröffentlicht: (2025)
von: Wang, Jiuniu, et al.
Veröffentlicht: (2025)
Gemini in Reasoning: Unveiling Commonsense in Multimodal Large Language Models
von: Wang, Yuqing, et al.
Veröffentlicht: (2023)
von: Wang, Yuqing, et al.
Veröffentlicht: (2023)
Visual Language Tracking with Multi-modal Interaction: A Robust Benchmark
von: Li, Xuchen, et al.
Veröffentlicht: (2024)
von: Li, Xuchen, et al.
Veröffentlicht: (2024)
MMGR: Multi-Modal Generative Reasoning
von: Cai, Zefan, et al.
Veröffentlicht: (2025)
von: Cai, Zefan, et al.
Veröffentlicht: (2025)
Plug-and-Play Grounding of Reasoning in Multimodal Large Language Models
von: Chen, Jiaxing, et al.
Veröffentlicht: (2024)
von: Chen, Jiaxing, et al.
Veröffentlicht: (2024)
Multi-modal, Multi-task, Multi-criteria Automatic Evaluation with Vision Language Models
von: Ohi, Masanari, et al.
Veröffentlicht: (2024)
von: Ohi, Masanari, et al.
Veröffentlicht: (2024)
OmAgent: A Multi-modal Agent Framework for Complex Video Understanding with Task Divide-and-Conquer
von: Zhang, Lu, et al.
Veröffentlicht: (2024)
von: Zhang, Lu, et al.
Veröffentlicht: (2024)
EmbodiedBench: Comprehensive Benchmarking Multi-modal Large Language Models for Vision-Driven Embodied Agents
von: Yang, Rui, et al.
Veröffentlicht: (2025)
von: Yang, Rui, et al.
Veröffentlicht: (2025)
Sparkle: Mastering Basic Spatial Capabilities in Vision Language Models Elicits Generalization to Spatial Reasoning
von: Tang, Yihong, et al.
Veröffentlicht: (2024)
von: Tang, Yihong, et al.
Veröffentlicht: (2024)
GeoGPT4V: Towards Geometric Multi-modal Large Language Models with Geometric Image Generation
von: Cai, Shihao, et al.
Veröffentlicht: (2024)
von: Cai, Shihao, et al.
Veröffentlicht: (2024)
MuDPT: Multi-modal Deep-symphysis Prompt Tuning for Large Pre-trained Vision-Language Models
von: Miao, Yongzhu, et al.
Veröffentlicht: (2023)
von: Miao, Yongzhu, et al.
Veröffentlicht: (2023)
Ähnliche Einträge
-
Benchmarking Multi-Image Understanding in Vision and Language Models: Perception, Knowledge, Reasoning, and Multi-Hop Reasoning
von: Zhao, Bingchen, et al.
Veröffentlicht: (2024) -
What Makes CLIP More Robust to Long-Tailed Pre-Training Data? A Controlled Study for Transferable Insights
von: Wen, Xin, et al.
Veröffentlicht: (2024) -
What if...?: Thinking Counterfactual Keywords Helps to Mitigate Hallucination in Large Multi-modal Models
von: Kim, Junho, et al.
Veröffentlicht: (2024) -
Customizing Visual-Language Foundation Models for Multi-modal Anomaly Detection and Reasoning
von: Xu, Xiaohao, et al.
Veröffentlicht: (2024) -
Structured and Abstractive Reasoning on Multi-modal Relational Knowledge Images
von: Zhang, Yichi, et al.
Veröffentlicht: (2025)