Retrieving Counterfactuals Improves Visual In-Context Learning
Fuente:
arXiv
Salvato in:
| Autori principali: | Xiong, Guangzhi, Sinha, Sanchit, He, Zhenghao, Zhang, Aidong |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Rethinking Visual Attribution for Chest X-ray Reasoning in Large Vision Language Models
di: Xiong, Guangzhi, et al.
Pubblicazione: (2026)
di: Xiong, Guangzhi, et al.
Pubblicazione: (2026)
Concept-RuleNet: Grounded Multi-Agent Neurosymbolic Reasoning in Vision Language Models
di: Sinha, Sanchit, et al.
Pubblicazione: (2025)
di: Sinha, Sanchit, et al.
Pubblicazione: (2025)
GCAV: A Global Concept Activation Vector Framework for Cross-Layer Consistency in Interpretability
di: He, Zhenghao, et al.
Pubblicazione: (2025)
di: He, Zhenghao, et al.
Pubblicazione: (2025)
Toward Faithful Retrieval-Augmented Generation with Sparse Autoencoders
di: Xiong, Guangzhi, et al.
Pubblicazione: (2025)
di: Xiong, Guangzhi, et al.
Pubblicazione: (2025)
Attention-guided Fine-tuning of Multimodal Large Language Models Improves Chain-of-Thought Reasoning
di: Sinha, Sanchit, et al.
Pubblicazione: (2026)
di: Sinha, Sanchit, et al.
Pubblicazione: (2026)
COCO-Tree: Compositional Hierarchical Concept Trees for Enhanced Reasoning in Vision Language Models
di: Sinha, Sanchit, et al.
Pubblicazione: (2025)
di: Sinha, Sanchit, et al.
Pubblicazione: (2025)
Reasoning Beyond Chain-of-Thought: A Latent Computational Mode in Large Language Models
di: He, Zhenghao, et al.
Pubblicazione: (2026)
di: He, Zhenghao, et al.
Pubblicazione: (2026)
ASCENT-ViT: Attention-based Scale-aware Concept Learning Framework for Enhanced Alignment in Vision Transformers
di: Sinha, Sanchit, et al.
Pubblicazione: (2025)
di: Sinha, Sanchit, et al.
Pubblicazione: (2025)
CASL: Concept-Aligned Sparse Latents for Interpreting Diffusion Models
di: He, Zhenghao, et al.
Pubblicazione: (2026)
di: He, Zhenghao, et al.
Pubblicazione: (2026)
Weaving Context Across Images: Improving Vision-Language Models through Focus-Centric Visual Chains
di: Zhang, Juntian, et al.
Pubblicazione: (2025)
di: Zhang, Juntian, et al.
Pubblicazione: (2025)
How Does the Textual Information Affect the Retrieval of Multimodal In-Context Learning?
di: Luo, Yang, et al.
Pubblicazione: (2024)
di: Luo, Yang, et al.
Pubblicazione: (2024)
EmoGist: Efficient In-Context Learning for Visual Emotion Understanding
di: Seoh, Ronald, et al.
Pubblicazione: (2025)
di: Seoh, Ronald, et al.
Pubblicazione: (2025)
Echo-4o: Harnessing the Power of GPT-4o Synthetic Images for Improved Image Generation
di: Ye, Junyan, et al.
Pubblicazione: (2025)
di: Ye, Junyan, et al.
Pubblicazione: (2025)
Nearest Neighbor Normalization Improves Multimodal Retrieval
di: Chowdhury, Neil, et al.
Pubblicazione: (2024)
di: Chowdhury, Neil, et al.
Pubblicazione: (2024)
Internalized Reasoning for Long-Context Visual Document Understanding
di: Veselka, Austin
Pubblicazione: (2026)
di: Veselka, Austin
Pubblicazione: (2026)
Benchmarking Retrieval-Augmented Generation for Medicine
di: Xiong, Guangzhi, et al.
Pubblicazione: (2024)
di: Xiong, Guangzhi, et al.
Pubblicazione: (2024)
VisionGraph: Leveraging Large Multimodal Models for Graph Theory Problems in Visual Context
di: Li, Yunxin, et al.
Pubblicazione: (2024)
di: Li, Yunxin, et al.
Pubblicazione: (2024)
How to Train Your Long-Context Visual Document Model
di: Veselka, Austin
Pubblicazione: (2026)
di: Veselka, Austin
Pubblicazione: (2026)
Neural Additive Experts: Context-Gated Experts for Controllable Model Additivity
di: Xiong, Guangzhi, et al.
Pubblicazione: (2026)
di: Xiong, Guangzhi, et al.
Pubblicazione: (2026)
Bayesian Example Selection Improves In-Context Learning for Speech, Text, and Visual Modalities
di: Wang, Siyin, et al.
Pubblicazione: (2024)
di: Wang, Siyin, et al.
Pubblicazione: (2024)
Watch Before You Answer: Learning from Visually Grounded Post-Training
di: Zhang, Yuxuan, et al.
Pubblicazione: (2026)
di: Zhang, Yuxuan, et al.
Pubblicazione: (2026)
RICO: Improving Accuracy and Completeness in Image Recaptioning via Visual Reconstruction
di: Wang, Yuchi, et al.
Pubblicazione: (2025)
di: Wang, Yuchi, et al.
Pubblicazione: (2025)
GIRAFFE: Design Choices for Extending the Context Length of Visual Language Models
di: Li, Mukai, et al.
Pubblicazione: (2024)
di: Li, Mukai, et al.
Pubblicazione: (2024)
Semantically-Prompted Language Models Improve Visual Descriptions
di: Ogezi, Michael, et al.
Pubblicazione: (2023)
di: Ogezi, Michael, et al.
Pubblicazione: (2023)
BEAT: Visual Backdoor Attacks on VLM-based Embodied Agents via Contrastive Trigger Learning
di: Zhan, Qiusi, et al.
Pubblicazione: (2025)
di: Zhan, Qiusi, et al.
Pubblicazione: (2025)
Jigsaw-R1: A Study of Rule-based Visual Reinforcement Learning with Jigsaw Puzzles
di: Wang, Zifu, et al.
Pubblicazione: (2025)
di: Wang, Zifu, et al.
Pubblicazione: (2025)
Query-centric Audio-Visual Cognition Network for Moment Retrieval, Segmentation and Step-Captioning
di: Tu, Yunbin, et al.
Pubblicazione: (2024)
di: Tu, Yunbin, et al.
Pubblicazione: (2024)
Improving Retrieval-Augmented Generation in Medicine with Iterative Follow-up Questions
di: Xiong, Guangzhi, et al.
Pubblicazione: (2024)
di: Xiong, Guangzhi, et al.
Pubblicazione: (2024)
Impact of Visual Context on Noisy Multimodal NMT: An Empirical Study for English to Indian Languages
di: Gain, Baban, et al.
Pubblicazione: (2023)
di: Gain, Baban, et al.
Pubblicazione: (2023)
Retrieval-Augmented Long-Context Translation for Cultural Image Captioning: Gators submission for AmericasNLP 2026 shared task
di: Dhawan, Aashish, et al.
Pubblicazione: (2026)
di: Dhawan, Aashish, et al.
Pubblicazione: (2026)
How Far Are We from Intelligent Visual Deductive Reasoning?
di: Zhang, Yizhe, et al.
Pubblicazione: (2024)
di: Zhang, Yizhe, et al.
Pubblicazione: (2024)
GSM8K-V: Can Vision Language Models Solve Grade School Math Word Problems in Visual Contexts
di: Yuan, Fan, et al.
Pubblicazione: (2025)
di: Yuan, Fan, et al.
Pubblicazione: (2025)
Losing Visual Needles in Image Haystacks: Vision Language Models are Easily Distracted in Short and Long Contexts
di: Sharma, Aditya, et al.
Pubblicazione: (2024)
di: Sharma, Aditya, et al.
Pubblicazione: (2024)
Learning Visual Composition through Improved Semantic Guidance
di: Stone, Austin, et al.
Pubblicazione: (2024)
di: Stone, Austin, et al.
Pubblicazione: (2024)
What if...?: Thinking Counterfactual Keywords Helps to Mitigate Hallucination in Large Multi-modal Models
di: Kim, Junho, et al.
Pubblicazione: (2024)
di: Kim, Junho, et al.
Pubblicazione: (2024)
Retrieval-based Disentangled Representation Learning with Natural Language Supervision
di: Zhou, Jiawei, et al.
Pubblicazione: (2022)
di: Zhou, Jiawei, et al.
Pubblicazione: (2022)
Eliciting In-Context Learning in Vision-Language Models for Videos Through Curated Data Distributional Properties
di: Yu, Keunwoo Peter, et al.
Pubblicazione: (2023)
di: Yu, Keunwoo Peter, et al.
Pubblicazione: (2023)
Guiding Vision-Language Model Selection for Visual Question-Answering Across Tasks, Domains, and Knowledge Types
di: Sinha, Neelabh, et al.
Pubblicazione: (2024)
di: Sinha, Neelabh, et al.
Pubblicazione: (2024)
Measuring Social Bias in Vision-Language Models with Face-Only Counterfactuals from Real Photos
di: Chen, Haodong, et al.
Pubblicazione: (2026)
di: Chen, Haodong, et al.
Pubblicazione: (2026)
What Factors Affect Multi-Modal In-Context Learning? An In-Depth Exploration
di: Qin, Libo, et al.
Pubblicazione: (2024)
di: Qin, Libo, et al.
Pubblicazione: (2024)
Documenti analoghi
-
Rethinking Visual Attribution for Chest X-ray Reasoning in Large Vision Language Models
di: Xiong, Guangzhi, et al.
Pubblicazione: (2026) -
Concept-RuleNet: Grounded Multi-Agent Neurosymbolic Reasoning in Vision Language Models
di: Sinha, Sanchit, et al.
Pubblicazione: (2025) -
GCAV: A Global Concept Activation Vector Framework for Cross-Layer Consistency in Interpretability
di: He, Zhenghao, et al.
Pubblicazione: (2025) -
Toward Faithful Retrieval-Augmented Generation with Sparse Autoencoders
di: Xiong, Guangzhi, et al.
Pubblicazione: (2025) -
Attention-guided Fine-tuning of Multimodal Large Language Models Improves Chain-of-Thought Reasoning
di: Sinha, Sanchit, et al.
Pubblicazione: (2026)