Escaping the SpuriVerse: Can Large Vision-Language Models Generalize Beyond Seen Spurious Correlations?
Fuente:
arXiv
Saved in:
| Main Authors: | Yang, Yiwei, Lee, Chung Peng, Feng, Shangbin, Zhao, Dora, Wen, Bingbing, Liu, Anthony Z., Tsvetkov, Yulia, Howe, Bill |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Know Your Limits: A Survey of Abstention in Large Language Models
by: Wen, Bingbing, et al.
Published: (2024)
by: Wen, Bingbing, et al.
Published: (2024)
Data Swarms: Optimizable Generation of Synthetic Evaluation Data
by: Feng, Shangbin, et al.
Published: (2025)
by: Feng, Shangbin, et al.
Published: (2025)
ScienceMeter: Tracking Scientific Knowledge Updates in Language Models
by: Wang, Yike, et al.
Published: (2025)
by: Wang, Yike, et al.
Published: (2025)
The Single-Multi Evolution Loop for Self-Improving Model Collaboration Systems
by: Feng, Shangbin, et al.
Published: (2026)
by: Feng, Shangbin, et al.
Published: (2026)
Among Us: Measuring and Mitigating Malicious Contributions in Model Collaboration Systems
by: Yang, Ziyuan, et al.
Published: (2026)
by: Yang, Ziyuan, et al.
Published: (2026)
MentorCollab: Selective Large-to-Small Inference-Time Guidance for Efficient Reasoning
by: Wang, Haojin, et al.
Published: (2026)
by: Wang, Haojin, et al.
Published: (2026)
Can Language Models Solve Graph Problems in Natural Language?
by: Wang, Heng, et al.
Published: (2023)
by: Wang, Heng, et al.
Published: (2023)
Can LLM Graph Reasoning Generalize beyond Pattern Memorization?
by: Zhang, Yizhuo, et al.
Published: (2024)
by: Zhang, Yizhuo, et al.
Published: (2024)
What Does the Bot Say? Opportunities and Risks of Large Language Models in Social Media Bot Detection
by: Feng, Shangbin, et al.
Published: (2024)
by: Feng, Shangbin, et al.
Published: (2024)
SPARTA ALIGNMENT: Collectively Aligning Multiple Language Models through Combat
by: Jiang, Yuru, et al.
Published: (2025)
by: Jiang, Yuru, et al.
Published: (2025)
Characterizing LLM Abstention Behavior in Science QA with Context Perturbations
by: Wen, Bingbing, et al.
Published: (2024)
by: Wen, Bingbing, et al.
Published: (2024)
Towards Zero-Shot Annotation of the Built Environment with Vision-Language Models (Vision Paper)
by: Han, Bin, et al.
Published: (2024)
by: Han, Bin, et al.
Published: (2024)
Varying Shades of Wrong: Aligning LLMs with Wrong Answers Only
by: Yao, Jihan, et al.
Published: (2024)
by: Yao, Jihan, et al.
Published: (2024)
Resolving Knowledge Conflicts in Large Language Models
by: Wang, Yike, et al.
Published: (2023)
by: Wang, Yike, et al.
Published: (2023)
Knowledge Crosswords: Geometric Knowledge Reasoning with Large Language Models
by: Ding, Wenxuan, et al.
Published: (2023)
by: Ding, Wenxuan, et al.
Published: (2023)
KGQuiz: Evaluating the Generalization of Encoded Knowledge in Large Language Models
by: Bai, Yuyang, et al.
Published: (2023)
by: Bai, Yuyang, et al.
Published: (2023)
Quantifying Language Models' Sensitivity to Spurious Features in Prompt Design or: How I learned to start worrying about prompt formatting
by: Sclar, Melanie, et al.
Published: (2023)
by: Sclar, Melanie, et al.
Published: (2023)
DELL: Generating Reactions and Explanations for LLM-Based Misinformation Detection
by: Wan, Herun, et al.
Published: (2024)
by: Wan, Herun, et al.
Published: (2024)
Generalizable LLM Learning of Graph Synthetic Data with Post-training Alignment
by: Zhang, Yizhuo, et al.
Published: (2025)
by: Zhang, Yizhuo, et al.
Published: (2025)
Don't Hallucinate, Abstain: Identifying LLM Knowledge Gaps via Multi-LLM Collaboration
by: Feng, Shangbin, et al.
Published: (2024)
by: Feng, Shangbin, et al.
Published: (2024)
Don't Throw Away Your Pretrained Model
by: Feng, Shangbin, et al.
Published: (2025)
by: Feng, Shangbin, et al.
Published: (2025)
Knowledge Card: Filling LLMs' Knowledge Gaps with Plug-in Specialized Language Models
by: Feng, Shangbin, et al.
Published: (2023)
by: Feng, Shangbin, et al.
Published: (2023)
GuessBench: Sensemaking Multimodal Creativity in the Wild
by: Zhu, Zifeng, et al.
Published: (2025)
by: Zhu, Zifeng, et al.
Published: (2025)
Small Reward Models via Backward Inference
by: Wang, Yike, et al.
Published: (2026)
by: Wang, Yike, et al.
Published: (2026)
FACTS&EVIDENCE: An Interactive Tool for Transparent Fine-Grained Factual Verification of Machine-Generated Text
by: Boonsanong, Varich, et al.
Published: (2025)
by: Boonsanong, Varich, et al.
Published: (2025)
Interactive Reasoning: Visualizing and Controlling Chain-of-Thought Reasoning in Large Language Models
by: Pang, Rock Yuren, et al.
Published: (2025)
by: Pang, Rock Yuren, et al.
Published: (2025)
P^3SUM: Preserving Author's Perspective in News Summarization with Diffusion Language Models
by: Liu, Yuhan, et al.
Published: (2023)
by: Liu, Yuhan, et al.
Published: (2023)
Modular Pluralism: Pluralistic Alignment via Multi-LLM Collaboration
by: Feng, Shangbin, et al.
Published: (2024)
by: Feng, Shangbin, et al.
Published: (2024)
Safety Mirage: How Spurious Correlations Undermine VLM Safety Fine-Tuning and Can Be Mitigated by Machine Unlearning
by: Chen, Yiwei, et al.
Published: (2025)
by: Chen, Yiwei, et al.
Published: (2025)
Can Large Language Models Integrate Spatial Data? Empirical Insights into Reasoning Strengths and Computational Weaknesses
by: Han, Bin, et al.
Published: (2025)
by: Han, Bin, et al.
Published: (2025)
MediQ: Question-Asking LLMs and a Benchmark for Reliable Interactive Clinical Reasoning
by: Li, Shuyue Stella, et al.
Published: (2024)
by: Li, Shuyue Stella, et al.
Published: (2024)
MMMG: a Comprehensive and Reliable Evaluation Suite for Multitask Multimodal Generation
by: Yao, Jihan, et al.
Published: (2025)
by: Yao, Jihan, et al.
Published: (2025)
Spurious Correlations in Concept Drift: Can Explanatory Interaction Help?
by: Lalletti, Cristiana, et al.
Published: (2024)
by: Lalletti, Cristiana, et al.
Published: (2024)
Stumbling Blocks: Stress Testing the Robustness of Machine-Generated Text Detectors Under Attacks
by: Wang, Yichen, et al.
Published: (2024)
by: Wang, Yichen, et al.
Published: (2024)
STOP: Structured On-Policy Pruning of Long-Form Reasoning in Low-Data Regimes
by: Xu, Chenjun, et al.
Published: (2026)
by: Xu, Chenjun, et al.
Published: (2026)
Do Language Models Mirror Human Confidence? Exploring Psychological Insights to Address Overconfidence in LLMs
by: Xu, Chenjun, et al.
Published: (2025)
by: Xu, Chenjun, et al.
Published: (2025)
Spurious Correlations and Beyond: Understanding and Mitigating Shortcut Learning in SDOH Extraction with Large Language Models
by: Sakib, Fardin Ahsan, et al.
Published: (2025)
by: Sakib, Fardin Ahsan, et al.
Published: (2025)
David helps Goliath: Inference-Time Collaboration Between Small Specialized and Large General Diffusion LMs
by: Han, Xiaochuang, et al.
Published: (2023)
by: Han, Xiaochuang, et al.
Published: (2023)
Mitigating Spurious Correlations for Self-supervised Recommendation
by: Lin, Xinyu, et al.
Published: (2022)
by: Lin, Xinyu, et al.
Published: (2022)
Spurious Rewards: Rethinking Training Signals in RLVR
by: Shao, Rulin, et al.
Published: (2025)
by: Shao, Rulin, et al.
Published: (2025)
Similar Items
-
Know Your Limits: A Survey of Abstention in Large Language Models
by: Wen, Bingbing, et al.
Published: (2024) -
Data Swarms: Optimizable Generation of Synthetic Evaluation Data
by: Feng, Shangbin, et al.
Published: (2025) -
ScienceMeter: Tracking Scientific Knowledge Updates in Language Models
by: Wang, Yike, et al.
Published: (2025) -
The Single-Multi Evolution Loop for Self-Improving Model Collaboration Systems
by: Feng, Shangbin, et al.
Published: (2026) -
Among Us: Measuring and Mitigating Malicious Contributions in Model Collaboration Systems
by: Yang, Ziyuan, et al.
Published: (2026)