Shortcut Mitigation via Spurious-Positive Samples
Fuente:
arXiv
Saved in:
| Main Authors: | Le, Phuong Quynh, Schlötterer, Jörg, Sadiya, Sari, Roig, Gemma, Seifert, Christin |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Efficient Unsupervised Shortcut Learning Detection and Mitigation in Transformers
by: Kuhn, Lukas, et al.
Published: (2025)
by: Kuhn, Lukas, et al.
Published: (2025)
Out of Spuriousity: Improving Robustness to Spurious Correlations without Group Annotations
by: Le, Phuong Quynh, et al.
Published: (2024)
by: Le, Phuong Quynh, et al.
Published: (2024)
An XAI-based Analysis of Shortcut Learning in Neural Networks
by: Le, Phuong Quynh, et al.
Published: (2025)
by: Le, Phuong Quynh, et al.
Published: (2025)
Is Last Layer Re-Training Truly Sufficient for Robustness to Spurious Correlations?
by: Le, Phuong Quynh, et al.
Published: (2023)
by: Le, Phuong Quynh, et al.
Published: (2023)
Invariant Learning with Annotation-free Environments
by: Le, Phuong Quynh, et al.
Published: (2025)
by: Le, Phuong Quynh, et al.
Published: (2025)
Towards Interpretable Deep Neural Networks for Tabular Data
by: Elhadri, Khawla, et al.
Published: (2025)
by: Elhadri, Khawla, et al.
Published: (2025)
XNNTab -- Interpretable Neural Networks for Tabular Data using Sparse Autoencoders
by: Elhadri, Khawla, et al.
Published: (2025)
by: Elhadri, Khawla, et al.
Published: (2025)
Persuasion Tokens for Editing Factual Knowledge in LLMs
by: Youssef, Paul, et al.
Published: (2026)
by: Youssef, Paul, et al.
Published: (2026)
Different Algorithms (Might) Uncover Different Patterns: A Brain-Age Prediction Case Study
by: Ettling, Tobias, et al.
Published: (2024)
by: Ettling, Tobias, et al.
Published: (2024)
A Second Look on BASS -- Boosting Abstractive Summarization with Unified Semantic Graphs -- A Replication Study
by: Koraş, Osman Alperen, et al.
Published: (2024)
by: Koraş, Osman Alperen, et al.
Published: (2024)
This looks like what? Challenges and Future Research Directions for Part-Prototype Models
by: Elhadri, Khawla, et al.
Published: (2025)
by: Elhadri, Khawla, et al.
Published: (2025)
Reasoning in Transformers -- Mitigating Spurious Correlations and Reasoning Shortcuts
by: Enström, Daniel, et al.
Published: (2024)
by: Enström, Daniel, et al.
Published: (2024)
Navigating Shortcuts, Spurious Correlations, and Confounders: From Origins via Detection to Mitigation
by: Steinmann, David, et al.
Published: (2024)
by: Steinmann, David, et al.
Published: (2024)
Limited but consistent gains in adversarial robustness by co-training object recognition models with human EEG
by: Guo, Manshan, et al.
Published: (2024)
by: Guo, Manshan, et al.
Published: (2024)
Shortcut to Nowhere: Demystifying Deep Spurious Regression
by: Xu, Guanrong, et al.
Published: (2026)
by: Xu, Guanrong, et al.
Published: (2026)
Cognitive Neural Architecture Search Reveals Hierarchical Entailment
by: Kuhn, Lukas, et al.
Published: (2025)
by: Kuhn, Lukas, et al.
Published: (2025)
The Queen of England is not England's Queen: On the Lack of Factual Coherency in PLMs
by: Youssef, Paul, et al.
Published: (2024)
by: Youssef, Paul, et al.
Published: (2024)
Enhancing Fact Retrieval in PLMs through Truthfulness
by: Youssef, Paul, et al.
Published: (2024)
by: Youssef, Paul, et al.
Published: (2024)
Models Know Their Shortcuts: Deployment-Time Shortcut Mitigation
by: Li, Jiayi, et al.
Published: (2026)
by: Li, Jiayi, et al.
Published: (2026)
Mitigating Spurious Correlations via Disagreement Probability
by: Han, Hyeonggeun, et al.
Published: (2024)
by: Han, Hyeonggeun, et al.
Published: (2024)
Guiding LLMs to Generate High-Fidelity and High-Quality Counterfactual Explanations for Text Classification
by: Nguyen, Van Bach, et al.
Published: (2025)
by: Nguyen, Van Bach, et al.
Published: (2025)
From Black Boxes to Conversations: Incorporating XAI in a Conversational Agent
by: Nguyen, Van Bach, et al.
Published: (2022)
by: Nguyen, Van Bach, et al.
Published: (2022)
CEval: A Benchmark for Evaluating Counterfactual Text Generation
by: Nguyen, Van Bach, et al.
Published: (2024)
by: Nguyen, Van Bach, et al.
Published: (2024)
Position: Editing Large Language Models Poses Serious Safety Risks
by: Youssef, Paul, et al.
Published: (2025)
by: Youssef, Paul, et al.
Published: (2025)
Mitigating Spurious Correlations in NLI via LLM-Synthesized Counterfactuals and Dynamic Balanced Sampling
by: Jaimes, Christopher Román
Published: (2025)
by: Jaimes, Christopher Román
Published: (2025)
Mitigating Spurious Correlations in LLMs via Causality-Aware Post-Training
by: Gui, Shurui, et al.
Published: (2025)
by: Gui, Shurui, et al.
Published: (2025)
Spurious Rewards Paradox: Mechanistically Understanding How RLVR Activates Memorization Shortcuts in LLMs
by: Yan, Lecheng, et al.
Published: (2026)
by: Yan, Lecheng, et al.
Published: (2026)
Norm-Hierarchy Transitions in Representation Learning: When and Why Neural Networks Abandon Shortcuts
by: Khanh, Truong Xuan, et al.
Published: (2026)
by: Khanh, Truong Xuan, et al.
Published: (2026)
Mitigating Spurious Correlation via Distributionally Robust Learning with Hierarchical Ambiguity Sets
by: Jo, Sung Ho, et al.
Published: (2025)
by: Jo, Sung Ho, et al.
Published: (2025)
Elastic Representation: Mitigating Spurious Correlations for Group Robustness
by: Wen, Tao, et al.
Published: (2025)
by: Wen, Tao, et al.
Published: (2025)
Comparative Explanations: Explanation Guided Decision Making for Human-in-the-Loop Preference Selection
by: Chakraborty, Tanmay, et al.
Published: (2025)
by: Chakraborty, Tanmay, et al.
Published: (2025)
Explainable Bayesian Optimization
by: Chakraborty, Tanmay, et al.
Published: (2024)
by: Chakraborty, Tanmay, et al.
Published: (2024)
Has this Fact been Edited? Detecting Knowledge Edits in Language Models
by: Youssef, Paul, et al.
Published: (2024)
by: Youssef, Paul, et al.
Published: (2024)
How to Make LLMs Forget: On Reversing In-Context Knowledge Edits
by: Youssef, Paul, et al.
Published: (2024)
by: Youssef, Paul, et al.
Published: (2024)
Tracing and Reversing Edits in LLMs
by: Youssef, Paul, et al.
Published: (2025)
by: Youssef, Paul, et al.
Published: (2025)
NeuronTune: Towards Self-Guided Spurious Bias Mitigation
by: Zheng, Guangtao, et al.
Published: (2025)
by: Zheng, Guangtao, et al.
Published: (2025)
One Mask to Rule Them All: On Hidden Facts after Editing and How to Find Them
by: Holmov, Ali, et al.
Published: (2026)
by: Holmov, Ali, et al.
Published: (2026)
Position: An Inner Interpretability Framework for AI Inspired by Lessons from Cognitive Neuroscience
by: Vilas, Martina G., et al.
Published: (2024)
by: Vilas, Martina G., et al.
Published: (2024)
Optimizing DDPM Sampling with Shortcut Fine-Tuning
by: Fan, Ying, et al.
Published: (2023)
by: Fan, Ying, et al.
Published: (2023)
Mitigating Shortcut Learning with InterpoLated Learning
by: Korakakis, Michalis, et al.
Published: (2025)
by: Korakakis, Michalis, et al.
Published: (2025)
Similar Items
-
Efficient Unsupervised Shortcut Learning Detection and Mitigation in Transformers
by: Kuhn, Lukas, et al.
Published: (2025) -
Out of Spuriousity: Improving Robustness to Spurious Correlations without Group Annotations
by: Le, Phuong Quynh, et al.
Published: (2024) -
An XAI-based Analysis of Shortcut Learning in Neural Networks
by: Le, Phuong Quynh, et al.
Published: (2025) -
Is Last Layer Re-Training Truly Sufficient for Robustness to Spurious Correlations?
by: Le, Phuong Quynh, et al.
Published: (2023) -
Invariant Learning with Annotation-free Environments
by: Le, Phuong Quynh, et al.
Published: (2025)