Short-circuiting Shortcuts: Mechanistic Investigation of Shortcuts in Text Classification
Fuente:
arXiv
Salvato in:
| Autori principali: | Eshuijs, Leon, Wang, Shihan, Fokkens, Antske |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Balancing the Scales: Reinforcement Learning for Fair Classification
di: Eshuijs, Leon, et al.
Pubblicazione: (2024)
di: Eshuijs, Leon, et al.
Pubblicazione: (2024)
Safety Training Modulates Harmful Misalignment Under On-Policy RL, But Direction Depends on Environment Design
di: Eshuijs, Leon, et al.
Pubblicazione: (2026)
di: Eshuijs, Leon, et al.
Pubblicazione: (2026)
Navigating the Shortcut Maze: A Comprehensive Analysis of Shortcut Learning in Text Classification by Language Models
di: Zhou, Yuqing, et al.
Pubblicazione: (2024)
di: Zhou, Yuqing, et al.
Pubblicazione: (2024)
Spurious Rewards Paradox: Mechanistically Understanding How RLVR Activates Memorization Shortcuts in LLMs
di: Yan, Lecheng, et al.
Pubblicazione: (2026)
di: Yan, Lecheng, et al.
Pubblicazione: (2026)
Do LLMs Overcome Shortcut Learning? An Evaluation of Shortcut Challenges in Large Language Models
di: Yuan, Yu, et al.
Pubblicazione: (2024)
di: Yuan, Yu, et al.
Pubblicazione: (2024)
Investigating the Robustness of Modelling Decisions for Few-Shot Cross-Topic Stance Detection: A Preregistered Study
di: Reuver, Myrthe, et al.
Pubblicazione: (2024)
di: Reuver, Myrthe, et al.
Pubblicazione: (2024)
Reasoning in Transformers -- Mitigating Spurious Correlations and Reasoning Shortcuts
di: Enström, Daniel, et al.
Pubblicazione: (2024)
di: Enström, Daniel, et al.
Pubblicazione: (2024)
Control the Temperature: Selective Sampling for Diverse and High-Quality LLM Outputs
di: Troshin, Sergey, et al.
Pubblicazione: (2025)
di: Troshin, Sergey, et al.
Pubblicazione: (2025)
Seeing to Generalize: How Visual Data Corrects Binding Shortcuts
di: Buzeta, Nicolas, et al.
Pubblicazione: (2026)
di: Buzeta, Nicolas, et al.
Pubblicazione: (2026)
MMLU-Pro+: Evaluating Higher-Order Reasoning and Shortcut Learning in LLMs
di: Taghanaki, Saeid Asgari, et al.
Pubblicazione: (2024)
di: Taghanaki, Saeid Asgari, et al.
Pubblicazione: (2024)
Learning Shortcuts: On the Misleading Promise of NLU in Language Models
di: Bihani, Geetanjali, et al.
Pubblicazione: (2024)
di: Bihani, Geetanjali, et al.
Pubblicazione: (2024)
Shortcut-connected Expert Parallelism for Accelerating Mixture-of-Experts
di: Cai, Weilin, et al.
Pubblicazione: (2024)
di: Cai, Weilin, et al.
Pubblicazione: (2024)
From Shortcuts to Triggers: Backdoor Defense with Denoised PoE
di: Liu, Qin, et al.
Pubblicazione: (2023)
di: Liu, Qin, et al.
Pubblicazione: (2023)
DefVerify: Do Hate Speech Models Reflect Their Dataset's Definition?
di: Khurana, Urja, et al.
Pubblicazione: (2024)
di: Khurana, Urja, et al.
Pubblicazione: (2024)
On the Low-Rank Parametrization of Reward Models for Controlled Language Generation
di: Troshin, Sergey, et al.
Pubblicazione: (2024)
di: Troshin, Sergey, et al.
Pubblicazione: (2024)
The Reliability Paradox: Exploring How Shortcut Learning Undermines Language Model Calibration
di: Bihani, Geetanjali, et al.
Pubblicazione: (2024)
di: Bihani, Geetanjali, et al.
Pubblicazione: (2024)
Learning from Sufficient Rationales: Analysing the Relationship Between Explanation Faithfulness and Token-level Regularisation Strategies
di: Kamp, Jonathan, et al.
Pubblicazione: (2025)
di: Kamp, Jonathan, et al.
Pubblicazione: (2025)
Improving Causal Interventions in Amnesic Probing with Mean Projection or LEACE
di: Dobrzeniecka, Alicja, et al.
Pubblicazione: (2025)
di: Dobrzeniecka, Alicja, et al.
Pubblicazione: (2025)
Not Eliminate but Aggregate: Post-Hoc Control over Mixture-of-Experts to Address Shortcut Shifts in Natural Language Understanding
di: Honda, Ukyo, et al.
Pubblicazione: (2024)
di: Honda, Ukyo, et al.
Pubblicazione: (2024)
The Role of Syntactic Span Preferences in Post-Hoc Explanation Disagreement
di: Kamp, Jonathan, et al.
Pubblicazione: (2024)
di: Kamp, Jonathan, et al.
Pubblicazione: (2024)
Adaptive Large Language Models By Layerwise Attention Shortcuts
di: Verma, Prateek, et al.
Pubblicazione: (2024)
di: Verma, Prateek, et al.
Pubblicazione: (2024)
CURE: Controlled Unlearning for Robust Embeddings -- Mitigating Conceptual Shortcuts in Pre-Trained Language Models
di: Kocak, Aysenur, et al.
Pubblicazione: (2025)
di: Kocak, Aysenur, et al.
Pubblicazione: (2025)
Investigating the Influence of Prompt-Specific Shortcuts in AI Generated Text Detection
di: Park, Choonghyun, et al.
Pubblicazione: (2024)
di: Park, Choonghyun, et al.
Pubblicazione: (2024)
On the Foundations of Shortcut Learning
di: Hermann, Katherine L., et al.
Pubblicazione: (2023)
di: Hermann, Katherine L., et al.
Pubblicazione: (2023)
Asking a Language Model for Diverse Responses
di: Troshin, Sergey, et al.
Pubblicazione: (2025)
di: Troshin, Sergey, et al.
Pubblicazione: (2025)
Crowd-Calibrator: Can Annotator Disagreement Inform Calibration in Subjective Tasks?
di: Khurana, Urja, et al.
Pubblicazione: (2024)
di: Khurana, Urja, et al.
Pubblicazione: (2024)
Models Know Their Shortcuts: Deployment-Time Shortcut Mitigation
di: Li, Jiayi, et al.
Pubblicazione: (2026)
di: Li, Jiayi, et al.
Pubblicazione: (2026)
Persian Slang Text Conversion to Formal and Deep Learning of Persian Short Texts on Social Media for Sentiment Classification
di: Khazeni, Mohsen, et al.
Pubblicazione: (2024)
di: Khazeni, Mohsen, et al.
Pubblicazione: (2024)
ShortcutProbe: Probing Prediction Shortcuts for Learning Robust Models
di: Zheng, Guangtao, et al.
Pubblicazione: (2025)
di: Zheng, Guangtao, et al.
Pubblicazione: (2025)
Shortcut Learning Susceptibility in Vision Classifiers
di: Suhail, Pirzada, et al.
Pubblicazione: (2025)
di: Suhail, Pirzada, et al.
Pubblicazione: (2025)
L3Cube-MahaNews: News-based Short Text and Long Document Classification Datasets in Marathi
di: Mittal, Saloni, et al.
Pubblicazione: (2024)
di: Mittal, Saloni, et al.
Pubblicazione: (2024)
L3Cube-IndicNews: News-based Short Text and Long Document Classification Datasets in Indic Languages
di: Mirashi, Aishwarya, et al.
Pubblicazione: (2024)
di: Mirashi, Aishwarya, et al.
Pubblicazione: (2024)
Shortcuts Arising from Contrast: Effective and Covert Clean-Label Attacks in Prompt-Based Learning
di: Xie, Xiaopeng, et al.
Pubblicazione: (2024)
di: Xie, Xiaopeng, et al.
Pubblicazione: (2024)
One Step Diffusion via Shortcut Models
di: Frans, Kevin, et al.
Pubblicazione: (2024)
di: Frans, Kevin, et al.
Pubblicazione: (2024)
The Readout Shortcut: Positional Number Copying Dominates Arithmetic CoT Readout in Small Language Models
di: Liu, Ming
Pubblicazione: (2026)
di: Liu, Ming
Pubblicazione: (2026)
Mechanistic origins of catastrophic forgetting: why RL preserves circuits better than SFT?
di: Nunez, Jeanmely Rojas, et al.
Pubblicazione: (2026)
di: Nunez, Jeanmely Rojas, et al.
Pubblicazione: (2026)
Generative Classifiers Avoid Shortcut Solutions
di: Li, Alexander C., et al.
Pubblicazione: (2025)
di: Li, Alexander C., et al.
Pubblicazione: (2025)
Bias in the Shadows: Explore Shortcuts in Encrypted Network Traffic Classification
di: Wang, Chuyi, et al.
Pubblicazione: (2026)
di: Wang, Chuyi, et al.
Pubblicazione: (2026)
Stochastic Adversarial Networks for Multi-Domain Text Classification
di: Wang, Xu, et al.
Pubblicazione: (2024)
di: Wang, Xu, et al.
Pubblicazione: (2024)
Understand the Effectiveness of Shortcuts through the Lens of DCA
di: Sun, Youran, et al.
Pubblicazione: (2024)
di: Sun, Youran, et al.
Pubblicazione: (2024)
Documenti analoghi
-
Balancing the Scales: Reinforcement Learning for Fair Classification
di: Eshuijs, Leon, et al.
Pubblicazione: (2024) -
Safety Training Modulates Harmful Misalignment Under On-Policy RL, But Direction Depends on Environment Design
di: Eshuijs, Leon, et al.
Pubblicazione: (2026) -
Navigating the Shortcut Maze: A Comprehensive Analysis of Shortcut Learning in Text Classification by Language Models
di: Zhou, Yuqing, et al.
Pubblicazione: (2024) -
Spurious Rewards Paradox: Mechanistically Understanding How RLVR Activates Memorization Shortcuts in LLMs
di: Yan, Lecheng, et al.
Pubblicazione: (2026) -
Do LLMs Overcome Shortcut Learning? An Evaluation of Shortcut Challenges in Large Language Models
di: Yuan, Yu, et al.
Pubblicazione: (2024)