Spurious Rewards Paradox: Mechanistically Understanding How RLVR Activates Memorization Shortcuts in LLMs
Fuente:
arXiv
Saved in:
| Main Authors: | Yan, Lecheng, Li, Ruizhe, Chen, Guanhua, Li, Qing, Geng, Jiahui, Li, Wenxi, Wang, Vincent, Lee, Chris |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Trust No Tool: Evaluating and Defending LLM Agents under Untrusted Tool Feedback
by: Yan, Lecheng, et al.
Published: (2026)
by: Yan, Lecheng, et al.
Published: (2026)
Exploration vs Exploitation: Rethinking RLVR through Clipping, Entropy, and Spurious Reward
by: Chen, Peter, et al.
Published: (2025)
by: Chen, Peter, et al.
Published: (2025)
Unveiling Over-Memorization in Finetuning LLMs for Reasoning Tasks
by: Ruan, Zhiwen, et al.
Published: (2025)
by: Ruan, Zhiwen, et al.
Published: (2025)
Spurious Rewards: Rethinking Training Signals in RLVR
by: Shao, Rulin, et al.
Published: (2025)
by: Shao, Rulin, et al.
Published: (2025)
Reasoning in Transformers -- Mitigating Spurious Correlations and Reasoning Shortcuts
by: Enström, Daniel, et al.
Published: (2024)
by: Enström, Daniel, et al.
Published: (2024)
Spurious Correlations and Beyond: Understanding and Mitigating Shortcut Learning in SDOH Extraction with Large Language Models
by: Sakib, Fardin Ahsan, et al.
Published: (2025)
by: Sakib, Fardin Ahsan, et al.
Published: (2025)
FairSteer: Inference Time Debiasing for LLMs with Dynamic Activation Steering
by: Li, Yichen, et al.
Published: (2025)
by: Li, Yichen, et al.
Published: (2025)
Mitigating Memorization in LLMs using Activation Steering
by: Suri, Manan, et al.
Published: (2025)
by: Suri, Manan, et al.
Published: (2025)
Knowledge-to-Verification: Exploring RLVR for LLMs in Knowledge-Intensive Domains
by: Yuan, Zhonghang, et al.
Published: (2026)
by: Yuan, Zhonghang, et al.
Published: (2026)
Short-circuiting Shortcuts: Mechanistic Investigation of Shortcuts in Text Classification
by: Eshuijs, Leon, et al.
Published: (2025)
by: Eshuijs, Leon, et al.
Published: (2025)
Memorization or Reasoning? Exploring the Idiom Understanding of LLMs
by: Kim, Jisu, et al.
Published: (2025)
by: Kim, Jisu, et al.
Published: (2025)
Positional Fragility in LLMs: How Offset Effects Reshape Our Understanding of Memorization Risks
by: Xu, Yixuan, et al.
Published: (2025)
by: Xu, Yixuan, et al.
Published: (2025)
The Reliability Paradox: Exploring How Shortcut Learning Undermines Language Model Calibration
by: Bihani, Geetanjali, et al.
Published: (2024)
by: Bihani, Geetanjali, et al.
Published: (2024)
Understanding Verbatim Memorization in LLMs Through Circuit Discovery
by: Lasy, Ilya, et al.
Published: (2025)
by: Lasy, Ilya, et al.
Published: (2025)
Evaluating the Effectiveness of Linguistic Knowledge in Pretrained Language Models: A Case Study of Universal Dependencies
by: Li, Wenxi
Published: (2025)
by: Li, Wenxi
Published: (2025)
Large Language Models as Code Executors: An Exploratory Study
by: Lyu, Chenyang, et al.
Published: (2024)
by: Lyu, Chenyang, et al.
Published: (2024)
Rewards as Labels: Revisiting RLVR from a Classification Perspective
by: Zhai, Zepeng, et al.
Published: (2026)
by: Zhai, Zepeng, et al.
Published: (2026)
When Bias Pretends to Be Truth: How Spurious Correlations Undermine Hallucination Detection in LLMs
by: Wang, Shaowen, et al.
Published: (2025)
by: Wang, Shaowen, et al.
Published: (2025)
Internal Activation Revision: Safeguarding Vision Language Models Without Parameter Update
by: Li, Qing, et al.
Published: (2025)
by: Li, Qing, et al.
Published: (2025)
LongRLVR: Long-Context Reinforcement Learning Requires Verifiable Context Rewards
by: Chen, Guanzheng, et al.
Published: (2026)
by: Chen, Guanzheng, et al.
Published: (2026)
Do LLMs Really Memorize Personally Identifiable Information? Revisiting PII Leakage with a Cue-Controlled Memorization Framework
by: Luo, Xiaoyu, et al.
Published: (2026)
by: Luo, Xiaoyu, et al.
Published: (2026)
Knowing the Facts but Choosing the Shortcut: Understanding How Large Language Models Compare Entities
by: Lehmann, Hans Hergen, et al.
Published: (2025)
by: Lehmann, Hans Hergen, et al.
Published: (2025)
SAGE: Shaping Anchors for Guided Exploration in RLVR of LLMs
by: Lee, Chanuk, et al.
Published: (2026)
by: Lee, Chanuk, et al.
Published: (2026)
Understanding How CodeLLMs (Mis)Predict Types with Activation Steering
by: Lucchetti, Francesca, et al.
Published: (2024)
by: Lucchetti, Francesca, et al.
Published: (2024)
How Far Can Unsupervised RLVR Scale LLM Training?
by: He, Bingxiang, et al.
Published: (2026)
by: He, Bingxiang, et al.
Published: (2026)
The Devil Is in the Details: Tackling Unimodal Spurious Correlations for Generalizable Multimodal Reward Models
by: Li, Zichao, et al.
Published: (2025)
by: Li, Zichao, et al.
Published: (2025)
Shared Path: Unraveling Memorization in Multilingual LLMs through Language Similarities
by: Luo, Xiaoyu, et al.
Published: (2025)
by: Luo, Xiaoyu, et al.
Published: (2025)
Memorization vs. Reasoning: Updating LLMs with New Knowledge
by: Li, Aochong Oliver, et al.
Published: (2025)
by: Li, Aochong Oliver, et al.
Published: (2025)
Skywork-Reward: Bag of Tricks for Reward Modeling in LLMs
by: Liu, Chris Yuhao, et al.
Published: (2024)
by: Liu, Chris Yuhao, et al.
Published: (2024)
Continual Memorization of Factoids in Language Models
by: Chen, Howard, et al.
Published: (2024)
by: Chen, Howard, et al.
Published: (2024)
Memorization $\neq$ Understanding: Do Large Language Models Have the Ability of Scenario Cognition?
by: Ma, Boxiang, et al.
Published: (2025)
by: Ma, Boxiang, et al.
Published: (2025)
How Emotion Shapes the Behavior of LLMs and Agents: A Mechanistic Study
by: Sun, Moran, et al.
Published: (2026)
by: Sun, Moran, et al.
Published: (2026)
Rewarding How Models Think Pedagogically: Integrating Pedagogical Reasoning and Thinking Rewards for LLMs in Education
by: Lee, Unggi, et al.
Published: (2026)
by: Lee, Unggi, et al.
Published: (2026)
STAPO: Stabilizing Reinforcement Learning for LLMs by Silencing Rare Spurious Tokens
by: Liu, Shiqi, et al.
Published: (2026)
by: Liu, Shiqi, et al.
Published: (2026)
Document Reconstruction Unlocks Scalable Long-Context RLVR
by: Xiao, Yao, et al.
Published: (2026)
by: Xiao, Yao, et al.
Published: (2026)
ImPart: Importance-Aware Delta-Sparsification for Improved Model Compression and Merging in LLMs
by: Yang, Yan, et al.
Published: (2025)
by: Yang, Yan, et al.
Published: (2025)
Benchmarking Chinese Commonsense Reasoning of LLMs: From Chinese-Specifics to Reasoning-Memorization Correlations
by: Sun, Jiaxing, et al.
Published: (2024)
by: Sun, Jiaxing, et al.
Published: (2024)
HD-NDEs: Neural Differential Equations for Hallucination Detection in LLMs
by: Li, Qing, et al.
Published: (2025)
by: Li, Qing, et al.
Published: (2025)
Be like a Goldfish, Don't Memorize! Mitigating Memorization in Generative LLMs
by: Hans, Abhimanyu, et al.
Published: (2024)
by: Hans, Abhimanyu, et al.
Published: (2024)
Towards Bridging the Reward-Generation Gap in Direct Alignment Algorithms
by: Xiao, Zeguan, et al.
Published: (2025)
by: Xiao, Zeguan, et al.
Published: (2025)
Similar Items
-
Trust No Tool: Evaluating and Defending LLM Agents under Untrusted Tool Feedback
by: Yan, Lecheng, et al.
Published: (2026) -
Exploration vs Exploitation: Rethinking RLVR through Clipping, Entropy, and Spurious Reward
by: Chen, Peter, et al.
Published: (2025) -
Unveiling Over-Memorization in Finetuning LLMs for Reasoning Tasks
by: Ruan, Zhiwen, et al.
Published: (2025) -
Spurious Rewards: Rethinking Training Signals in RLVR
by: Shao, Rulin, et al.
Published: (2025) -
Reasoning in Transformers -- Mitigating Spurious Correlations and Reasoning Shortcuts
by: Enström, Daniel, et al.
Published: (2024)