Understanding and Mitigating Spurious Signal Amplification in Test-Time Reinforcement Learning for Math Reasoning
Fuente:
arXiv
Saved in:
| Main Authors: | Yu, Yongcan, He, Lingxiao, Liang, Jian, Guo, Kuangpu, Wang, Meng, Xie, Qianlong, Wang, Xingxing, He, Ran |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Reassessing the Role of Supervised Fine-Tuning: An Empirical Study in VLM Reasoning
by: Yu, Yongcan, et al.
Published: (2025)
by: Yu, Yongcan, et al.
Published: (2025)
Cooperative Pseudo Labeling for Unsupervised Federated Classification
by: Guo, Kuangpu, et al.
Published: (2025)
by: Guo, Kuangpu, et al.
Published: (2025)
Test-Time Immunization: A Universal Defense Framework Against Jailbreaks for (Multimodal) Large Language Models
by: Yu, Yongcan, et al.
Published: (2025)
by: Yu, Yongcan, et al.
Published: (2025)
STAMP: Outlier-Aware Test-Time Adaptation with Stable Memory Replay
by: Yu, Yongcan, et al.
Published: (2024)
by: Yu, Yongcan, et al.
Published: (2024)
Do MLLMs Really Understand Space? A Mathematical Reasoning Evaluation
by: Lu, Shuo, et al.
Published: (2026)
by: Lu, Shuo, et al.
Published: (2026)
A Comprehensive Survey on Trustworthiness in Reasoning with Large Language Models
by: Wang, Yanbo, et al.
Published: (2025)
by: Wang, Yanbo, et al.
Published: (2025)
Mitigating the Safety-utility Trade-off in LLM Alignment via Adaptive Safe Context Learning
by: Wang, Yanbo, et al.
Published: (2026)
by: Wang, Yanbo, et al.
Published: (2026)
Stay Unique, Stay Efficient: Preserving Model Personality in Multi-Task Merging
by: Guo, Kuangpu, et al.
Published: (2025)
by: Guo, Kuangpu, et al.
Published: (2025)
Exploring Vacant Classes in Label-Skewed Federated Learning
by: Guo, Kuangpu, et al.
Published: (2024)
by: Guo, Kuangpu, et al.
Published: (2024)
How to Train Your Deep Research Agent? Prompt, Reward, and Policy Optimization in Search-R1
by: Xu, Yinuo, et al.
Published: (2026)
by: Xu, Yinuo, et al.
Published: (2026)
Safe Offline Reinforcement Learning with Real-Time Budget Constraints
by: Lin, Qian, et al.
Published: (2023)
by: Lin, Qian, et al.
Published: (2023)
What If Consensus Lies? Selective-Complementary Reinforcement Learning at Test Time
by: Yan, Dong, et al.
Published: (2026)
by: Yan, Dong, et al.
Published: (2026)
An Enhanced Federated Prototype Learning Method under Domain Shift
by: Kuang, Liang, et al.
Published: (2024)
by: Kuang, Liang, et al.
Published: (2024)
Amplification Effects in Test-Time Reinforcement Learning: Safety and Reasoning Vulnerabilities
by: Khattar, Vanshaj, et al.
Published: (2026)
by: Khattar, Vanshaj, et al.
Published: (2026)
Personalized Federated Learning via Dual-Prompt Optimization and Cross Fusion
by: Zhang, Yuguang, et al.
Published: (2025)
by: Zhang, Yuguang, et al.
Published: (2025)
Off-Policy Primal-Dual Safe Reinforcement Learning
by: Wu, Zifan, et al.
Published: (2024)
by: Wu, Zifan, et al.
Published: (2024)
A Comprehensive Survey on Test-Time Adaptation under Distribution Shifts
by: Liang, Jian, et al.
Published: (2023)
by: Liang, Jian, et al.
Published: (2023)
TabularMath: Understanding Math Reasoning over Tables with Large Language Models
by: Tian, Shi-Yu, et al.
Published: (2025)
by: Tian, Shi-Yu, et al.
Published: (2025)
R-TPT: Improving Adversarial Robustness of Vision-Language Models through Test-Time Prompt Tuning
by: Sheng, Lijun, et al.
Published: (2025)
by: Sheng, Lijun, et al.
Published: (2025)
CodeReasoner: Enhancing the Code Reasoning Ability with Reinforcement Learning
by: Tang, Lingxiao, et al.
Published: (2025)
by: Tang, Lingxiao, et al.
Published: (2025)
MathScale: Scaling Instruction Tuning for Mathematical Reasoning
by: Tang, Zhengyang, et al.
Published: (2024)
by: Tang, Zhengyang, et al.
Published: (2024)
RL-MPCA: A Reinforcement Learning Based Multi-Phase Computation Allocation Approach for Recommender Systems
by: Zhou, Jiahong, et al.
Published: (2023)
by: Zhou, Jiahong, et al.
Published: (2023)
AUTO: Adaptive Outlier Optimization for Test-Time OOD Detection
by: Yang, Puning, et al.
Published: (2023)
by: Yang, Puning, et al.
Published: (2023)
Reasoning in Transformers -- Mitigating Spurious Correlations and Reasoning Shortcuts
by: Enström, Daniel, et al.
Published: (2024)
by: Enström, Daniel, et al.
Published: (2024)
The Illusion of Progress? A Critical Look at Test-Time Adaptation for Vision-Language Models
by: Sheng, Lijun, et al.
Published: (2025)
by: Sheng, Lijun, et al.
Published: (2025)
FITRep: Attention-Guided Item Representation via MLLMs
by: Zhang, Guoxiao, et al.
Published: (2025)
by: Zhang, Guoxiao, et al.
Published: (2025)
NFT: Bridging Supervised Learning and Reinforcement Learning in Math Reasoning
by: Chen, Huayu, et al.
Published: (2025)
by: Chen, Huayu, et al.
Published: (2025)
The Trap of Trajectory: Towards Understanding and Mitigating Spurious Correlations in Agentic Memory
by: Tang, Luoxi, et al.
Published: (2026)
by: Tang, Luoxi, et al.
Published: (2026)
Mitigating Spurious Correlations with Causal Logit Perturbation
by: Zhou, Xiaoling, et al.
Published: (2025)
by: Zhou, Xiaoling, et al.
Published: (2025)
Understanding and Mitigating Spurious Correlations in Text Classification with Neighborhood Analysis
by: Chew, Oscar, et al.
Published: (2023)
by: Chew, Oscar, et al.
Published: (2023)
HiBid: A Cross-Channel Constrained Bidding System with Budget Allocation by Hierarchical Offline Deep Reinforcement Learning
by: Wang, Hao, et al.
Published: (2023)
by: Wang, Hao, et al.
Published: (2023)
DeepResearch-Slice: Bridging the Retrieval-Utilization Gap via Explicit Text Slicing
by: Lu, Shuo, et al.
Published: (2025)
by: Lu, Shuo, et al.
Published: (2025)
STAPO: Stabilizing Reinforcement Learning for LLMs by Silencing Rare Spurious Tokens
by: Liu, Shiqi, et al.
Published: (2026)
by: Liu, Shiqi, et al.
Published: (2026)
Generative Large-Scale Pre-trained Models for Automated Ad Bidding Optimization
by: Lei, Yu, et al.
Published: (2025)
by: Lei, Yu, et al.
Published: (2025)
Mitigating Spurious Correlations for Self-supervised Recommendation
by: Lin, Xinyu, et al.
Published: (2022)
by: Lin, Xinyu, et al.
Published: (2022)
Precedent-Informed Reasoning: Mitigating Overthinking in Large Reasoning Models via Test-Time Precedent Learning
by: Wang, Qianyue, et al.
Published: (2026)
by: Wang, Qianyue, et al.
Published: (2026)
Towards Eliminating Hard Label Constraints in Gradient Inversion Attacks
by: Wang, Yanbo, et al.
Published: (2024)
by: Wang, Yanbo, et al.
Published: (2024)
VideoChat-R1.5: Visual Test-Time Scaling to Reinforce Multimodal Reasoning by Iterative Perception
by: Yan, Ziang, et al.
Published: (2025)
by: Yan, Ziang, et al.
Published: (2025)
Spurious Correlations and Beyond: Understanding and Mitigating Shortcut Learning in SDOH Extraction with Large Language Models
by: Sakib, Fardin Ahsan, et al.
Published: (2025)
by: Sakib, Fardin Ahsan, et al.
Published: (2025)
Elastic Representation: Mitigating Spurious Correlations for Group Robustness
by: Wen, Tao, et al.
Published: (2025)
by: Wen, Tao, et al.
Published: (2025)
Similar Items
-
Reassessing the Role of Supervised Fine-Tuning: An Empirical Study in VLM Reasoning
by: Yu, Yongcan, et al.
Published: (2025) -
Cooperative Pseudo Labeling for Unsupervised Federated Classification
by: Guo, Kuangpu, et al.
Published: (2025) -
Test-Time Immunization: A Universal Defense Framework Against Jailbreaks for (Multimodal) Large Language Models
by: Yu, Yongcan, et al.
Published: (2025) -
STAMP: Outlier-Aware Test-Time Adaptation with Stable Memory Replay
by: Yu, Yongcan, et al.
Published: (2024) -
Do MLLMs Really Understand Space? A Mathematical Reasoning Evaluation
by: Lu, Shuo, et al.
Published: (2026)