Monitorability as a Free Gift: How RLVR Spontaneously Aligns Reasoning
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Xiong, Zidi, Chen, Shan, Lakkaraju, Himabindu |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Measuring the Faithfulness of Thinking Drafts in Large Reasoning Models
von: Xiong, Zidi, et al.
Veröffentlicht: (2025)
von: Xiong, Zidi, et al.
Veröffentlicht: (2025)
Learning Recourse Costs from Pairwise Feature Comparisons
von: Rawal, Kaivalya, et al.
Veröffentlicht: (2024)
von: Rawal, Kaivalya, et al.
Veröffentlicht: (2024)
In-Context Unlearning: Language Models as Few Shot Unlearners
von: Pawelczyk, Martin, et al.
Veröffentlicht: (2023)
von: Pawelczyk, Martin, et al.
Veröffentlicht: (2023)
Towards Unified Attribution in Explainable AI, Data-Centric AI, and Mechanistic Interpretability
von: Zhang, Shichang, et al.
Veröffentlicht: (2025)
von: Zhang, Shichang, et al.
Veröffentlicht: (2025)
Who Gets Credit or Blame? Attributing Accountability in Modern AI Systems
von: Zhang, Shichang, et al.
Veröffentlicht: (2025)
von: Zhang, Shichang, et al.
Veröffentlicht: (2025)
Operationalizing the Blueprint for an AI Bill of Rights: Recommendations for Practitioners, Researchers, and Policy Makers
von: Oesterling, Alex, et al.
Veröffentlicht: (2024)
von: Oesterling, Alex, et al.
Veröffentlicht: (2024)
Generalized Group Data Attribution
von: Ley, Dan, et al.
Veröffentlicht: (2024)
von: Ley, Dan, et al.
Veröffentlicht: (2024)
Generalizing Trust: Weak-to-Strong Trustworthiness in Language Models
von: Pawelczyk, Martin, et al.
Veröffentlicht: (2024)
von: Pawelczyk, Martin, et al.
Veröffentlicht: (2024)
Evaluating Adversarial Robustness of Concept Representations in Sparse Autoencoders
von: Li, Aaron J., et al.
Veröffentlicht: (2025)
von: Li, Aaron J., et al.
Veröffentlicht: (2025)
The Disagreement Problem in Explainable Machine Learning: A Practitioner's Perspective
von: Krishna, Satyapriya, et al.
Veröffentlicht: (2022)
von: Krishna, Satyapriya, et al.
Veröffentlicht: (2022)
Data Poisoning Attacks on Off-Policy Policy Evaluation Methods
von: Lobo, Elita, et al.
Veröffentlicht: (2024)
von: Lobo, Elita, et al.
Veröffentlicht: (2024)
In-Context Explainers: Harnessing LLMs for Explaining Black Box Models
von: Kroeger, Nicholas, et al.
Veröffentlicht: (2023)
von: Kroeger, Nicholas, et al.
Veröffentlicht: (2023)
On the Direction of RLVR Updates for LLM Reasoning: Identification and Exploitation
von: Huang, Kexin, et al.
Veröffentlicht: (2026)
von: Huang, Kexin, et al.
Veröffentlicht: (2026)
How Post-Training Reshapes LLMs: A Mechanistic View on Knowledge, Truthfulness, Refusal, and Confidence
von: Du, Hongzhe, et al.
Veröffentlicht: (2025)
von: Du, Hongzhe, et al.
Veröffentlicht: (2025)
How Memory Management Impacts LLM Agents: An Empirical Study of Experience-Following Behavior
von: Xiong, Zidi, et al.
Veröffentlicht: (2025)
von: Xiong, Zidi, et al.
Veröffentlicht: (2025)
Quantile Advantage Estimation: Stabilizing RLVR for LLM Reasoning
von: Wu, Junkang, et al.
Veröffentlicht: (2025)
von: Wu, Junkang, et al.
Veröffentlicht: (2025)
Temporal Sparse Autoencoders: Leveraging the Sequential Nature of Language for Interpretability
von: Bhalla, Usha, et al.
Veröffentlicht: (2025)
von: Bhalla, Usha, et al.
Veröffentlicht: (2025)
Follow My Instruction and Spill the Beans: Scalable Data Extraction from Retrieval-Augmented Generation Systems
von: Qi, Zhenting, et al.
Veröffentlicht: (2024)
von: Qi, Zhenting, et al.
Veröffentlicht: (2024)
Controllable Exploration in Hybrid-Policy RLVR for Multi-Modal Reasoning
von: Huang, Zhuoxu, et al.
Veröffentlicht: (2026)
von: Huang, Zhuoxu, et al.
Veröffentlicht: (2026)
OpenXAI: Towards a Transparent Evaluation of Model Explanations
von: Agarwal, Chirag, et al.
Veröffentlicht: (2022)
von: Agarwal, Chirag, et al.
Veröffentlicht: (2022)
Generalization of RLVR Using Causal Reasoning as a Testbed
von: Lu, Brian, et al.
Veröffentlicht: (2025)
von: Lu, Brian, et al.
Veröffentlicht: (2025)
Adaptive Negative Reinforcement for LLM Reasoning:Dynamically Balancing Correction and Diversity in RLVR
von: Ingle, Yash, et al.
Veröffentlicht: (2026)
von: Ingle, Yash, et al.
Veröffentlicht: (2026)
Depth-Breadth Synergy in RLVR: Unlocking LLM Reasoning Gains with Adaptive Exploration
von: Yang, Zhicheng, et al.
Veröffentlicht: (2025)
von: Yang, Zhicheng, et al.
Veröffentlicht: (2025)
RLVR without Ineffective Samples: Group Prioritized Off-Policy Optimization for LLM Reasoning
von: Mao, Yixiu, et al.
Veröffentlicht: (2026)
von: Mao, Yixiu, et al.
Veröffentlicht: (2026)
Shorter but not Worse: Frugal Reasoning via Easy Samples as Length Regularizers in Math RLVR
von: Bounhar, Abdelaziz, et al.
Veröffentlicht: (2025)
von: Bounhar, Abdelaziz, et al.
Veröffentlicht: (2025)
Quantifying Empirical Compute-Supervision Tradeoffs in RLVR
von: Mitsuhashi, Ryo, et al.
Veröffentlicht: (2026)
von: Mitsuhashi, Ryo, et al.
Veröffentlicht: (2026)
Open-Medical-R1: How to Choose Data for RLVR Training at Medicine Domain
von: Qiu, Zhongxi, et al.
Veröffentlicht: (2025)
von: Qiu, Zhongxi, et al.
Veröffentlicht: (2025)
Certifying LLM Safety against Adversarial Prompting
von: Kumar, Aounon, et al.
Veröffentlicht: (2023)
von: Kumar, Aounon, et al.
Veröffentlicht: (2023)
JudgeRLVR: Judge First, Generate Second for Efficient Reasoning
von: Duo, Jiangshan, et al.
Veröffentlicht: (2026)
von: Duo, Jiangshan, et al.
Veröffentlicht: (2026)
D-PACE: Dynamic Position-Aware Cross-Entropy for Parallel Speculative Drafting
von: Wu, Tianyu, et al.
Veröffentlicht: (2026)
von: Wu, Tianyu, et al.
Veröffentlicht: (2026)
Rethinking Entropy Interventions in RLVR: An Entropy Change Perspective
von: Hao, Zhezheng, et al.
Veröffentlicht: (2025)
von: Hao, Zhezheng, et al.
Veröffentlicht: (2025)
Beyond Uniform Credit Assignment: Selective Eligibility Traces for RLVR
von: Mou, Chaoli, et al.
Veröffentlicht: (2026)
von: Mou, Chaoli, et al.
Veröffentlicht: (2026)
The Debate on RLVR Reasoning Capability Boundary: Shrinkage, Expansion, or Both? A Two-Stage Dynamic View
von: Yao, Xinhao, et al.
Veröffentlicht: (2025)
von: Yao, Xinhao, et al.
Veröffentlicht: (2025)
A Study on the Calibration of In-context Learning
von: Zhang, Hanlin, et al.
Veröffentlicht: (2023)
von: Zhang, Hanlin, et al.
Veröffentlicht: (2023)
Mitigating Distribution Sharpening in Math RLVR via Distribution-Aligned Hint Synthesis and Backward Hint Annealing
von: Xie, Pei-Xi, et al.
Veröffentlicht: (2026)
von: Xie, Pei-Xi, et al.
Veröffentlicht: (2026)
VL Norm: Rethink Loss Aggregation in RLVR
von: He, Zhiyuan, et al.
Veröffentlicht: (2025)
von: He, Zhiyuan, et al.
Veröffentlicht: (2025)
Spurious Rewards: Rethinking Training Signals in RLVR
von: Shao, Rulin, et al.
Veröffentlicht: (2025)
von: Shao, Rulin, et al.
Veröffentlicht: (2025)
A State-of-the-Art SQL Reasoning Model using RLVR
von: Ali, Alnur, et al.
Veröffentlicht: (2025)
von: Ali, Alnur, et al.
Veröffentlicht: (2025)
Manipulating Large Language Models to Increase Product Visibility
von: Kumar, Aounon, et al.
Veröffentlicht: (2024)
von: Kumar, Aounon, et al.
Veröffentlicht: (2024)
Part II: ROLL Flash -- Accelerating RLVR and Agentic Training with Asynchrony
von: Lu, Han, et al.
Veröffentlicht: (2025)
von: Lu, Han, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Measuring the Faithfulness of Thinking Drafts in Large Reasoning Models
von: Xiong, Zidi, et al.
Veröffentlicht: (2025) -
Learning Recourse Costs from Pairwise Feature Comparisons
von: Rawal, Kaivalya, et al.
Veröffentlicht: (2024) -
In-Context Unlearning: Language Models as Few Shot Unlearners
von: Pawelczyk, Martin, et al.
Veröffentlicht: (2023) -
Towards Unified Attribution in Explainable AI, Data-Centric AI, and Mechanistic Interpretability
von: Zhang, Shichang, et al.
Veröffentlicht: (2025) -
Who Gets Credit or Blame? Attributing Accountability in Modern AI Systems
von: Zhang, Shichang, et al.
Veröffentlicht: (2025)