Monitorability as a Free Gift: How RLVR Spontaneously Aligns Reasoning
Fuente:
arXiv
Guardado en:
| Autores principales: | Xiong, Zidi, Chen, Shan, Lakkaraju, Himabindu |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Measuring the Faithfulness of Thinking Drafts in Large Reasoning Models
por: Xiong, Zidi, et al.
Publicado: (2025)
por: Xiong, Zidi, et al.
Publicado: (2025)
Learning Recourse Costs from Pairwise Feature Comparisons
por: Rawal, Kaivalya, et al.
Publicado: (2024)
por: Rawal, Kaivalya, et al.
Publicado: (2024)
In-Context Unlearning: Language Models as Few Shot Unlearners
por: Pawelczyk, Martin, et al.
Publicado: (2023)
por: Pawelczyk, Martin, et al.
Publicado: (2023)
Towards Unified Attribution in Explainable AI, Data-Centric AI, and Mechanistic Interpretability
por: Zhang, Shichang, et al.
Publicado: (2025)
por: Zhang, Shichang, et al.
Publicado: (2025)
Who Gets Credit or Blame? Attributing Accountability in Modern AI Systems
por: Zhang, Shichang, et al.
Publicado: (2025)
por: Zhang, Shichang, et al.
Publicado: (2025)
Operationalizing the Blueprint for an AI Bill of Rights: Recommendations for Practitioners, Researchers, and Policy Makers
por: Oesterling, Alex, et al.
Publicado: (2024)
por: Oesterling, Alex, et al.
Publicado: (2024)
Generalized Group Data Attribution
por: Ley, Dan, et al.
Publicado: (2024)
por: Ley, Dan, et al.
Publicado: (2024)
Generalizing Trust: Weak-to-Strong Trustworthiness in Language Models
por: Pawelczyk, Martin, et al.
Publicado: (2024)
por: Pawelczyk, Martin, et al.
Publicado: (2024)
Evaluating Adversarial Robustness of Concept Representations in Sparse Autoencoders
por: Li, Aaron J., et al.
Publicado: (2025)
por: Li, Aaron J., et al.
Publicado: (2025)
The Disagreement Problem in Explainable Machine Learning: A Practitioner's Perspective
por: Krishna, Satyapriya, et al.
Publicado: (2022)
por: Krishna, Satyapriya, et al.
Publicado: (2022)
Data Poisoning Attacks on Off-Policy Policy Evaluation Methods
por: Lobo, Elita, et al.
Publicado: (2024)
por: Lobo, Elita, et al.
Publicado: (2024)
In-Context Explainers: Harnessing LLMs for Explaining Black Box Models
por: Kroeger, Nicholas, et al.
Publicado: (2023)
por: Kroeger, Nicholas, et al.
Publicado: (2023)
On the Direction of RLVR Updates for LLM Reasoning: Identification and Exploitation
por: Huang, Kexin, et al.
Publicado: (2026)
por: Huang, Kexin, et al.
Publicado: (2026)
How Post-Training Reshapes LLMs: A Mechanistic View on Knowledge, Truthfulness, Refusal, and Confidence
por: Du, Hongzhe, et al.
Publicado: (2025)
por: Du, Hongzhe, et al.
Publicado: (2025)
How Memory Management Impacts LLM Agents: An Empirical Study of Experience-Following Behavior
por: Xiong, Zidi, et al.
Publicado: (2025)
por: Xiong, Zidi, et al.
Publicado: (2025)
Quantile Advantage Estimation: Stabilizing RLVR for LLM Reasoning
por: Wu, Junkang, et al.
Publicado: (2025)
por: Wu, Junkang, et al.
Publicado: (2025)
Temporal Sparse Autoencoders: Leveraging the Sequential Nature of Language for Interpretability
por: Bhalla, Usha, et al.
Publicado: (2025)
por: Bhalla, Usha, et al.
Publicado: (2025)
Follow My Instruction and Spill the Beans: Scalable Data Extraction from Retrieval-Augmented Generation Systems
por: Qi, Zhenting, et al.
Publicado: (2024)
por: Qi, Zhenting, et al.
Publicado: (2024)
Controllable Exploration in Hybrid-Policy RLVR for Multi-Modal Reasoning
por: Huang, Zhuoxu, et al.
Publicado: (2026)
por: Huang, Zhuoxu, et al.
Publicado: (2026)
OpenXAI: Towards a Transparent Evaluation of Model Explanations
por: Agarwal, Chirag, et al.
Publicado: (2022)
por: Agarwal, Chirag, et al.
Publicado: (2022)
Generalization of RLVR Using Causal Reasoning as a Testbed
por: Lu, Brian, et al.
Publicado: (2025)
por: Lu, Brian, et al.
Publicado: (2025)
Adaptive Negative Reinforcement for LLM Reasoning:Dynamically Balancing Correction and Diversity in RLVR
por: Ingle, Yash, et al.
Publicado: (2026)
por: Ingle, Yash, et al.
Publicado: (2026)
Depth-Breadth Synergy in RLVR: Unlocking LLM Reasoning Gains with Adaptive Exploration
por: Yang, Zhicheng, et al.
Publicado: (2025)
por: Yang, Zhicheng, et al.
Publicado: (2025)
RLVR without Ineffective Samples: Group Prioritized Off-Policy Optimization for LLM Reasoning
por: Mao, Yixiu, et al.
Publicado: (2026)
por: Mao, Yixiu, et al.
Publicado: (2026)
Shorter but not Worse: Frugal Reasoning via Easy Samples as Length Regularizers in Math RLVR
por: Bounhar, Abdelaziz, et al.
Publicado: (2025)
por: Bounhar, Abdelaziz, et al.
Publicado: (2025)
Quantifying Empirical Compute-Supervision Tradeoffs in RLVR
por: Mitsuhashi, Ryo, et al.
Publicado: (2026)
por: Mitsuhashi, Ryo, et al.
Publicado: (2026)
Open-Medical-R1: How to Choose Data for RLVR Training at Medicine Domain
por: Qiu, Zhongxi, et al.
Publicado: (2025)
por: Qiu, Zhongxi, et al.
Publicado: (2025)
Certifying LLM Safety against Adversarial Prompting
por: Kumar, Aounon, et al.
Publicado: (2023)
por: Kumar, Aounon, et al.
Publicado: (2023)
JudgeRLVR: Judge First, Generate Second for Efficient Reasoning
por: Duo, Jiangshan, et al.
Publicado: (2026)
por: Duo, Jiangshan, et al.
Publicado: (2026)
D-PACE: Dynamic Position-Aware Cross-Entropy for Parallel Speculative Drafting
por: Wu, Tianyu, et al.
Publicado: (2026)
por: Wu, Tianyu, et al.
Publicado: (2026)
Rethinking Entropy Interventions in RLVR: An Entropy Change Perspective
por: Hao, Zhezheng, et al.
Publicado: (2025)
por: Hao, Zhezheng, et al.
Publicado: (2025)
Beyond Uniform Credit Assignment: Selective Eligibility Traces for RLVR
por: Mou, Chaoli, et al.
Publicado: (2026)
por: Mou, Chaoli, et al.
Publicado: (2026)
The Debate on RLVR Reasoning Capability Boundary: Shrinkage, Expansion, or Both? A Two-Stage Dynamic View
por: Yao, Xinhao, et al.
Publicado: (2025)
por: Yao, Xinhao, et al.
Publicado: (2025)
A Study on the Calibration of In-context Learning
por: Zhang, Hanlin, et al.
Publicado: (2023)
por: Zhang, Hanlin, et al.
Publicado: (2023)
Mitigating Distribution Sharpening in Math RLVR via Distribution-Aligned Hint Synthesis and Backward Hint Annealing
por: Xie, Pei-Xi, et al.
Publicado: (2026)
por: Xie, Pei-Xi, et al.
Publicado: (2026)
VL Norm: Rethink Loss Aggregation in RLVR
por: He, Zhiyuan, et al.
Publicado: (2025)
por: He, Zhiyuan, et al.
Publicado: (2025)
Spurious Rewards: Rethinking Training Signals in RLVR
por: Shao, Rulin, et al.
Publicado: (2025)
por: Shao, Rulin, et al.
Publicado: (2025)
A State-of-the-Art SQL Reasoning Model using RLVR
por: Ali, Alnur, et al.
Publicado: (2025)
por: Ali, Alnur, et al.
Publicado: (2025)
Manipulating Large Language Models to Increase Product Visibility
por: Kumar, Aounon, et al.
Publicado: (2024)
por: Kumar, Aounon, et al.
Publicado: (2024)
Part II: ROLL Flash -- Accelerating RLVR and Agentic Training with Asynchrony
por: Lu, Han, et al.
Publicado: (2025)
por: Lu, Han, et al.
Publicado: (2025)
Ejemplares similares
-
Measuring the Faithfulness of Thinking Drafts in Large Reasoning Models
por: Xiong, Zidi, et al.
Publicado: (2025) -
Learning Recourse Costs from Pairwise Feature Comparisons
por: Rawal, Kaivalya, et al.
Publicado: (2024) -
In-Context Unlearning: Language Models as Few Shot Unlearners
por: Pawelczyk, Martin, et al.
Publicado: (2023) -
Towards Unified Attribution in Explainable AI, Data-Centric AI, and Mechanistic Interpretability
por: Zhang, Shichang, et al.
Publicado: (2025) -
Who Gets Credit or Blame? Attributing Accountability in Modern AI Systems
por: Zhang, Shichang, et al.
Publicado: (2025)