The Path Not Taken: RLVR Provably Learns Off the Principals
Fuente:
arXiv
Saved in:
| Main Authors: | Zhu, Hanqing, Zhang, Zhenyu, Huang, Hanxian, Su, DiJia, Liu, Zechun, Zhao, Jiawei, Fedorov, Igor, Pirsiavash, Hamed, Sha, Zhizhou, Lee, Jinwon, Pan, David Z., Wang, Zhangyang, Tian, Yuandong, Tai, Kai Sheng |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
R-Sparse: Rank-Aware Activation Sparsity for Efficient LLM Inference
by: Zhang, Zhenyu, et al.
Published: (2025)
by: Zhang, Zhenyu, et al.
Published: (2025)
GaLore 2: Large-Scale LLM Pre-Training by Gradient Low-Rank Projection
by: Su, DiJia, et al.
Published: (2025)
by: Su, DiJia, et al.
Published: (2025)
gen2seg: Generative Models Enable Generalizable Instance Segmentation
by: Khangaonkar, Om, et al.
Published: (2025)
by: Khangaonkar, Om, et al.
Published: (2025)
Dualformer: Controllable Fast and Slow Thinking by Learning with Randomized Reasoning Traces
by: Su, DiJia, et al.
Published: (2024)
by: Su, DiJia, et al.
Published: (2024)
GaLore: Memory-Efficient LLM Training by Gradient Low-Rank Projection
by: Zhao, Jiawei, et al.
Published: (2024)
by: Zhao, Jiawei, et al.
Published: (2024)
SimA: Simple Softmax-free Attention for Vision Transformers
by: Koohpayegani, Soroush Abbasi, et al.
Published: (2022)
by: Koohpayegani, Soroush Abbasi, et al.
Published: (2022)
Param$Δ$ for Direct Weight Mixing: Post-Train Large Language Model at Zero Cost
by: Cao, Sheng, et al.
Published: (2025)
by: Cao, Sheng, et al.
Published: (2025)
APOLLO: SGD-like Memory, AdamW-level Performance
by: Zhu, Hanqing, et al.
Published: (2024)
by: Zhu, Hanqing, et al.
Published: (2024)
Token Assorted: Mixing Latent and Text Tokens for Improved Language Model Reasoning
by: Su, DiJia, et al.
Published: (2025)
by: Su, DiJia, et al.
Published: (2025)
Why Data Anonymization Has Not Taken Off
by: Schneider, Matthew J., et al.
Published: (2025)
by: Schneider, Matthew J., et al.
Published: (2025)
You Only Use Reactive Attention Slice For Long Context Retrieval
by: Soh, Yun Joon, et al.
Published: (2024)
by: Soh, Yun Joon, et al.
Published: (2024)
Provable Scaling Laws of Feature Emergence from Learning Dynamics of Grokking
by: Tian, Yuandong
Published: (2025)
by: Tian, Yuandong
Published: (2025)
Q-GaLore: Quantized GaLore with INT4 Projection and Layer-Adaptive Low-Rank Gradients
by: Zhang, Zhenyu, et al.
Published: (2024)
by: Zhang, Zhenyu, et al.
Published: (2024)
SpinQuant: LLM quantization with learned rotations
by: Liu, Zechun, et al.
Published: (2024)
by: Liu, Zechun, et al.
Published: (2024)
Multimodal Language Models Cannot Spot Spatial Inconsistencies
by: Khangaonkar, Om, et al.
Published: (2026)
by: Khangaonkar, Om, et al.
Published: (2026)
Paths Not Taken: A Secure Computing Tutorial
by: Boebert, William Earl
Published: (2025)
by: Boebert, William Earl
Published: (2025)
On Computational Limits and Provably Efficient Criteria of Visual Autoregressive Models: A Fine-Grained Complexity Analysis
by: Ke, Yekun, et al.
Published: (2025)
by: Ke, Yekun, et al.
Published: (2025)
Training Large Language Models to Reason in a Continuous Latent Space
by: Hao, Shibo, et al.
Published: (2024)
by: Hao, Shibo, et al.
Published: (2024)
Beyond A*: Better Planning with Transformers via Search Dynamics Bootstrapping
by: Lehnert, Lucas, et al.
Published: (2024)
by: Lehnert, Lucas, et al.
Published: (2024)
One Category One Prompt: Dataset Distillation using Diffusion Models
by: Abbasi, Ali, et al.
Published: (2024)
by: Abbasi, Ali, et al.
Published: (2024)
Vector-Quantized Soft Label Compression for Dataset Distillation
by: Abbasi, Ali, et al.
Published: (2026)
by: Abbasi, Ali, et al.
Published: (2026)
The Path Not Taken: Duality in Reasoning about Program Execution
by: Hasanov, Eshgin, et al.
Published: (2026)
by: Hasanov, Eshgin, et al.
Published: (2026)
Paths Not Taken: Understanding and Mending the Multilingual Factual Recall Pipeline
by: Lu, Meng, et al.
Published: (2025)
by: Lu, Meng, et al.
Published: (2025)
Australia and the Path Not Taken: The Declining Independence and Influence of Middle Powers
by: Mark Beeson
Published: (2025)
by: Mark Beeson
Published: (2025)
Provable Model-Parallel Distributed Principal Component Analysis with Parallel Deflation
by: Liao, Fangshuo, et al.
Published: (2025)
by: Liao, Fangshuo, et al.
Published: (2025)
LoCoCo: Dropping In Convolutions for Long Context Compression
by: Cai, Ruisi, et al.
Published: (2024)
by: Cai, Ruisi, et al.
Published: (2024)
SHARP: Accelerating Language Model Inference by SHaring Adjacent layers with Recovery Parameters
by: Wang, Yiping, et al.
Published: (2025)
by: Wang, Yiping, et al.
Published: (2025)
Data-Efficient RLVR via Off-Policy Influence Guidance
by: Zhu, Erle, et al.
Published: (2025)
by: Zhu, Erle, et al.
Published: (2025)
Conformal Selective Acting: Anytime-Valid Risk Control for RLVR-Trained LLMs
by: Khosravi, Hamed, et al.
Published: (2026)
by: Khosravi, Hamed, et al.
Published: (2026)
Sinkhorn-Drifting Generative Models
by: He, Ping, et al.
Published: (2026)
by: He, Ping, et al.
Published: (2026)
MoIN: Mixture of Introvert Experts to Upcycle an LLM
by: Tejankar, Ajinkya, et al.
Published: (2024)
by: Tejankar, Ajinkya, et al.
Published: (2024)
CompGS: Smaller and Faster Gaussian Splatting with Vector Quantization
by: Navaneet, KL, et al.
Published: (2023)
by: Navaneet, KL, et al.
Published: (2023)
IO-SVD: Input-Output Whitened SVD for Adaptive-Rank LLM Compression
by: Abbasi, Ali, et al.
Published: (2026)
by: Abbasi, Ali, et al.
Published: (2026)
MobileLLM: Optimizing Sub-billion Parameter Language Models for On-Device Use Cases
by: Liu, Zechun, et al.
Published: (2024)
by: Liu, Zechun, et al.
Published: (2024)
MobileMoE: Scaling On-Device Mixture of Experts
by: Chen, Yanbei, et al.
Published: (2026)
by: Chen, Yanbei, et al.
Published: (2026)
PubSwap: Public-Data Off-Policy Coordination for Federated RLVR
by: Nayak, Anupam, et al.
Published: (2026)
by: Nayak, Anupam, et al.
Published: (2026)
Multi-modal Learning for WebAssembly Reverse Engineering
by: Huang, Hanxian, et al.
Published: (2024)
by: Huang, Hanxian, et al.
Published: (2024)
Intellectual Freedom Stands of American Bible College Libraries: Taken or Not Taken.
by: Dahl, Katherine
Published: (1988)
by: Dahl, Katherine
Published: (1988)
ParetoQ: Improving Scaling Laws in Extremely Low-bit LLM Quantization
by: Liu, Zechun, et al.
Published: (2025)
by: Liu, Zechun, et al.
Published: (2025)
NOLA: Compressing LoRA using Linear Combination of Random Basis
by: Koohpayegani, Soroush Abbasi, et al.
Published: (2023)
by: Koohpayegani, Soroush Abbasi, et al.
Published: (2023)
Similar Items
-
R-Sparse: Rank-Aware Activation Sparsity for Efficient LLM Inference
by: Zhang, Zhenyu, et al.
Published: (2025) -
GaLore 2: Large-Scale LLM Pre-Training by Gradient Low-Rank Projection
by: Su, DiJia, et al.
Published: (2025) -
gen2seg: Generative Models Enable Generalizable Instance Segmentation
by: Khangaonkar, Om, et al.
Published: (2025) -
Dualformer: Controllable Fast and Slow Thinking by Learning with Randomized Reasoning Traces
by: Su, DiJia, et al.
Published: (2024) -
GaLore: Memory-Efficient LLM Training by Gradient Low-Rank Projection
by: Zhao, Jiawei, et al.
Published: (2024)