CurveRL: Principled Distribution-Aware Context Reweighting for LLM Reasoning
Fuente:
arXiv
Saved in:
| Main Authors: | Sun, Ke, Zhao, Yizhou, Xin, Jiayi, Long, Qi, Su, Weijie |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Mitigating Think-Answer Mismatch in LLM Reasoning Through Noise-Aware Advantage Reweighting
by: Shen, Si, et al.
Published: (2025)
by: Shen, Si, et al.
Published: (2025)
Robust Spectral Watermark for Synthetic Tabular Data
by: Zhao, Yizhou, et al.
Published: (2025)
by: Zhao, Yizhou, et al.
Published: (2025)
UCS: Estimating Unseen Coverage for Improved In-Context Learning
by: Xin, Jiayi, et al.
Published: (2026)
by: Xin, Jiayi, et al.
Published: (2026)
Evaluating the Unseen Capabilities: How Many Theorems Do LLMs Know?
by: Li, Xiang, et al.
Published: (2025)
by: Li, Xiang, et al.
Published: (2025)
Topology-Aware Dynamic Reweighting for Distribution Shifts on Graph
by: Zheng, Weihuang, et al.
Published: (2024)
by: Zheng, Weihuang, et al.
Published: (2024)
The Impact of Language Mixing on Bilingual LLM Reasoning
by: Li, Yihao, et al.
Published: (2025)
by: Li, Yihao, et al.
Published: (2025)
A Principled Path to Fitted Distributional Evaluation
by: Hong, Sungee, et al.
Published: (2025)
by: Hong, Sungee, et al.
Published: (2025)
The CAP Principle for LLM Serving: A Survey of Long-Context Large Language Model Serving
by: Zeng, Pai, et al.
Published: (2024)
by: Zeng, Pai, et al.
Published: (2024)
FlowRL: Matching Reward Distributions for LLM Reasoning
by: Zhu, Xuekai, et al.
Published: (2025)
by: Zhu, Xuekai, et al.
Published: (2025)
Bridging the Gap: Rademacher Complexity in Robust and Standard Generalization
by: Xiao, Jiancong, et al.
Published: (2024)
by: Xiao, Jiancong, et al.
Published: (2024)
DAJ: Data-Reweighted LLM Judge for Test-Time Scaling in Code Generation
by: Qin, Peijia, et al.
Published: (2026)
by: Qin, Peijia, et al.
Published: (2026)
Don't Waste Mistakes: Leveraging Negative RL-Groups via Confidence Reweighting
by: Feng, Yunzhen, et al.
Published: (2025)
by: Feng, Yunzhen, et al.
Published: (2025)
On Context-Content Uncertainty Principle
by: Li, Xin
Published: (2025)
by: Li, Xin
Published: (2025)
PickLLM: Context-Aware RL-Assisted Large Language Model Routing
by: Sikeridis, Dimitrios, et al.
Published: (2024)
by: Sikeridis, Dimitrios, et al.
Published: (2024)
RaaS: Reasoning-Aware Attention Sparsity for Efficient LLM Reasoning
by: Hu, Junhao, et al.
Published: (2025)
by: Hu, Junhao, et al.
Published: (2025)
Online LLM watermark detection via e-processes
by: Su, Weijie, et al.
Published: (2026)
by: Su, Weijie, et al.
Published: (2026)
Constrained Reweighting of Distributions: an Optimal Transport Approach
by: Chakraborty, Abhisek, et al.
Published: (2023)
by: Chakraborty, Abhisek, et al.
Published: (2023)
Architecture-Aware Learning Curve Extrapolation via Graph Ordinary Differential Equation
by: Ding, Yanna, et al.
Published: (2024)
by: Ding, Yanna, et al.
Published: (2024)
DreamPRM: Domain-Reweighted Process Reward Model for Multimodal Reasoning
by: Cao, Qi, et al.
Published: (2025)
by: Cao, Qi, et al.
Published: (2025)
Enhancing RL Safety with Counterfactual LLM Reasoning
by: Gross, Dennis, et al.
Published: (2024)
by: Gross, Dennis, et al.
Published: (2024)
Accelerating RL for LLM Reasoning with Optimal Advantage Regression
by: Brantley, Kianté, et al.
Published: (2025)
by: Brantley, Kianté, et al.
Published: (2025)
Improving RL Exploration for LLM Reasoning through Retrospective Replay
by: Dou, Shihan, et al.
Published: (2025)
by: Dou, Shihan, et al.
Published: (2025)
Token-Efficient RL for LLM Reasoning
by: Lee, Alan, et al.
Published: (2025)
by: Lee, Alan, et al.
Published: (2025)
Minimax Estimation for Personalized Federated Learning: An Alternative between FedAvg and Local Training?
by: Chen, Shuxiao, et al.
Published: (2021)
by: Chen, Shuxiao, et al.
Published: (2021)
Fundamental Limits of Game-Theoretic LLM Alignment: Smith Consistency and Preference Matching
by: Shi, Zhekun, et al.
Published: (2025)
by: Shi, Zhekun, et al.
Published: (2025)
Addressing Performance Saturation for LLM RL via Precise Entropy Curve Control
by: Li, Bolian, et al.
Published: (2026)
by: Li, Bolian, et al.
Published: (2026)
PolarGrad: A Class of Matrix-Gradient Optimizers from a Unifying Preconditioning Perspective
by: Lau, Tim Tsz-Kit, et al.
Published: (2025)
by: Lau, Tim Tsz-Kit, et al.
Published: (2025)
An Empirical Study on the Effectiveness of Incorporating Offline RL As Online RL Subroutines
by: Su, Jianhai, et al.
Published: (2025)
by: Su, Jianhai, et al.
Published: (2025)
Exploring Criteria of Loss Reweighting to Enhance LLM Unlearning
by: Yang, Puning, et al.
Published: (2025)
by: Yang, Puning, et al.
Published: (2025)
On the Necessity of Output Distribution Reweighting for Effective Class Unlearning
by: Ebrahimpour-Boroojeny, Ali, et al.
Published: (2025)
by: Ebrahimpour-Boroojeny, Ali, et al.
Published: (2025)
DISA: Offline Importance Sampling for Distribution-Matching LLM-RL
by: Wang, Shaobo, et al.
Published: (2026)
by: Wang, Shaobo, et al.
Published: (2026)
Learning to Extract Context for Context-Aware LLM Inference
by: Kim, Minseon, et al.
Published: (2025)
by: Kim, Minseon, et al.
Published: (2025)
Out-of-Distribution Adaptation in Offline RL: Counterfactual Reasoning via Causal Normalizing Flows
by: Cho, Minjae, et al.
Published: (2024)
by: Cho, Minjae, et al.
Published: (2024)
On the Limits of Sparse Autoencoders: A Theoretical Framework and Reweighted Remedy
by: Cui, Jingyi, et al.
Published: (2025)
by: Cui, Jingyi, et al.
Published: (2025)
Beat the long tail: Distribution-Aware Speculative Decoding for RL Training
by: Shao, Zelei, et al.
Published: (2025)
by: Shao, Zelei, et al.
Published: (2025)
Contextual Bandits for Unbounded Context Distributions
by: Zhao, Puning, et al.
Published: (2024)
by: Zhao, Puning, et al.
Published: (2024)
Beyond High-Entropy Exploration: Correctness-Aware Low-Entropy Segment-Based Advantage Shaping for Reasoning LLMs
by: Chen, Xinzhu, et al.
Published: (2025)
by: Chen, Xinzhu, et al.
Published: (2025)
Restoring Calibration for Aligned Large Language Models: A Calibration-Aware Fine-Tuning Approach
by: Xiao, Jiancong, et al.
Published: (2025)
by: Xiao, Jiancong, et al.
Published: (2025)
Part I: Tricks or Traps? A Deep Dive into RL for LLM Reasoning
by: Liu, Zihe, et al.
Published: (2025)
by: Liu, Zihe, et al.
Published: (2025)
DFedReweighting: A Unified Framework for Objective-Oriented Reweighting in Decentralized Federated Learning
by: Zhang, Kaichuang, et al.
Published: (2025)
by: Zhang, Kaichuang, et al.
Published: (2025)
Similar Items
-
Mitigating Think-Answer Mismatch in LLM Reasoning Through Noise-Aware Advantage Reweighting
by: Shen, Si, et al.
Published: (2025) -
Robust Spectral Watermark for Synthetic Tabular Data
by: Zhao, Yizhou, et al.
Published: (2025) -
UCS: Estimating Unseen Coverage for Improved In-Context Learning
by: Xin, Jiayi, et al.
Published: (2026) -
Evaluating the Unseen Capabilities: How Many Theorems Do LLMs Know?
by: Li, Xiang, et al.
Published: (2025) -
Topology-Aware Dynamic Reweighting for Distribution Shifts on Graph
by: Zheng, Weihuang, et al.
Published: (2024)