CurveRL: Principled Distribution-Aware Context Reweighting for LLM Reasoning
Fuente:
arXiv
Salvato in:
| Autori principali: | Sun, Ke, Zhao, Yizhou, Xin, Jiayi, Long, Qi, Su, Weijie |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Mitigating Think-Answer Mismatch in LLM Reasoning Through Noise-Aware Advantage Reweighting
di: Shen, Si, et al.
Pubblicazione: (2025)
di: Shen, Si, et al.
Pubblicazione: (2025)
Robust Spectral Watermark for Synthetic Tabular Data
di: Zhao, Yizhou, et al.
Pubblicazione: (2025)
di: Zhao, Yizhou, et al.
Pubblicazione: (2025)
UCS: Estimating Unseen Coverage for Improved In-Context Learning
di: Xin, Jiayi, et al.
Pubblicazione: (2026)
di: Xin, Jiayi, et al.
Pubblicazione: (2026)
Evaluating the Unseen Capabilities: How Many Theorems Do LLMs Know?
di: Li, Xiang, et al.
Pubblicazione: (2025)
di: Li, Xiang, et al.
Pubblicazione: (2025)
Topology-Aware Dynamic Reweighting for Distribution Shifts on Graph
di: Zheng, Weihuang, et al.
Pubblicazione: (2024)
di: Zheng, Weihuang, et al.
Pubblicazione: (2024)
The Impact of Language Mixing on Bilingual LLM Reasoning
di: Li, Yihao, et al.
Pubblicazione: (2025)
di: Li, Yihao, et al.
Pubblicazione: (2025)
A Principled Path to Fitted Distributional Evaluation
di: Hong, Sungee, et al.
Pubblicazione: (2025)
di: Hong, Sungee, et al.
Pubblicazione: (2025)
The CAP Principle for LLM Serving: A Survey of Long-Context Large Language Model Serving
di: Zeng, Pai, et al.
Pubblicazione: (2024)
di: Zeng, Pai, et al.
Pubblicazione: (2024)
FlowRL: Matching Reward Distributions for LLM Reasoning
di: Zhu, Xuekai, et al.
Pubblicazione: (2025)
di: Zhu, Xuekai, et al.
Pubblicazione: (2025)
Bridging the Gap: Rademacher Complexity in Robust and Standard Generalization
di: Xiao, Jiancong, et al.
Pubblicazione: (2024)
di: Xiao, Jiancong, et al.
Pubblicazione: (2024)
DAJ: Data-Reweighted LLM Judge for Test-Time Scaling in Code Generation
di: Qin, Peijia, et al.
Pubblicazione: (2026)
di: Qin, Peijia, et al.
Pubblicazione: (2026)
Don't Waste Mistakes: Leveraging Negative RL-Groups via Confidence Reweighting
di: Feng, Yunzhen, et al.
Pubblicazione: (2025)
di: Feng, Yunzhen, et al.
Pubblicazione: (2025)
On Context-Content Uncertainty Principle
di: Li, Xin
Pubblicazione: (2025)
di: Li, Xin
Pubblicazione: (2025)
PickLLM: Context-Aware RL-Assisted Large Language Model Routing
di: Sikeridis, Dimitrios, et al.
Pubblicazione: (2024)
di: Sikeridis, Dimitrios, et al.
Pubblicazione: (2024)
RaaS: Reasoning-Aware Attention Sparsity for Efficient LLM Reasoning
di: Hu, Junhao, et al.
Pubblicazione: (2025)
di: Hu, Junhao, et al.
Pubblicazione: (2025)
Online LLM watermark detection via e-processes
di: Su, Weijie, et al.
Pubblicazione: (2026)
di: Su, Weijie, et al.
Pubblicazione: (2026)
Constrained Reweighting of Distributions: an Optimal Transport Approach
di: Chakraborty, Abhisek, et al.
Pubblicazione: (2023)
di: Chakraborty, Abhisek, et al.
Pubblicazione: (2023)
Architecture-Aware Learning Curve Extrapolation via Graph Ordinary Differential Equation
di: Ding, Yanna, et al.
Pubblicazione: (2024)
di: Ding, Yanna, et al.
Pubblicazione: (2024)
DreamPRM: Domain-Reweighted Process Reward Model for Multimodal Reasoning
di: Cao, Qi, et al.
Pubblicazione: (2025)
di: Cao, Qi, et al.
Pubblicazione: (2025)
Enhancing RL Safety with Counterfactual LLM Reasoning
di: Gross, Dennis, et al.
Pubblicazione: (2024)
di: Gross, Dennis, et al.
Pubblicazione: (2024)
Accelerating RL for LLM Reasoning with Optimal Advantage Regression
di: Brantley, Kianté, et al.
Pubblicazione: (2025)
di: Brantley, Kianté, et al.
Pubblicazione: (2025)
Improving RL Exploration for LLM Reasoning through Retrospective Replay
di: Dou, Shihan, et al.
Pubblicazione: (2025)
di: Dou, Shihan, et al.
Pubblicazione: (2025)
Token-Efficient RL for LLM Reasoning
di: Lee, Alan, et al.
Pubblicazione: (2025)
di: Lee, Alan, et al.
Pubblicazione: (2025)
Minimax Estimation for Personalized Federated Learning: An Alternative between FedAvg and Local Training?
di: Chen, Shuxiao, et al.
Pubblicazione: (2021)
di: Chen, Shuxiao, et al.
Pubblicazione: (2021)
Fundamental Limits of Game-Theoretic LLM Alignment: Smith Consistency and Preference Matching
di: Shi, Zhekun, et al.
Pubblicazione: (2025)
di: Shi, Zhekun, et al.
Pubblicazione: (2025)
Addressing Performance Saturation for LLM RL via Precise Entropy Curve Control
di: Li, Bolian, et al.
Pubblicazione: (2026)
di: Li, Bolian, et al.
Pubblicazione: (2026)
PolarGrad: A Class of Matrix-Gradient Optimizers from a Unifying Preconditioning Perspective
di: Lau, Tim Tsz-Kit, et al.
Pubblicazione: (2025)
di: Lau, Tim Tsz-Kit, et al.
Pubblicazione: (2025)
An Empirical Study on the Effectiveness of Incorporating Offline RL As Online RL Subroutines
di: Su, Jianhai, et al.
Pubblicazione: (2025)
di: Su, Jianhai, et al.
Pubblicazione: (2025)
Exploring Criteria of Loss Reweighting to Enhance LLM Unlearning
di: Yang, Puning, et al.
Pubblicazione: (2025)
di: Yang, Puning, et al.
Pubblicazione: (2025)
On the Necessity of Output Distribution Reweighting for Effective Class Unlearning
di: Ebrahimpour-Boroojeny, Ali, et al.
Pubblicazione: (2025)
di: Ebrahimpour-Boroojeny, Ali, et al.
Pubblicazione: (2025)
DISA: Offline Importance Sampling for Distribution-Matching LLM-RL
di: Wang, Shaobo, et al.
Pubblicazione: (2026)
di: Wang, Shaobo, et al.
Pubblicazione: (2026)
Learning to Extract Context for Context-Aware LLM Inference
di: Kim, Minseon, et al.
Pubblicazione: (2025)
di: Kim, Minseon, et al.
Pubblicazione: (2025)
Out-of-Distribution Adaptation in Offline RL: Counterfactual Reasoning via Causal Normalizing Flows
di: Cho, Minjae, et al.
Pubblicazione: (2024)
di: Cho, Minjae, et al.
Pubblicazione: (2024)
On the Limits of Sparse Autoencoders: A Theoretical Framework and Reweighted Remedy
di: Cui, Jingyi, et al.
Pubblicazione: (2025)
di: Cui, Jingyi, et al.
Pubblicazione: (2025)
Beat the long tail: Distribution-Aware Speculative Decoding for RL Training
di: Shao, Zelei, et al.
Pubblicazione: (2025)
di: Shao, Zelei, et al.
Pubblicazione: (2025)
Contextual Bandits for Unbounded Context Distributions
di: Zhao, Puning, et al.
Pubblicazione: (2024)
di: Zhao, Puning, et al.
Pubblicazione: (2024)
Beyond High-Entropy Exploration: Correctness-Aware Low-Entropy Segment-Based Advantage Shaping for Reasoning LLMs
di: Chen, Xinzhu, et al.
Pubblicazione: (2025)
di: Chen, Xinzhu, et al.
Pubblicazione: (2025)
Restoring Calibration for Aligned Large Language Models: A Calibration-Aware Fine-Tuning Approach
di: Xiao, Jiancong, et al.
Pubblicazione: (2025)
di: Xiao, Jiancong, et al.
Pubblicazione: (2025)
Part I: Tricks or Traps? A Deep Dive into RL for LLM Reasoning
di: Liu, Zihe, et al.
Pubblicazione: (2025)
di: Liu, Zihe, et al.
Pubblicazione: (2025)
DFedReweighting: A Unified Framework for Objective-Oriented Reweighting in Decentralized Federated Learning
di: Zhang, Kaichuang, et al.
Pubblicazione: (2025)
di: Zhang, Kaichuang, et al.
Pubblicazione: (2025)
Documenti analoghi
-
Mitigating Think-Answer Mismatch in LLM Reasoning Through Noise-Aware Advantage Reweighting
di: Shen, Si, et al.
Pubblicazione: (2025) -
Robust Spectral Watermark for Synthetic Tabular Data
di: Zhao, Yizhou, et al.
Pubblicazione: (2025) -
UCS: Estimating Unseen Coverage for Improved In-Context Learning
di: Xin, Jiayi, et al.
Pubblicazione: (2026) -
Evaluating the Unseen Capabilities: How Many Theorems Do LLMs Know?
di: Li, Xiang, et al.
Pubblicazione: (2025) -
Topology-Aware Dynamic Reweighting for Distribution Shifts on Graph
di: Zheng, Weihuang, et al.
Pubblicazione: (2024)