Enregistré dans:
| Auteurs principaux: | Zhu, Youheng, Lu, Yiping |
|---|---|
| Format: | Preprint |
| Publié: |
2026
|
| Sujets: | |
| Accès en ligne: | https://arxiv.org/abs/2602.01381 |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
A Covering Framework for Offline POMDPs Learning using Belief Space Metric
par: Zhu, Youheng, et autres
Publié: (2026)
par: Zhu, Youheng, et autres
Publié: (2026)
Inference-Time Scaling for Generalist Reward Modeling
par: Liu, Zijun, et autres
Publié: (2025)
par: Liu, Zijun, et autres
Publié: (2025)
Entropy Centroids as Intrinsic Rewards for Test-Time Scaling
par: Zhao, Wenshuo, et autres
Publié: (2026)
par: Zhao, Wenshuo, et autres
Publié: (2026)
RFG: Test-Time Scaling for Diffusion Large Language Model Reasoning with Reward-Free Guidance
par: Chen, Tianlang, et autres
Publié: (2025)
par: Chen, Tianlang, et autres
Publié: (2025)
TDRM: Smooth Reward Models with Temporal Difference for LLM RL and Inference
par: Zhang, Dan, et autres
Publié: (2025)
par: Zhang, Dan, et autres
Publié: (2025)
Inference-Time Hyper-Scaling with KV Cache Compression
par: Łańcucki, Adrian, et autres
Publié: (2025)
par: Łańcucki, Adrian, et autres
Publié: (2025)
Scaling Inference-Time Search with Vision Value Model for Improved Visual Comprehension
par: Wang, Xiyao, et autres
Publié: (2024)
par: Wang, Xiyao, et autres
Publié: (2024)
What is a Sketch-and-Precondition Derivation for Low-Rank Approximation? Inverse Power Error or Inverse Power Estimation?
par: Xu, Ruihan, et autres
Publié: (2025)
par: Xu, Ruihan, et autres
Publié: (2025)
Simple Approximation and Derivative Free Inference-Time Scaling for Diffusion Models via Sequential Monte Carlo on Path Measures
par: Wang, Chenyang, et autres
Publié: (2026)
par: Wang, Chenyang, et autres
Publié: (2026)
Reward-Robust RLHF in LLMs
par: Yan, Yuzi, et autres
Publié: (2024)
par: Yan, Yuzi, et autres
Publié: (2024)
Temper and Tilt Lead to SLOP: Reward Hacking Mitigation with Inference-Time Alignment
par: Wang, Ye, et autres
Publié: (2026)
par: Wang, Ye, et autres
Publié: (2026)
GRAM-R$^2$: Self-Training Generative Foundation Reward Models for Reward Reasoning
par: Wang, Chenglong, et autres
Publié: (2025)
par: Wang, Chenglong, et autres
Publié: (2025)
Bayesian Preference Learning for Test-Time Steerable Reward Models
par: Hong, Jiwoo, et autres
Publié: (2026)
par: Hong, Jiwoo, et autres
Publié: (2026)
MarkovScale: Towards Optimal Sequential Scaling at Inference Time
par: Wang, Youkang, et autres
Publié: (2026)
par: Wang, Youkang, et autres
Publié: (2026)
T1: Advancing Language Model Reasoning through Reinforcement Learning and Inference Scaling
par: Hou, Zhenyu, et autres
Publié: (2025)
par: Hou, Zhenyu, et autres
Publié: (2025)
Noise Contrastive Alignment of Language Models with Explicit Rewards
par: Chen, Huayu, et autres
Publié: (2024)
par: Chen, Huayu, et autres
Publié: (2024)
GTS: Inference-Time Scaling of Latent Reasoning with a Learnable Gaussian Thought Sampler
par: Wang, Minghan, et autres
Publié: (2026)
par: Wang, Minghan, et autres
Publié: (2026)
Scaling Laws for Reward Model Overoptimization in Direct Alignment Algorithms
par: Rafailov, Rafael, et autres
Publié: (2024)
par: Rafailov, Rafael, et autres
Publié: (2024)
Are More Tokens Rational? Inference-Time Scaling in Language Models as Adaptive Resource Rationality
par: Hu, Zhimin, et autres
Publié: (2026)
par: Hu, Zhimin, et autres
Publié: (2026)
Rewarding Progress: Scaling Automated Process Verifiers for LLM Reasoning
par: Setlur, Amrith, et autres
Publié: (2024)
par: Setlur, Amrith, et autres
Publié: (2024)
Duel-Evolve: Reward-Free Test-Time Scaling via LLM Self-Preferences
par: Karlekar, Sweta, et autres
Publié: (2026)
par: Karlekar, Sweta, et autres
Publié: (2026)
Scaling over Scaling: Exploring Test-Time Scaling Plateau in Large Reasoning Models
par: Wang, Jian, et autres
Publié: (2025)
par: Wang, Jian, et autres
Publié: (2025)
Scaling Inference-Efficient Language Models
par: Bian, Song, et autres
Publié: (2025)
par: Bian, Song, et autres
Publié: (2025)
Cascade Reward Sampling for Efficient Decoding-Time Alignment
par: Li, Bolian, et autres
Publié: (2024)
par: Li, Bolian, et autres
Publié: (2024)
Process Reward Models That Think
par: Khalifa, Muhammad, et autres
Publié: (2025)
par: Khalifa, Muhammad, et autres
Publié: (2025)
Personalized RewardBench: Evaluating Reward Models with Human Aligned Personalization
par: Ma, Qiyao, et autres
Publié: (2026)
par: Ma, Qiyao, et autres
Publié: (2026)
Beyond Chinchilla-Optimal: Accounting for Inference in Language Model Scaling Laws
par: Sardana, Nikhil, et autres
Publié: (2023)
par: Sardana, Nikhil, et autres
Publié: (2023)
Fast-Slow Thinking RM: Efficient Integration of Scalar and Generative Reward Models
par: Wu, Jiayun, et autres
Publié: (2026)
par: Wu, Jiayun, et autres
Publié: (2026)
Latent Thought Models with Variational Bayes Inference-Time Computation
par: Kong, Deqian, et autres
Publié: (2025)
par: Kong, Deqian, et autres
Publié: (2025)
Test-Time Scaling with Reflective Generative Model
par: Wang, Zixiao, et autres
Publié: (2025)
par: Wang, Zixiao, et autres
Publié: (2025)
Entropy-Regularized Process Reward Model
par: Zhang, Hanning, et autres
Publié: (2024)
par: Zhang, Hanning, et autres
Publié: (2024)
On Almost Surely Safe Alignment of Large Language Models at Inference-Time
par: Ji, Xiaotong, et autres
Publié: (2025)
par: Ji, Xiaotong, et autres
Publié: (2025)
Auto MC-Reward: Automated Dense Reward Design with Large Language Models for Minecraft
par: Li, Hao, et autres
Publié: (2023)
par: Li, Hao, et autres
Publié: (2023)
Seed Diffusion: A Large-Scale Diffusion Language Model with High-Speed Inference
par: Song, Yuxuan, et autres
Publié: (2025)
par: Song, Yuxuan, et autres
Publié: (2025)
Mitigating Premature Exploitation in Particle-based Monte Carlo for Inference-Time Scaling
par: Giannone, Giorgio, et autres
Publié: (2025)
par: Giannone, Giorgio, et autres
Publié: (2025)
How to Evaluate Reward Models for RLHF
par: Frick, Evan, et autres
Publié: (2024)
par: Frick, Evan, et autres
Publié: (2024)
SDiaReward: Modeling and Benchmarking Spoken Dialogue Rewards with Modality and Colloquialness
par: Lu, Jingyu, et autres
Publié: (2026)
par: Lu, Jingyu, et autres
Publié: (2026)
Expected Reward Prediction, with Applications to Model Routing
par: Hasanaliyev, Kenan, et autres
Publié: (2026)
par: Hasanaliyev, Kenan, et autres
Publié: (2026)
Bootstrapping Language Models with DPO Implicit Rewards
par: Chen, Changyu, et autres
Publié: (2024)
par: Chen, Changyu, et autres
Publié: (2024)
RewardAnything: Generalizable Principle-Following Reward Models
par: Yu, Zhuohao, et autres
Publié: (2025)
par: Yu, Zhuohao, et autres
Publié: (2025)
Documents similaires
-
A Covering Framework for Offline POMDPs Learning using Belief Space Metric
par: Zhu, Youheng, et autres
Publié: (2026) -
Inference-Time Scaling for Generalist Reward Modeling
par: Liu, Zijun, et autres
Publié: (2025) -
Entropy Centroids as Intrinsic Rewards for Test-Time Scaling
par: Zhao, Wenshuo, et autres
Publié: (2026) -
RFG: Test-Time Scaling for Diffusion Large Language Model Reasoning with Reward-Free Guidance
par: Chen, Tianlang, et autres
Publié: (2025) -
TDRM: Smooth Reward Models with Temporal Difference for LLM RL and Inference
par: Zhang, Dan, et autres
Publié: (2025)