Reward-Shifted Speculative Sampling Is An Efficient Test-Time Weak-to-Strong Aligner
Fuente:
arXiv
Saved in:
| Main Authors: | Li, Bolian, Wu, Yanran, Luo, Xinyu, Zhang, Ruqi |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Cascade Reward Sampling for Efficient Decoding-Time Alignment
by: Li, Bolian, et al.
Published: (2024)
by: Li, Bolian, et al.
Published: (2024)
ETA: Evaluating Then Aligning Safety of Vision Language Models at Inference Time
by: Ding, Yi, et al.
Published: (2024)
by: Ding, Yi, et al.
Published: (2024)
Uniform-Correct Policy Optimization: Breaking RLVR's Indifference to Diversity
by: Lochab, Anamika, et al.
Published: (2026)
by: Lochab, Anamika, et al.
Published: (2026)
Learning Self-Correction in Vision-Language Models via Rollout Augmentation
by: Ding, Yi, et al.
Published: (2026)
by: Ding, Yi, et al.
Published: (2026)
Entropy-MCMC: Sampling from Flat Basins with Ease
by: Li, Bolian, et al.
Published: (2023)
by: Li, Bolian, et al.
Published: (2023)
Addressing Performance Saturation for LLM RL via Precise Entropy Curve Control
by: Li, Bolian, et al.
Published: (2026)
by: Li, Bolian, et al.
Published: (2026)
Accelerated Test-Time Scaling with Model-Free Speculative Sampling
by: Song, Woomin, et al.
Published: (2025)
by: Song, Woomin, et al.
Published: (2025)
Energy-Based Reward Models for Robust Language Model Alignment
by: Lochab, Anamika, et al.
Published: (2025)
by: Lochab, Anamika, et al.
Published: (2025)
When In-Distribution Gains Fail: Evaluating Weak-to-Strong Reward Models under Preference Shift
by: Le, Khoi, et al.
Published: (2026)
by: Le, Khoi, et al.
Published: (2026)
Calibrate, Don't Curate: Label-Efficient Estimation from Noisy LLM Judges
by: Li, Yanran
Published: (2026)
by: Li, Yanran
Published: (2026)
DRIFT: Learning from Abundant User Dissatisfaction in Real-World Preference Learning
by: Wang, Yifan, et al.
Published: (2025)
by: Wang, Yifan, et al.
Published: (2025)
Test-Time Speculation
by: Kumar, Avinash, et al.
Published: (2026)
by: Kumar, Avinash, et al.
Published: (2026)
Weak-to-Strong Preference Optimization: Stealing Reward from Weak Aligned Model
by: Zhu, Wenhong, et al.
Published: (2024)
by: Zhu, Wenhong, et al.
Published: (2024)
Reward-Guided Speculative Decoding for Efficient LLM Reasoning
by: Liao, Baohao, et al.
Published: (2025)
by: Liao, Baohao, et al.
Published: (2025)
Alignment-Augmented Speculative Decoding with Alignment Sampling and Conditional Verification
by: Wang, Jikai, et al.
Published: (2025)
by: Wang, Jikai, et al.
Published: (2025)
VCap: Hypergeometric Rewards for Weak-to-Strong Visual Captioning
by: Lu, Xingyu, et al.
Published: (2026)
by: Lu, Xingyu, et al.
Published: (2026)
TransAlign: Machine Translation Encoders are Strong Word Aligners, Too
by: Ebing, Benedikt, et al.
Published: (2025)
by: Ebing, Benedikt, et al.
Published: (2025)
Making Reliable and Flexible Decisions in Long-tailed Classification
by: Li, Bolian, et al.
Published: (2025)
by: Li, Bolian, et al.
Published: (2025)
Aligner: Efficient Alignment by Learning to Correct
by: Ji, Jiaming, et al.
Published: (2024)
by: Ji, Jiaming, et al.
Published: (2024)
Rethinking Sample Polarity in Reinforcement Learning with Verifiable Rewards
by: Tang, Xinyu, et al.
Published: (2025)
by: Tang, Xinyu, et al.
Published: (2025)
InferAligner: Inference-Time Alignment for Harmlessness through Cross-Model Guidance
by: Wang, Pengyu, et al.
Published: (2024)
by: Wang, Pengyu, et al.
Published: (2024)
ArcAligner: Adaptive Recursive Aligner for Compressed Context Embeddings in RAG
by: Li, Jianbo, et al.
Published: (2026)
by: Li, Jianbo, et al.
Published: (2026)
Efficient Reasoning for LLMs through Speculative Chain-of-Thought
by: Wang, Jikai, et al.
Published: (2025)
by: Wang, Jikai, et al.
Published: (2025)
Speculative Decoding for Multi-Sample Inference
by: Li, Yiwei, et al.
Published: (2025)
by: Li, Yiwei, et al.
Published: (2025)
Scaling Up, Speeding Up: A Benchmark of Speculative Decoding for Efficient LLM Test-Time Scaling
by: Sun, Shengyin, et al.
Published: (2025)
by: Sun, Shengyin, et al.
Published: (2025)
EAGLE: Speculative Sampling Requires Rethinking Feature Uncertainty
by: Li, Yuhui, et al.
Published: (2024)
by: Li, Yuhui, et al.
Published: (2024)
Multi-Candidate Speculative Decoding
by: Yang, Sen, et al.
Published: (2024)
by: Yang, Sen, et al.
Published: (2024)
Distribution-Aware Reward Estimation for Test-Time Reinforcement Learning
by: Du, Bodong, et al.
Published: (2026)
by: Du, Bodong, et al.
Published: (2026)
Learning Harmonized Representations for Speculative Sampling
by: Zhang, Lefan, et al.
Published: (2024)
by: Zhang, Lefan, et al.
Published: (2024)
CoT-UQ: Improving Response-wise Uncertainty Quantification in LLMs with Chain-of-Thought
by: Zhang, Boxuan, et al.
Published: (2025)
by: Zhang, Boxuan, et al.
Published: (2025)
Sampling-Efficient Test-Time Scaling: Self-Estimating the Best-of-N Sampling in Early Decoding
by: Wang, Yiming, et al.
Published: (2025)
by: Wang, Yiming, et al.
Published: (2025)
AdaSPEC: Selective Knowledge Distillation for Efficient Speculative Decoders
by: Hu, Yuezhou, et al.
Published: (2025)
by: Hu, Yuezhou, et al.
Published: (2025)
NeMo-Aligner: Scalable Toolkit for Efficient Model Alignment
by: Shen, Gerald, et al.
Published: (2024)
by: Shen, Gerald, et al.
Published: (2024)
Selective Weak-to-Strong Generalization
by: Lang, Hao, et al.
Published: (2025)
by: Lang, Hao, et al.
Published: (2025)
Out-of-Vocabulary Sampling Boosts Speculative Decoding
by: Timor, Nadav, et al.
Published: (2025)
by: Timor, Nadav, et al.
Published: (2025)
Controlled LLM Decoding via Discrete Auto-regressive Biasing
by: Pynadath, Patrick, et al.
Published: (2025)
by: Pynadath, Patrick, et al.
Published: (2025)
Why Any-Order Autoregressive Models Need Two-Stream Attention: A Structural-Semantic Tradeoff
by: Pynadath, Patrick, et al.
Published: (2026)
by: Pynadath, Patrick, et al.
Published: (2026)
Adaptive Draft-Verification for Efficient Large Language Model Decoding
by: Liu, Xukun, et al.
Published: (2024)
by: Liu, Xukun, et al.
Published: (2024)
Tutorial Proposal: Speculative Decoding for Efficient LLM Inference
by: Xia, Heming, et al.
Published: (2025)
by: Xia, Heming, et al.
Published: (2025)
Superfiltering: Weak-to-Strong Data Filtering for Fast Instruction-Tuning
by: Li, Ming, et al.
Published: (2024)
by: Li, Ming, et al.
Published: (2024)
Similar Items
-
Cascade Reward Sampling for Efficient Decoding-Time Alignment
by: Li, Bolian, et al.
Published: (2024) -
ETA: Evaluating Then Aligning Safety of Vision Language Models at Inference Time
by: Ding, Yi, et al.
Published: (2024) -
Uniform-Correct Policy Optimization: Breaking RLVR's Indifference to Diversity
by: Lochab, Anamika, et al.
Published: (2026) -
Learning Self-Correction in Vision-Language Models via Rollout Augmentation
by: Ding, Yi, et al.
Published: (2026) -
Entropy-MCMC: Sampling from Flat Basins with Ease
by: Li, Bolian, et al.
Published: (2023)