Sampling-Efficient Test-Time Scaling: Self-Estimating the Best-of-N Sampling in Early Decoding
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Wang, Yiming, Zhang, Pei, Huang, Siyuan, Yang, Baosong, Zhang, Zhuosheng, Huang, Fei, Wang, Rui |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
On the Overscaling Curse of Parallel Thinking: System Efficacy Contradicts Sample Efficiency
von: Wang, Yiming, et al.
Veröffentlicht: (2026)
von: Wang, Yiming, et al.
Veröffentlicht: (2026)
Embedding Trajectory for Out-of-Distribution Detection in Mathematical Reasoning
von: Wang, Yiming, et al.
Veröffentlicht: (2024)
von: Wang, Yiming, et al.
Veröffentlicht: (2024)
Latent Space Chain-of-Embedding Enables Output-free LLM Self-Evaluation
von: Wang, Yiming, et al.
Veröffentlicht: (2024)
von: Wang, Yiming, et al.
Veröffentlicht: (2024)
Adaptive Rectification Sampling for Test-Time Compute Scaling
von: Tan, Zhendong, et al.
Veröffentlicht: (2025)
von: Tan, Zhendong, et al.
Veröffentlicht: (2025)
Meta-Reasoning: Semantics-Symbol Deconstruction for Large Language Models
von: Wang, Yiming, et al.
Veröffentlicht: (2023)
von: Wang, Yiming, et al.
Veröffentlicht: (2023)
TreeBoN: Enhancing Inference-Time Alignment with Speculative Tree-Search and Best-of-N Sampling
von: Qiu, Jiahao, et al.
Veröffentlicht: (2024)
von: Qiu, Jiahao, et al.
Veröffentlicht: (2024)
Efficient Test-Time Scaling via Self-Calibration
von: Huang, Chengsong, et al.
Veröffentlicht: (2025)
von: Huang, Chengsong, et al.
Veröffentlicht: (2025)
CultureSynth: A Hierarchical Taxonomy-Guided and Retrieval-Augmented Framework for Cultural Question-Answer Synthesis
von: Zhang, Xinyu, et al.
Veröffentlicht: (2025)
von: Zhang, Xinyu, et al.
Veröffentlicht: (2025)
Iterative Deepening Sampling as Efficient Test-Time Scaling
von: Chen, Weizhe, et al.
Veröffentlicht: (2025)
von: Chen, Weizhe, et al.
Veröffentlicht: (2025)
Scaling Up, Speeding Up: A Benchmark of Speculative Decoding for Efficient LLM Test-Time Scaling
von: Sun, Shengyin, et al.
Veröffentlicht: (2025)
von: Sun, Shengyin, et al.
Veröffentlicht: (2025)
Satori-SWE: Evolutionary Test-Time Scaling for Sample-Efficient Software Engineering
von: Zeng, Guangtao, et al.
Veröffentlicht: (2025)
von: Zeng, Guangtao, et al.
Veröffentlicht: (2025)
Mining Intrinsic Rewards from LLM Hidden States for Efficient Best-of-N Sampling
von: Guo, Jizhou, et al.
Veröffentlicht: (2025)
von: Guo, Jizhou, et al.
Veröffentlicht: (2025)
Speculative Decoding for Multi-Sample Inference
von: Li, Yiwei, et al.
Veröffentlicht: (2025)
von: Li, Yiwei, et al.
Veröffentlicht: (2025)
On the Role of Temperature Sampling in Test-Time Scaling
von: Wu, Yuheng, et al.
Veröffentlicht: (2025)
von: Wu, Yuheng, et al.
Veröffentlicht: (2025)
Scaling LLM Inference with Optimized Sample Compute Allocation
von: Zhang, Kexun, et al.
Veröffentlicht: (2024)
von: Zhang, Kexun, et al.
Veröffentlicht: (2024)
Reward Difference Optimization For Sample Reweighting In Offline RLHF
von: Wang, Shiqi, et al.
Veröffentlicht: (2024)
von: Wang, Shiqi, et al.
Veröffentlicht: (2024)
MetaScale: Test-Time Scaling with Evolving Meta-Thoughts
von: Liu, Qin, et al.
Veröffentlicht: (2025)
von: Liu, Qin, et al.
Veröffentlicht: (2025)
Towards Cross-lingual Values Judgment: A Consensus-Pluralism Perspective
von: Chen, Yukun, et al.
Veröffentlicht: (2026)
von: Chen, Yukun, et al.
Veröffentlicht: (2026)
Self-Prompting Large Language Models for Zero-Shot Open-Domain QA
von: Li, Junlong, et al.
Veröffentlicht: (2022)
von: Li, Junlong, et al.
Veröffentlicht: (2022)
DynScaling: Efficient Verifier-free Inference Scaling via Dynamic and Integrated Sampling
von: Wang, Fei, et al.
Veröffentlicht: (2025)
von: Wang, Fei, et al.
Veröffentlicht: (2025)
Improving Machine Translation with Human Feedback: An Exploration of Quality Estimation as a Reward Model
von: He, Zhiwei, et al.
Veröffentlicht: (2024)
von: He, Zhiwei, et al.
Veröffentlicht: (2024)
MatryoshkaThinking: Recursive Test-Time Scaling Enables Efficient Reasoning
von: Chen, Hongwei, et al.
Veröffentlicht: (2025)
von: Chen, Hongwei, et al.
Veröffentlicht: (2025)
Enhancing Persona Following at Decoding Time via Dynamic Importance Estimation for Role-Playing Agents
von: Liu, Yuxin, et al.
Veröffentlicht: (2026)
von: Liu, Yuxin, et al.
Veröffentlicht: (2026)
Gracefully Filtering Backdoor Samples for Generative Large Language Models without Retraining
von: Wu, Zongru, et al.
Veröffentlicht: (2024)
von: Wu, Zongru, et al.
Veröffentlicht: (2024)
Scaling Laws for Speculative Decoding
von: Yan, Siyuan, et al.
Veröffentlicht: (2025)
von: Yan, Siyuan, et al.
Veröffentlicht: (2025)
Regularized Best-of-N Sampling with Minimum Bayes Risk Objective for Language Model Alignment
von: Jinnai, Yuu, et al.
Veröffentlicht: (2024)
von: Jinnai, Yuu, et al.
Veröffentlicht: (2024)
FlashSampling: Fast and Memory-Efficient Exact Sampling
von: Ruiz, Tomas, et al.
Veröffentlicht: (2026)
von: Ruiz, Tomas, et al.
Veröffentlicht: (2026)
Human-Instruction-Free LLM Self-Alignment with Limited Samples
von: Guo, Hongyi, et al.
Veröffentlicht: (2024)
von: Guo, Hongyi, et al.
Veröffentlicht: (2024)
Language Control Diffusion: Efficiently Scaling through Space, Time, and Tasks
von: Zhang, Edwin, et al.
Veröffentlicht: (2022)
von: Zhang, Edwin, et al.
Veröffentlicht: (2022)
SLOT: Sample-specific Language Model Optimization at Test-time
von: Hu, Yang, et al.
Veröffentlicht: (2025)
von: Hu, Yang, et al.
Veröffentlicht: (2025)
Time-Annealed Perturbation Sampling: Diverse Generation for Diffusion Language Models
von: Wu, Jingxuan, et al.
Veröffentlicht: (2026)
von: Wu, Jingxuan, et al.
Veröffentlicht: (2026)
Inference-Aware Fine-Tuning for Best-of-N Sampling in Large Language Models
von: Chow, Yinlam, et al.
Veröffentlicht: (2024)
von: Chow, Yinlam, et al.
Veröffentlicht: (2024)
The Diminishing Returns of Early-Exit Decoding in Modern LLMs
von: Wei, Rui, et al.
Veröffentlicht: (2026)
von: Wei, Rui, et al.
Veröffentlicht: (2026)
AdaFuse: Adaptive Ensemble Decoding with Test-Time Scaling for LLMs
von: Cui, Chengming, et al.
Veröffentlicht: (2026)
von: Cui, Chengming, et al.
Veröffentlicht: (2026)
TTCS: Test-Time Curriculum Synthesis for Self-Evolving
von: Yang, Chengyi, et al.
Veröffentlicht: (2026)
von: Yang, Chengyi, et al.
Veröffentlicht: (2026)
SETS: Leveraging Self-Verification and Self-Correction for Improved Test-Time Scaling
von: Chen, Jiefeng, et al.
Veröffentlicht: (2025)
von: Chen, Jiefeng, et al.
Veröffentlicht: (2025)
S$^4$C: Speculative Sampling with Syntactic and Semantic Coherence for Efficient Inference of Large Language Models
von: He, Tao, et al.
Veröffentlicht: (2025)
von: He, Tao, et al.
Veröffentlicht: (2025)
END: Early Noise Dropping for Efficient and Effective Context Denoising
von: Jin, Hongye, et al.
Veröffentlicht: (2025)
von: Jin, Hongye, et al.
Veröffentlicht: (2025)
Sentinel: Decoding Context Utilization via Attention Probing for Efficient LLM Context Compression
von: Zhang, Yong, et al.
Veröffentlicht: (2025)
von: Zhang, Yong, et al.
Veröffentlicht: (2025)
Adaptive Decoding via Test-Time Policy Learning for Self-Improving Generation
von: Bhardwaj, Asmita, et al.
Veröffentlicht: (2026)
von: Bhardwaj, Asmita, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
On the Overscaling Curse of Parallel Thinking: System Efficacy Contradicts Sample Efficiency
von: Wang, Yiming, et al.
Veröffentlicht: (2026) -
Embedding Trajectory for Out-of-Distribution Detection in Mathematical Reasoning
von: Wang, Yiming, et al.
Veröffentlicht: (2024) -
Latent Space Chain-of-Embedding Enables Output-free LLM Self-Evaluation
von: Wang, Yiming, et al.
Veröffentlicht: (2024) -
Adaptive Rectification Sampling for Test-Time Compute Scaling
von: Tan, Zhendong, et al.
Veröffentlicht: (2025) -
Meta-Reasoning: Semantics-Symbol Deconstruction for Large Language Models
von: Wang, Yiming, et al.
Veröffentlicht: (2023)