More Test-Time Compute Can Hurt: Overestimation Bias in LLM Beam Search
Fuente:
arXiv
Saved in:
| Main Authors: | Dalal, Gal, Hallak, Assaf, Chechik, Gal, Ziser, Yftah |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Policy Gradient with Tree Expansion
by: Dalal, Gal, et al.
Published: (2023)
by: Dalal, Gal, et al.
Published: (2023)
Policy Optimized Text-to-Image Pipeline Design
by: Gadot, Uri, et al.
Published: (2025)
by: Gadot, Uri, et al.
Published: (2025)
PlaMo: Plan and Move in Rich 3D Physical Environments
by: Hallak, Assaf, et al.
Published: (2024)
by: Hallak, Assaf, et al.
Published: (2024)
Implementing Reinforcement Learning Datacenter Congestion Control in NVIDIA NICs
by: Fuhrer, Benjamin, et al.
Published: (2022)
by: Fuhrer, Benjamin, et al.
Published: (2022)
AutoLoop: Fast Visual SLAM Fine-tuning through Agentic Curriculum Learning
by: Lahiany, Assaf, et al.
Published: (2025)
by: Lahiany, Assaf, et al.
Published: (2025)
From Actions to Words: Towards Abstractive-Textual Policy Summarization in RL
by: Admoni, Sahar, et al.
Published: (2025)
by: Admoni, Sahar, et al.
Published: (2025)
Tree Search-Based Policy Optimization under Stochastic Execution Delay
by: Valensi, David, et al.
Published: (2024)
by: Valensi, David, et al.
Published: (2024)
RL-RC-DoT: A Block-level RL agent for Task-Aware Video Compression
by: Gadot, Uri, et al.
Published: (2025)
by: Gadot, Uri, et al.
Published: (2025)
Efficient GNN Training Through Structure-Aware Randomized Mini-Batching
by: Balaji, Vignesh, et al.
Published: (2025)
by: Balaji, Vignesh, et al.
Published: (2025)
Knowing Before Saying: LLM Representations Encode Information About Chain-of-Thought Success Before Completion
by: Afzal, Anum, et al.
Published: (2025)
by: Afzal, Anum, et al.
Published: (2025)
Assessing Image Quality Using a Simple Generative Representation
by: Raviv, Simon, et al.
Published: (2024)
by: Raviv, Simon, et al.
Published: (2024)
Gradient Boosting Reinforcement Learning
by: Fuhrer, Benjamin, et al.
Published: (2024)
by: Fuhrer, Benjamin, et al.
Published: (2024)
Beyond Next Token Probabilities: Learnable, Fast Detection of Hallucinations and Data Contamination on LLM Output Distributions
by: Bar-Shalom, Guy, et al.
Published: (2025)
by: Bar-Shalom, Guy, et al.
Published: (2025)
Add-it: Training-Free Object Insertion in Images With Pretrained Diffusion Models
by: Tewel, Yoad, et al.
Published: (2024)
by: Tewel, Yoad, et al.
Published: (2024)
Spanning the Visual Analogy Space with a Weight Basis of LoRAs
by: Manor, Hila, et al.
Published: (2026)
by: Manor, Hila, et al.
Published: (2026)
Training-Free Consistent Text-to-Image Generation
by: Tewel, Yoad, et al.
Published: (2024)
by: Tewel, Yoad, et al.
Published: (2024)
Spectral Editing of Activations for Large Language Model Alignment
by: Qiu, Yifu, et al.
Published: (2024)
by: Qiu, Yifu, et al.
Published: (2024)
Stabilizing Policy Gradients for Sample-Efficient Reinforcement Learning in LLM Reasoning
by: Melo, Luckeciano C., et al.
Published: (2025)
by: Melo, Luckeciano C., et al.
Published: (2025)
Improved Generalization of Weight Space Networks via Augmentations
by: Shamsian, Aviv, et al.
Published: (2024)
by: Shamsian, Aviv, et al.
Published: (2024)
Exploration-Driven Policy Optimization in RLHF: Theoretical Insights on Efficient Data Utilization
by: Du, Yihan, et al.
Published: (2024)
by: Du, Yihan, et al.
Published: (2024)
Self-Improving World Modelling with Latent Actions
by: Qiu, Yifu, et al.
Published: (2026)
by: Qiu, Yifu, et al.
Published: (2026)
When More Thinking Hurts: Overthinking in LLM Test-Time Compute Scaling
by: Zhou, Shu, et al.
Published: (2026)
by: Zhou, Shu, et al.
Published: (2026)
Aligning What LLMs Do and Say: Towards Self-Consistent Explanations
by: Admoni, Sahar, et al.
Published: (2025)
by: Admoni, Sahar, et al.
Published: (2025)
Reliable Policy Iteration: Performance Robustness Across Architecture and Environment Perturbations
by: Eshwar, S. R., et al.
Published: (2025)
by: Eshwar, S. R., et al.
Published: (2025)
WMINet: A Wheel-Mounted Inertial Learning Approach For Mobile-Robot Positioning
by: Versano, Gal, et al.
Published: (2025)
by: Versano, Gal, et al.
Published: (2025)
Adversarial Robustness Overestimation and Instability in TRADES
by: Li, Jonathan Weiping, et al.
Published: (2024)
by: Li, Jonathan Weiping, et al.
Published: (2024)
Simple Baselines are Competitive with Code Evolution
by: Gideoni, Yonatan, et al.
Published: (2026)
by: Gideoni, Yonatan, et al.
Published: (2026)
DDTR: Diffusion Denoising Trace Recovery
by: Matyash, Maximilian, et al.
Published: (2025)
by: Matyash, Maximilian, et al.
Published: (2025)
Superscopes: Amplifying Internal Feature Representations for Language Model Interpretation
by: Jacobi, Jonathan, et al.
Published: (2025)
by: Jacobi, Jonathan, et al.
Published: (2025)
Leveraging Deep Learning for Physical Model Bias of Global Air Quality Estimates
by: Doerksen, Kelsey, et al.
Published: (2025)
by: Doerksen, Kelsey, et al.
Published: (2025)
Large Language Models for Water Distribution Systems Modeling and Decision-Making
by: Goldshtein, Yinon, et al.
Published: (2025)
by: Goldshtein, Yinon, et al.
Published: (2025)
Temporal-Difference Variational Continual Learning
by: Melo, Luckeciano C., et al.
Published: (2024)
by: Melo, Luckeciano C., et al.
Published: (2024)
Split and Conquer Partial Deepfake Speech
by: Rimon, Inbal, et al.
Published: (2026)
by: Rimon, Inbal, et al.
Published: (2026)
Suppressing Overestimation in Q-Learning through Adversarial Behaviors
by: Lee, HyeAnn, et al.
Published: (2023)
by: Lee, HyeAnn, et al.
Published: (2023)
Reasoning Introduces New Poisoning Attacks Yet Makes Them More Complicated
by: Foerster, Hanna, et al.
Published: (2025)
by: Foerster, Hanna, et al.
Published: (2025)
Revisiting Meta-Learning with Noisy Labels: Reweighting Dynamics and Theoretical Guarantees
by: Zhang, Yiming, et al.
Published: (2025)
by: Zhang, Yiming, et al.
Published: (2025)
Robust Monocular Visual Odometry using Curriculum Learning
by: Lahiany, Assaf, et al.
Published: (2024)
by: Lahiany, Assaf, et al.
Published: (2024)
Key-Locked Rank One Editing for Text-to-Image Personalization
by: Tewel, Yoad, et al.
Published: (2023)
by: Tewel, Yoad, et al.
Published: (2023)
From Glucose Patterns to Health Outcomes: A Generalizable Foundation Model for Continuous Glucose Monitor Data Analysis
by: Lutsker, Guy, et al.
Published: (2024)
by: Lutsker, Guy, et al.
Published: (2024)
The Pitfalls of Memorization: When Memorization Hurts Generalization
by: Bayat, Reza, et al.
Published: (2024)
by: Bayat, Reza, et al.
Published: (2024)
Similar Items
-
Policy Gradient with Tree Expansion
by: Dalal, Gal, et al.
Published: (2023) -
Policy Optimized Text-to-Image Pipeline Design
by: Gadot, Uri, et al.
Published: (2025) -
PlaMo: Plan and Move in Rich 3D Physical Environments
by: Hallak, Assaf, et al.
Published: (2024) -
Implementing Reinforcement Learning Datacenter Congestion Control in NVIDIA NICs
by: Fuhrer, Benjamin, et al.
Published: (2022) -
AutoLoop: Fast Visual SLAM Fine-tuning through Agentic Curriculum Learning
by: Lahiany, Assaf, et al.
Published: (2025)