Reasoning with Sampling: Cutting at Decision Points
Fuente:
arXiv
Saved in:
| Main Authors: | Zhou, Felix, Mehrotra, Anay, Liu, Quanquan C. |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Transformers as Decision Makers: Provable In-Context Reinforcement Learning via Supervised Pretraining
by: Lin, Licong, et al.
Published: (2023)
by: Lin, Licong, et al.
Published: (2023)
Improved Guarantees for Heterogeneous Treatment-Effect Estimation via Matrix Completion
by: Mehrotra, Anay, et al.
Published: (2026)
by: Mehrotra, Anay, et al.
Published: (2026)
Differentially Private Language Generation and Identification in the Limit
by: Mehrotra, Anay, et al.
Published: (2026)
by: Mehrotra, Anay, et al.
Published: (2026)
Language Generation with Infinite Contamination
by: Mehrotra, Anay, et al.
Published: (2025)
by: Mehrotra, Anay, et al.
Published: (2025)
Reject, Resample, Repeat: Understanding Parallel Reasoning in Language Model Inference
by: Golowich, Noah, et al.
Published: (2026)
by: Golowich, Noah, et al.
Published: (2026)
Can SGD Select Good Fishermen? Local Convergence under Self-Selection Biases and Beyond
by: Kalavasis, Alkis, et al.
Published: (2025)
by: Kalavasis, Alkis, et al.
Published: (2025)
On the Optimal Sample Complexity of Offline Multi-Armed Bandits with KL Regularization
by: Ji, Kaixuan, et al.
Published: (2026)
by: Ji, Kaixuan, et al.
Published: (2026)
On Characterizations for Language Generation: Interplay of Hallucinations, Breadth, and Stability
by: Kalavasis, Alkis, et al.
Published: (2024)
by: Kalavasis, Alkis, et al.
Published: (2024)
On the Limits of Language Generation: Trade-Offs Between Hallucination and Mode Collapse
by: Kalavasis, Alkis, et al.
Published: (2024)
by: Kalavasis, Alkis, et al.
Published: (2024)
Counterfactual reasoning: an analysis of in-context emergence
by: Miller, Moritz, et al.
Published: (2025)
by: Miller, Moritz, et al.
Published: (2025)
Don't Pass@k: A Bayesian Framework for Large Language Model Evaluation
by: Hariri, Mohsen, et al.
Published: (2025)
by: Hariri, Mohsen, et al.
Published: (2025)
Towards Efficient Online Exploration for Reinforcement Learning with Human Feedback
by: Li, Gen, et al.
Published: (2025)
by: Li, Gen, et al.
Published: (2025)
Is a Good Foundation Necessary for Efficient Reinforcement Learning? The Computational Role of the Base Model in Exploration
by: Foster, Dylan J., et al.
Published: (2025)
by: Foster, Dylan J., et al.
Published: (2025)
The Coverage Principle: How Pre-Training Enables Post-Training
by: Chen, Fan, et al.
Published: (2025)
by: Chen, Fan, et al.
Published: (2025)
Unveiling the Statistical Foundations of Chain-of-Thought Prompting Methods
by: Hu, Xinyang, et al.
Published: (2024)
by: Hu, Xinyang, et al.
Published: (2024)
Smoothed Analysis of Learning from Positive Samples
by: Lee, Jane H., et al.
Published: (2025)
by: Lee, Jane H., et al.
Published: (2025)
Efficient Statistics With Unknown Truncation, Polynomial Time Algorithms, Beyond Gaussians
by: Lee, Jane H., et al.
Published: (2024)
by: Lee, Jane H., et al.
Published: (2024)
Sample Complexity of Bias Detection with Subsampled Point-to-Subspace Distances
by: Matilla, German Martinez, et al.
Published: (2025)
by: Matilla, German Martinez, et al.
Published: (2025)
Near-Optimal Regret for KL-Regularized Multi-Armed Bandits
by: Ji, Kaixuan, et al.
Published: (2026)
by: Ji, Kaixuan, et al.
Published: (2026)
Towards a Sharp Analysis of Offline Policy Learning for $f$-Divergence-Regularized Contextual Bandits
by: Zhao, Qingyue, et al.
Published: (2025)
by: Zhao, Qingyue, et al.
Published: (2025)
On Language Generation in the Limit with Bounded Memory
by: Kleinberg, Jon, et al.
Published: (2026)
by: Kleinberg, Jon, et al.
Published: (2026)
Fast Rates for Offline Contextual Bandits with Forward-KL Regularization under Single-Policy Concentrability
by: Zhao, Qingyue, et al.
Published: (2026)
by: Zhao, Qingyue, et al.
Published: (2026)
What is Learnable in Valiant's Theory of the Learnable?
by: Hanneke, Steve, et al.
Published: (2026)
by: Hanneke, Steve, et al.
Published: (2026)
Smaller Confidence Intervals From IPW Estimators via Data-Dependent Coarsening
by: Kalavasis, Alkis, et al.
Published: (2024)
by: Kalavasis, Alkis, et al.
Published: (2024)
Mean Estimation from Coarse Data: Characterizations and Efficient Algorithms
by: Kalavasis, Alkis, et al.
Published: (2026)
by: Kalavasis, Alkis, et al.
Published: (2026)
A Statistical Hypothesis Testing Framework for Data Misappropriation Detection in Large Language Models
by: Cai, Yinpeng, et al.
Published: (2025)
by: Cai, Yinpeng, et al.
Published: (2025)
Retrieval-Augmented Generation as Noisy In-Context Learning: A Unified Theory and Risk Bounds
by: Guo, Yang, et al.
Published: (2025)
by: Guo, Yang, et al.
Published: (2025)
A Note on Non-Negative $L_1$-Approximating Polynomials
by: Lee, Jane H., et al.
Published: (2026)
by: Lee, Jane H., et al.
Published: (2026)
Tree of Attacks: Jailbreaking Black-Box LLMs Automatically
by: Mehrotra, Anay, et al.
Published: (2023)
by: Mehrotra, Anay, et al.
Published: (2023)
Adaptive Sample Aggregation In Transfer Learning
by: Hanneke, Steve, et al.
Published: (2024)
by: Hanneke, Steve, et al.
Published: (2024)
Diffusion Posterior Sampling is Computationally Intractable
by: Gupta, Shivam, et al.
Published: (2024)
by: Gupta, Shivam, et al.
Published: (2024)
On the Design of KL-Regularized Policy Gradient Algorithms for LLM Reasoning
by: Zhang, Yifan, et al.
Published: (2025)
by: Zhang, Yifan, et al.
Published: (2025)
Fair Classification with Partial Feedback: An Exploration-Based Data Collection Approach
by: Keswani, Vijay, et al.
Published: (2024)
by: Keswani, Vijay, et al.
Published: (2024)
On the Provable Performance Guarantee of Efficient Reasoning Models
by: Zeng, Hao, et al.
Published: (2025)
by: Zeng, Hao, et al.
Published: (2025)
Linear Regression with Unknown Truncation Beyond Gaussian Features
by: Kouridakis, Alexandros, et al.
Published: (2026)
by: Kouridakis, Alexandros, et al.
Published: (2026)
DDPM Score Matching and Distribution Learning
by: Chewi, Sinho, et al.
Published: (2025)
by: Chewi, Sinho, et al.
Published: (2025)
Unified Algorithms for RL with Decision-Estimation Coefficients: PAC, Reward-Free, Preference-Based Learning, and Beyond
by: Chen, Fan, et al.
Published: (2022)
by: Chen, Fan, et al.
Published: (2022)
The Good, the Bad, and the Sampled: a No-Regret Approach to Safe Online Classification
by: Baharav, Tavor Z., et al.
Published: (2025)
by: Baharav, Tavor Z., et al.
Published: (2025)
Causal Sufficiency and Necessity Improves Chain-of-Thought Reasoning
by: Yu, Xiangning, et al.
Published: (2025)
by: Yu, Xiangning, et al.
Published: (2025)
Decision Making in Changing Environments: Robustness, Query-Based Learning, and Differential Privacy
by: Chen, Fan, et al.
Published: (2025)
by: Chen, Fan, et al.
Published: (2025)
Similar Items
-
Transformers as Decision Makers: Provable In-Context Reinforcement Learning via Supervised Pretraining
by: Lin, Licong, et al.
Published: (2023) -
Improved Guarantees for Heterogeneous Treatment-Effect Estimation via Matrix Completion
by: Mehrotra, Anay, et al.
Published: (2026) -
Differentially Private Language Generation and Identification in the Limit
by: Mehrotra, Anay, et al.
Published: (2026) -
Language Generation with Infinite Contamination
by: Mehrotra, Anay, et al.
Published: (2025) -
Reject, Resample, Repeat: Understanding Parallel Reasoning in Language Model Inference
by: Golowich, Noah, et al.
Published: (2026)