Entropy After </Think> for reasoning model early exiting
Fuente:
arXiv
Saved in:
| Main Authors: | Wang, Xi, McInerney, James, Wang, Lequn, Kallus, Nathan |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Variation Due to Regularization Tractably Recovers Bayesian Deep Learning
by: McInerney, James, et al.
Published: (2024)
by: McInerney, James, et al.
Published: (2024)
Optimization of Epsilon-Greedy Exploration
by: Che, Ethan, et al.
Published: (2025)
by: Che, Ethan, et al.
Published: (2025)
Adjusting Regression Models for Conditional Uncertainty Calibration
by: Gao, Ruijiang, et al.
Published: (2024)
by: Gao, Ruijiang, et al.
Published: (2024)
Switching the Loss Reduces the Cost in Batch (Offline) Reinforcement Learning
by: Ayoub, Alex, et al.
Published: (2024)
by: Ayoub, Alex, et al.
Published: (2024)
A Statistical-Modelling Approach to Feedforward Neural Network Model Selection
by: McInerney, Andrew, et al.
Published: (2022)
by: McInerney, Andrew, et al.
Published: (2022)
Semiparametric Preference Optimization: Your Language Model is Secretly a Single-Index Model
by: Kallus, Nathan
Published: (2025)
by: Kallus, Nathan
Published: (2025)
Estimating Heterogeneous Treatment Effects by Combining Weak Instruments and Observational Data
by: Oprescu, Miruna, et al.
Published: (2024)
by: Oprescu, Miruna, et al.
Published: (2024)
On the role of surrogates in the efficient estimation of treatment effects with limited outcome data
by: Kallus, Nathan, et al.
Published: (2020)
by: Kallus, Nathan, et al.
Published: (2020)
The Central Role of the Loss Function in Reinforcement Learning
by: Wang, Kaiwen, et al.
Published: (2024)
by: Wang, Kaiwen, et al.
Published: (2024)
Robust and Agnostic Learning of Conditional Distributional Treatment Effects
by: Kallus, Nathan, et al.
Published: (2022)
by: Kallus, Nathan, et al.
Published: (2022)
A Reductions Approach to Risk-Sensitive Reinforcement Learning with Optimized Certainty Equivalents
by: Wang, Kaiwen, et al.
Published: (2024)
by: Wang, Kaiwen, et al.
Published: (2024)
Demistifying Inference after Adaptive Experiments
by: Bibaut, Aurélien, et al.
Published: (2024)
by: Bibaut, Aurélien, et al.
Published: (2024)
Stationary Reweighting Yields Local Convergence of Soft Fitted Q-Iteration
by: van der Laan, Lars, et al.
Published: (2025)
by: van der Laan, Lars, et al.
Published: (2025)
Fitted $Q$ Evaluation Without Bellman Completeness via Stationary Weighting
by: van der Laan, Lars, et al.
Published: (2025)
by: van der Laan, Lars, et al.
Published: (2025)
Multi-Armed Bandits with Interference
by: Jia, Su, et al.
Published: (2024)
by: Jia, Su, et al.
Published: (2024)
ERDE: Entropy-Regularized Distillation for Early-exit
by: Guidez, Martial, et al.
Published: (2025)
by: Guidez, Martial, et al.
Published: (2025)
Exploration in the Limit
by: Cho, Brian M., et al.
Published: (2025)
by: Cho, Brian M., et al.
Published: (2025)
Bellman Calibration for $V$-Learning in Offline Reinforcement Learning
by: van der Laan, Lars, et al.
Published: (2025)
by: van der Laan, Lars, et al.
Published: (2025)
Peeking with PEAK: Sequential, Nonparametric Composite Hypothesis Tests for Means of Multiple Data Streams
by: Cho, Brian, et al.
Published: (2024)
by: Cho, Brian, et al.
Published: (2024)
Long-term Causal Inference Under Persistent Confounding via Data Combination
by: Imbens, Guido, et al.
Published: (2022)
by: Imbens, Guido, et al.
Published: (2022)
Teaching Knowledge Management (SIG KM).
by: McInerney, Claire
Published: (2000)
by: McInerney, Claire
Published: (2000)
Confidence-gated training for efficient early-exit neural networks
by: Mokssit, Saad, et al.
Published: (2025)
by: Mokssit, Saad, et al.
Published: (2025)
The Context Gathering Decision Process: A POMDP Framework for Agentic Search
by: Kausik, Chinmaya, et al.
Published: (2026)
by: Kausik, Chinmaya, et al.
Published: (2026)
Low-Rank MDPs with Continuous Action Spaces
by: Bennett, Andrew, et al.
Published: (2023)
by: Bennett, Andrew, et al.
Published: (2023)
Is Cosine-Similarity of Embeddings Really About Similarity?
by: Steck, Harald, et al.
Published: (2024)
by: Steck, Harald, et al.
Published: (2024)
More Benefits of Being Distributional: Second-Order Bounds for Reinforcement Learning
by: Wang, Kaiwen, et al.
Published: (2024)
by: Wang, Kaiwen, et al.
Published: (2024)
Efficient Adaptive Experimentation with Noncompliance
by: Oprescu, Miruna, et al.
Published: (2025)
by: Oprescu, Miruna, et al.
Published: (2025)
Simulation-Based Inference for Adaptive Experiments
by: Cho, Brian M, et al.
Published: (2025)
by: Cho, Brian M, et al.
Published: (2025)
GAAVI: Global Asymptotic Anytime Valid Inference for the Conditional Mean Function
by: Cho, Brian M, et al.
Published: (2026)
by: Cho, Brian M, et al.
Published: (2026)
Functional Natural Policy Gradients
by: Bibaut, Aurelien, et al.
Published: (2026)
by: Bibaut, Aurelien, et al.
Published: (2026)
Clustered Switchback Designs for Experimentation Under Spatio-temporal Interference
by: Jia, Su, et al.
Published: (2023)
by: Jia, Su, et al.
Published: (2023)
Near-Optimal Non-Parametric Sequential Tests and Confidence Sequences with Possibly Dependent Observations
by: Bibaut, Aurelien, et al.
Published: (2022)
by: Bibaut, Aurelien, et al.
Published: (2022)
Environmental Scanning and the Information Manager.
by: Newsome, James, et al.
Published: (1990)
by: Newsome, James, et al.
Published: (1990)
Efficient and Sharp Off-Policy Evaluation in Robust Markov Decision Processes
by: Bennett, Andrew, et al.
Published: (2024)
by: Bennett, Andrew, et al.
Published: (2024)
Causal Inference on Networks under Misspecified Exposure Mappings: A Partial Identification Framework
by: Schröder, Maresa, et al.
Published: (2026)
by: Schröder, Maresa, et al.
Published: (2026)
Contextual Linear Optimization with Partial Feedback
by: Hu, Yichun, et al.
Published: (2024)
by: Hu, Yichun, et al.
Published: (2024)
Reward Maximization for Pure Exploration: Minimax Optimal Good Arm Identification for Nonparametric Multi-Armed Bandits
by: Cho, Brian, et al.
Published: (2024)
by: Cho, Brian, et al.
Published: (2024)
Efficient Inference for Inverse Reinforcement Learning and Dynamic Discrete Choice Models
by: van der Laan, Lars, et al.
Published: (2025)
by: van der Laan, Lars, et al.
Published: (2025)
Reproductive behaviour of the blackspotted stickleback, Gasterostomus wheatlandi
by: McInerney, J. E
Published: (1969)
by: McInerney, J. E
Published: (1969)
Justice, Complexity and Effective Governance in the Twenty-First Century
by: Thomas F. McInerney
Published: (2021)
by: Thomas F. McInerney
Published: (2021)
Similar Items
-
Variation Due to Regularization Tractably Recovers Bayesian Deep Learning
by: McInerney, James, et al.
Published: (2024) -
Optimization of Epsilon-Greedy Exploration
by: Che, Ethan, et al.
Published: (2025) -
Adjusting Regression Models for Conditional Uncertainty Calibration
by: Gao, Ruijiang, et al.
Published: (2024) -
Switching the Loss Reduces the Cost in Batch (Offline) Reinforcement Learning
by: Ayoub, Alex, et al.
Published: (2024) -
A Statistical-Modelling Approach to Feedforward Neural Network Model Selection
by: McInerney, Andrew, et al.
Published: (2022)