More Bang for the Buck: Improving the Inference of Large Language Models at a Fixed Budget using Reset and Discard (ReD)
Fuente:
arXiv
Saved in:
| Main Authors: | Meir, Sagi, Keidar, Tommer D., Levi, Noam, Reuveni, Shlomi, Hirshberg, Barak |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
First-Passage Approach to Optimizing Perturbations for Improved Training of Machine Learning Models
by: Meir, Sagi, et al.
Published: (2025)
by: Meir, Sagi, et al.
Published: (2025)
Adaptive Resetting for Informed Search Strategies and the Design of Non-equilibrium Steady-states
by: Keidar, Tommer D., et al.
Published: (2024)
by: Keidar, Tommer D., et al.
Published: (2024)
Accelerating Molecular Dynamics through Informed Resetting
by: Church, Jonathan R., et al.
Published: (2024)
by: Church, Jonathan R., et al.
Published: (2024)
Universal Linear Response of First-Passage Kinetics: A Framework for Prediction and Inference
by: Keidar, Tommer D., et al.
Published: (2024)
by: Keidar, Tommer D., et al.
Published: (2024)
Learning Shrinks the Hard Tail: Training-Dependent Inference Scaling in a Solvable Linear Model
by: Levi, Noam
Published: (2026)
by: Levi, Noam
Published: (2026)
Short-Time Infrequent Metadynamics for Improved Kinetics Inference
by: Blumer, Ofir, et al.
Published: (2024)
by: Blumer, Ofir, et al.
Published: (2024)
Smart Resetting: An Energy-Efficient Strategy for Stochastic Search Processes
by: Tal-Friedman, Ofir, et al.
Published: (2024)
by: Tal-Friedman, Ofir, et al.
Published: (2024)
Inference of non-exponential kinetics through stochastic resetting
by: Blumer, Ofir, et al.
Published: (2024)
by: Blumer, Ofir, et al.
Published: (2024)
Probing the Latent Hierarchical Structure of Data via Diffusion Models
by: Sclocchi, Antonio, et al.
Published: (2024)
by: Sclocchi, Antonio, et al.
Published: (2024)
Grokking at the Edge of Linear Separability
by: Beck, Alon, et al.
Published: (2024)
by: Beck, Alon, et al.
Published: (2024)
Grokking in Linear Estimators -- A Solvable Model that Groks without Understanding
by: Levi, Noam, et al.
Published: (2023)
by: Levi, Noam, et al.
Published: (2023)
The Underlying Scaling Laws and Universal Statistical Structure of Complex Datasets
by: Levi, Noam, et al.
Published: (2023)
by: Levi, Noam, et al.
Published: (2023)
Qubit dephasing by spectrally diffusing quantum two-level systems
by: Matityahu, Shlomi, et al.
Published: (2023)
by: Matityahu, Shlomi, et al.
Published: (2023)
Sampling Data with Chains of Forward-Backward Diffusion Steps
by: Kang, Hyunmo, et al.
Published: (2026)
by: Kang, Hyunmo, et al.
Published: (2026)
Continuous Specialization Transition in the Soft Committee Machine with ReLU Activation
by: Afanah, Assem, et al.
Published: (2026)
by: Afanah, Assem, et al.
Published: (2026)
Thermally activated particle motion in biased correlated Gaussian disorder potentials
by: Valov, Alexander, et al.
Published: (2024)
by: Valov, Alexander, et al.
Published: (2024)
Improving deep neural network performance through sampling
by: Ghantasala, Lakshmi A., et al.
Published: (2025)
by: Ghantasala, Lakshmi A., et al.
Published: (2025)
Neuromodulation via Krotov-Hopfield Improves Accuracy and Robustness of RBMs
by: Tambaş, Başer, et al.
Published: (2025)
by: Tambaş, Başer, et al.
Published: (2025)
Exact Fixed-Point Constraints in Neural-ODEs with Provable Universality
by: Pacifico, Feliciano Giuseppe, et al.
Published: (2026)
by: Pacifico, Feliciano Giuseppe, et al.
Published: (2026)
Counting and Hardness-of-Finding Fixed Points in Cellular Automata on Random Graphs
by: Koller, Cédric, et al.
Published: (2024)
by: Koller, Cédric, et al.
Published: (2024)
Critical Phase Transition in Large Language Models
by: Nakaishi, Kai, et al.
Published: (2024)
by: Nakaishi, Kai, et al.
Published: (2024)
Large time effective kinetics $β$-functions for quantum (2+p)-spin glass
by: Lahoche, Vincent, et al.
Published: (2024)
by: Lahoche, Vincent, et al.
Published: (2024)
Re-entrant phase transitions induced by localization of zero-modes
by: Morone, Flaviano, et al.
Published: (2024)
by: Morone, Flaviano, et al.
Published: (2024)
Re-entrant localization induced by short-range hopping in the fractal Rosenzweig-Porter Model
by: Ghosh, Roopayan, et al.
Published: (2024)
by: Ghosh, Roopayan, et al.
Published: (2024)
Population Level Activity in Large Random Neural Networks
by: MacLaurin, James, et al.
Published: (2024)
by: MacLaurin, James, et al.
Published: (2024)
Stochastic Resetting Accelerates Policy Convergence in Reinforcement Learning
by: Zhou, Jello, et al.
Published: (2026)
by: Zhou, Jello, et al.
Published: (2026)
Sharp description of local minima in the loss landscape of high-dimensional two-layer ReLU neural networks
by: Huang, Jie, et al.
Published: (2026)
by: Huang, Jie, et al.
Published: (2026)
Financial instability transition under heterogeneous investments and portfolio diversification
by: Forer, Preben, et al.
Published: (2025)
by: Forer, Preben, et al.
Published: (2025)
Top eigenpair statistics of diluted Wishart matrices
by: Budnick, Barak, et al.
Published: (2025)
by: Budnick, Barak, et al.
Published: (2025)
Predictive Coding Networks and Inference Learning: Tutorial and Survey
by: van Zwol, Björn, et al.
Published: (2024)
by: van Zwol, Björn, et al.
Published: (2024)
Injectivity of ReLU networks: perspectives from statistical physics
by: Maillard, Antoine, et al.
Published: (2023)
by: Maillard, Antoine, et al.
Published: (2023)
How Feature Learning Can Improve Neural Scaling Laws
by: Bordelon, Blake, et al.
Published: (2024)
by: Bordelon, Blake, et al.
Published: (2024)
Injectivity capacity of ReLU gates
by: Stojnic, Mihailo
Published: (2024)
by: Stojnic, Mihailo
Published: (2024)
Evidence of Scaling Regimes in the Hopfield Dynamics of Whole Brain Model
by: Gosti, Giorgio, et al.
Published: (2024)
by: Gosti, Giorgio, et al.
Published: (2024)
Rare Event Analysis of Large Language Models
by: Dorman, Jake McAllister, et al.
Published: (2026)
by: Dorman, Jake McAllister, et al.
Published: (2026)
Learning the Electronic Hamiltonian of Large Atomic Structures
by: Xia, Chen Hao, et al.
Published: (2025)
by: Xia, Chen Hao, et al.
Published: (2025)
When Less is More: Approximating the Quantum Geometric Tensor with Block Structures
by: Shokry, Ahmedeo, et al.
Published: (2025)
by: Shokry, Ahmedeo, et al.
Published: (2025)
Robustness of the Random Language Model
by: Lalegani, Fatemeh, et al.
Published: (2023)
by: Lalegani, Fatemeh, et al.
Published: (2023)
Multi-Output Convolutional Neural Network for Improved Parameter Extraction in Time-Resolved Electrostatic Force Microscopy Data
by: Breshears, Madeleine D., et al.
Published: (2025)
by: Breshears, Madeleine D., et al.
Published: (2025)
Electric field-dependent conductivity as probe for charge carrier delocalization and morphology in organic semiconductors
by: Shokrani, Morteza, et al.
Published: (2025)
by: Shokrani, Morteza, et al.
Published: (2025)
Similar Items
-
First-Passage Approach to Optimizing Perturbations for Improved Training of Machine Learning Models
by: Meir, Sagi, et al.
Published: (2025) -
Adaptive Resetting for Informed Search Strategies and the Design of Non-equilibrium Steady-states
by: Keidar, Tommer D., et al.
Published: (2024) -
Accelerating Molecular Dynamics through Informed Resetting
by: Church, Jonathan R., et al.
Published: (2024) -
Universal Linear Response of First-Passage Kinetics: A Framework for Prediction and Inference
by: Keidar, Tommer D., et al.
Published: (2024) -
Learning Shrinks the Hard Tail: Training-Dependent Inference Scaling in a Solvable Linear Model
by: Levi, Noam
Published: (2026)