Zero-Overhead Introspection for Adaptive Test-Time Compute
Fuente:
arXiv
Saved in:
| Main Authors: | Manvi, Rohin, Hong, Joey, Seyde, Tim, Labonne, Maxime, Lechner, Mathias, Levine, Sergey |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Adaptive Inference-Time Compute: LLMs Can Predict if They Can Do Better, Even Mid-Generation
by: Manvi, Rohin, et al.
Published: (2024)
by: Manvi, Rohin, et al.
Published: (2024)
Q-SFT: Q-Learning for Language Models via Supervised Fine-Tuning
by: Hong, Joey, et al.
Published: (2024)
by: Hong, Joey, et al.
Published: (2024)
Interactive Dialogue Agents via Reinforcement Learning on Hindsight Regenerations
by: Hong, Joey, et al.
Published: (2024)
by: Hong, Joey, et al.
Published: (2024)
Large Language Models are Geographically Biased
by: Manvi, Rohin, et al.
Published: (2024)
by: Manvi, Rohin, et al.
Published: (2024)
Merging in a Bottle: Differentiable Adaptive Merging (DAM) and the Path from Averaging to Automation
by: Gauthier-Caron, Thomas, et al.
Published: (2024)
by: Gauthier-Caron, Thomas, et al.
Published: (2024)
CARE: Decoding Time Safety Alignment via Rollback and Introspection Intervention
by: Hu, Xiaomeng, et al.
Published: (2025)
by: Hu, Xiaomeng, et al.
Published: (2025)
IntroLM: Introspective Language Models via Prefilling-Time Self-Evaluation
by: Kasnavieh, Hossein Hosseini, et al.
Published: (2026)
by: Kasnavieh, Hossein Hosseini, et al.
Published: (2026)
BEST-Route: Adaptive LLM Routing with Test-Time Optimal Compute
by: Ding, Dujian, et al.
Published: (2025)
by: Ding, Dujian, et al.
Published: (2025)
Planning without Search: Refining Frontier LLMs with Offline Goal-Conditioned RL
by: Hong, Joey, et al.
Published: (2025)
by: Hong, Joey, et al.
Published: (2025)
Introspective Planning: Aligning Robots' Uncertainty with Inherent Task Ambiguity
by: Liang, Kaiqu, et al.
Published: (2024)
by: Liang, Kaiqu, et al.
Published: (2024)
Beyond Introspection: Reinforcing Thinking via Externalist Behavioral Feedback
by: Yang, Diji, et al.
Published: (2024)
by: Yang, Diji, et al.
Published: (2024)
QJL: 1-Bit Quantized JL Transform for KV Cache Quantization with Zero Overhead
by: Zandieh, Amir, et al.
Published: (2024)
by: Zandieh, Amir, et al.
Published: (2024)
Recursive Introspection: Teaching Language Model Agents How to Self-Improve
by: Qu, Yuxiao, et al.
Published: (2024)
by: Qu, Yuxiao, et al.
Published: (2024)
Latency and Token-Aware Test-Time Compute
by: Huang, Jenny Y., et al.
Published: (2025)
by: Huang, Jenny Y., et al.
Published: (2025)
Log-Augmented Generation: Scaling Test-Time Reasoning with Reusable Computation
by: Chen, Peter Baile, et al.
Published: (2025)
by: Chen, Peter Baile, et al.
Published: (2025)
Provable Scaling Laws for the Test-Time Compute of Large Language Models
by: Chen, Yanxi, et al.
Published: (2024)
by: Chen, Yanxi, et al.
Published: (2024)
Optimizing Test-Time Compute via Meta Reinforcement Fine-Tuning
by: Qu, Yuxiao, et al.
Published: (2025)
by: Qu, Yuxiao, et al.
Published: (2025)
Strategic Scaling of Test-Time Compute: A Bandit Learning Approach
by: Zuo, Bowen, et al.
Published: (2025)
by: Zuo, Bowen, et al.
Published: (2025)
Adaptive Test-Time Reasoning via Reward-Guided Dual-Phase Search
by: Cui, Yingqian, et al.
Published: (2025)
by: Cui, Yingqian, et al.
Published: (2025)
Unfamiliar Finetuning Examples Control How Language Models Hallucinate
by: Kang, Katie, et al.
Published: (2024)
by: Kang, Katie, et al.
Published: (2024)
ArCHer: Training Language Model Agents via Hierarchical Multi-Turn RL
by: Zhou, Yifei, et al.
Published: (2024)
by: Zhou, Yifei, et al.
Published: (2024)
RLAC: Reinforcement Learning with Adversarial Critic for Free-Form Generation Tasks
by: Wu, Mian, et al.
Published: (2025)
by: Wu, Mian, et al.
Published: (2025)
A Survey of Test-Time Compute: From Intuitive Inference to Deliberate Reasoning
by: Ji, Yixin, et al.
Published: (2025)
by: Ji, Yixin, et al.
Published: (2025)
Scaling Test-Time Compute to Achieve IOI Gold Medal with Open-Weight Models
by: Samadi, Mehrzad, et al.
Published: (2025)
by: Samadi, Mehrzad, et al.
Published: (2025)
Reflect then Learn: Active Prompting for Information Extraction Guided by Introspective Confusion
by: Zhao, Dong, et al.
Published: (2025)
by: Zhao, Dong, et al.
Published: (2025)
Scaling Test-Time Compute Without Verification or RL is Suboptimal
by: Setlur, Amrith, et al.
Published: (2025)
by: Setlur, Amrith, et al.
Published: (2025)
Evaluating & Reducing Deceptive Dialogue From Language Models with Multi-turn RL
by: Abdulhai, Marwa, et al.
Published: (2025)
by: Abdulhai, Marwa, et al.
Published: (2025)
DynaPrompt: Dynamic Test-Time Prompt Tuning
by: Xiao, Zehao, et al.
Published: (2025)
by: Xiao, Zehao, et al.
Published: (2025)
Scaling Test-Time Compute for Agentic Coding
by: Kim, Joongwon, et al.
Published: (2026)
by: Kim, Joongwon, et al.
Published: (2026)
LocMoE: A Low-Overhead MoE for Large Language Model Training
by: Li, Jing, et al.
Published: (2024)
by: Li, Jing, et al.
Published: (2024)
Towards a Zero-Data, Controllable, Adaptive Dialog System
by: Väth, Dirk, et al.
Published: (2024)
by: Väth, Dirk, et al.
Published: (2024)
Growing Q-Networks: Solving Continuous Control Tasks with Adaptive Control Resolution
by: Seyde, Tim, et al.
Published: (2024)
by: Seyde, Tim, et al.
Published: (2024)
Evaluating the Goal-Directedness of Large Language Models
by: Everitt, Tom, et al.
Published: (2025)
by: Everitt, Tom, et al.
Published: (2025)
Building Production-Ready Probes For Gemini
by: Kramár, János, et al.
Published: (2026)
by: Kramár, János, et al.
Published: (2026)
RIV: Recursive Introspection Mask Diffusion Vision Language Model
by: Li, YuQian, et al.
Published: (2025)
by: Li, YuQian, et al.
Published: (2025)
Mechanistic Interpretability as Statistical Estimation: A Variance Analysis
by: Méloux, Maxime, et al.
Published: (2025)
by: Méloux, Maxime, et al.
Published: (2025)
Adaptive Computation Pruning for the Forgetting Transformer
by: Lin, Zhixuan, et al.
Published: (2025)
by: Lin, Zhixuan, et al.
Published: (2025)
The Right Time Matters: Data Arrangement Affects Zero-Shot Generalization in Instruction Tuning
by: He, Bingxiang, et al.
Published: (2024)
by: He, Bingxiang, et al.
Published: (2024)
In-Place Test-Time Training
by: Feng, Guhao, et al.
Published: (2026)
by: Feng, Guhao, et al.
Published: (2026)
Test-Time Safety Alignment
by: Saglam, Baturay, et al.
Published: (2026)
by: Saglam, Baturay, et al.
Published: (2026)
Similar Items
-
Adaptive Inference-Time Compute: LLMs Can Predict if They Can Do Better, Even Mid-Generation
by: Manvi, Rohin, et al.
Published: (2024) -
Q-SFT: Q-Learning for Language Models via Supervised Fine-Tuning
by: Hong, Joey, et al.
Published: (2024) -
Interactive Dialogue Agents via Reinforcement Learning on Hindsight Regenerations
by: Hong, Joey, et al.
Published: (2024) -
Large Language Models are Geographically Biased
by: Manvi, Rohin, et al.
Published: (2024) -
Merging in a Bottle: Differentiable Adaptive Merging (DAM) and the Path from Averaging to Automation
by: Gauthier-Caron, Thomas, et al.
Published: (2024)