Hidden States as Early Signals: Step-level Trace Evaluation and Pruning for Efficient Test-Time Scaling
Fuente:
arXiv
Saved in:
| Main Authors: | Liang, Zhixiang, Huang, Beichen, Wang, Zheng, Zhang, Minjia |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
MiLo: Efficient Quantized MoE Inference with Mixture of Low-Rank Compensators
by: Huang, Beichen, et al.
Published: (2025)
by: Huang, Beichen, et al.
Published: (2025)
PuzzleMoE: Efficient Compression of Large Mixture-of-Experts Models via Sparse Expert Merging and Bit-packed inference
by: Zhao, Yushu, et al.
Published: (2025)
by: Zhao, Yushu, et al.
Published: (2025)
Revisiting the Effectiveness of LLM Pruning for Test-Time Scaling
by: Monjur, Ocean, et al.
Published: (2026)
by: Monjur, Ocean, et al.
Published: (2026)
Inference-Time Chain-of-Thought Pruning with Latent Informativeness Signals
by: Li, Sophie, et al.
Published: (2025)
by: Li, Sophie, et al.
Published: (2025)
Learning to (Learn at Test Time): RNNs with Expressive Hidden States
by: Sun, Yu, et al.
Published: (2024)
by: Sun, Yu, et al.
Published: (2024)
Downstream-Pretext Domain Knowledge Traceback for Active Learning
by: Zhang, Beichen, et al.
Published: (2024)
by: Zhang, Beichen, et al.
Published: (2024)
Efficient Test-Time Scaling via Self-Calibration
by: Huang, Chengsong, et al.
Published: (2025)
by: Huang, Chengsong, et al.
Published: (2025)
Towards Efficient Automatic Self-Pruning of Large Language Models
by: Huang, Weizhong, et al.
Published: (2025)
by: Huang, Weizhong, et al.
Published: (2025)
Test-Time Scaling in Diffusion LLMs via Hidden Semi-Autoregressive Experts
by: Lee, Jihoon, et al.
Published: (2025)
by: Lee, Jihoon, et al.
Published: (2025)
Critique to Verify: Accurate and Honest Test-Time Scaling with RL-Trained Verifiers
by: Yang, Zhicheng, et al.
Published: (2025)
by: Yang, Zhicheng, et al.
Published: (2025)
When Fewer Layers Break More Chains: Layer Pruning Harms Test-Time Scaling in LLMs
by: Wang, Keyu, et al.
Published: (2025)
by: Wang, Keyu, et al.
Published: (2025)
Evaluating Judges as Evaluators: The JETTS Benchmark of LLM-as-Judges as Test-Time Scaling Evaluators
by: Zhou, Yilun, et al.
Published: (2025)
by: Zhou, Yilun, et al.
Published: (2025)
VoltanaLLM: Feedback-Driven Frequency Control and State-Space Routing for Energy-Efficient LLM Serving
by: Yu, Jiahuan, et al.
Published: (2025)
by: Yu, Jiahuan, et al.
Published: (2025)
FLOP-Efficient Training: Early Stopping Based on Test-Time Compute Awareness
by: Amer, Hossam, et al.
Published: (2026)
by: Amer, Hossam, et al.
Published: (2026)
EfficientLLM: Scalable Pruning-Aware Pretraining for Architecture-Agnostic Edge Language Models
by: Xing, Xingrun, et al.
Published: (2025)
by: Xing, Xingrun, et al.
Published: (2025)
Cut Your Losses! Learning to Prune Paths Early for Efficient Parallel Reasoning
by: Bi, Jiaxi, et al.
Published: (2026)
by: Bi, Jiaxi, et al.
Published: (2026)
State Tuning: State-based Test-Time Scaling on RWKV-7
by: Xiao, Liu, et al.
Published: (2025)
by: Xiao, Liu, et al.
Published: (2025)
Beyond What Seems Necessary: Hidden Gains from Scaling Training-Time Reasoning Length under Outcome Supervision
by: Xue, Yihao, et al.
Published: (2026)
by: Xue, Yihao, et al.
Published: (2026)
Plan and Budget: Effective and Efficient Test-Time Scaling on Reasoning Large Language Models
by: Lin, Junhong, et al.
Published: (2025)
by: Lin, Junhong, et al.
Published: (2025)
Early Detection of Misinformation for Infodemic Management: A Domain Adaptation Approach
by: Mao, Minjia, et al.
Published: (2024)
by: Mao, Minjia, et al.
Published: (2024)
Every Rollout Counts: Optimal Resource Allocation for Efficient Test-Time Scaling
by: Wang, Xinglin, et al.
Published: (2025)
by: Wang, Xinglin, et al.
Published: (2025)
PrunePEFT: Iterative Hybrid Pruning for Parameter-Efficient Fine-tuning of LLMs
by: Yu, Tongzhou, et al.
Published: (2025)
by: Yu, Tongzhou, et al.
Published: (2025)
Do Not Waste Your Rollouts: Recycling Search Experience for Efficient Test-Time Scaling
by: Wang, Xinglin, et al.
Published: (2026)
by: Wang, Xinglin, et al.
Published: (2026)
SparseSSM: Efficient Selective Structured State Space Models Can Be Pruned in One-Shot
by: Tuo, Kaiwen, et al.
Published: (2025)
by: Tuo, Kaiwen, et al.
Published: (2025)
S2O: Early Stopping for Sparse Attention via Online Permutation
by: Zhang, Yu, et al.
Published: (2026)
by: Zhang, Yu, et al.
Published: (2026)
A Fine Evaluation Method for Cube Copying Test for Early Detection of Alzheimer's Disease
by: Jiang, Xinyu, et al.
Published: (2025)
by: Jiang, Xinyu, et al.
Published: (2025)
Test-Time Scaling with Reflective Generative Model
by: Wang, Zixiao, et al.
Published: (2025)
by: Wang, Zixiao, et al.
Published: (2025)
SWE-Replay: Efficient Test-Time Scaling for Software Engineering Agents
by: Ding, Yifeng, et al.
Published: (2026)
by: Ding, Yifeng, et al.
Published: (2026)
Beyond the Frontier: Stochastic Backtracking for Efficient Test-Time Scaling
by: Tran, Dao, et al.
Published: (2026)
by: Tran, Dao, et al.
Published: (2026)
MetaScale: Test-Time Scaling with Evolving Meta-Thoughts
by: Liu, Qin, et al.
Published: (2025)
by: Liu, Qin, et al.
Published: (2025)
SuperOffload: Unleashing the Power of Large-Scale LLM Training on Superchips
by: Lian, Xinyu, et al.
Published: (2025)
by: Lian, Xinyu, et al.
Published: (2025)
FastForward Pruning: Efficient LLM Pruning via Single-Step Reinforcement Learning
by: Yuan, Xin, et al.
Published: (2025)
by: Yuan, Xin, et al.
Published: (2025)
Kinetics: Rethinking Test-Time Scaling Laws
by: Sadhukhan, Ranajoy, et al.
Published: (2025)
by: Sadhukhan, Ranajoy, et al.
Published: (2025)
CipherPrune: Efficient and Scalable Private Transformer Inference
by: Zhang, Yancheng, et al.
Published: (2025)
by: Zhang, Yancheng, et al.
Published: (2025)
PaCoRe: Learning to Scale Test-Time Compute with Parallel Coordinated Reasoning
by: Hu, Jingcheng, et al.
Published: (2026)
by: Hu, Jingcheng, et al.
Published: (2026)
TraceNAS: Zero-shot LLM Pruning via Gradient Trace Correlation
by: Malettira, Prajna G., et al.
Published: (2026)
by: Malettira, Prajna G., et al.
Published: (2026)
EPSD: Early Pruning with Self-Distillation for Efficient Model Compression
by: Chen, Dong, et al.
Published: (2024)
by: Chen, Dong, et al.
Published: (2024)
DART-ing Through the Drift: Dynamic Tracing of Knowledge Neurons for Adaptive Inference-Time Pruning
by: Tyagi, Abhishek, et al.
Published: (2026)
by: Tyagi, Abhishek, et al.
Published: (2026)
Pyramidal Hidden Markov Model For Multivariate Time Series Forecasting
by: Huang, YeXin
Published: (2023)
by: Huang, YeXin
Published: (2023)
SalienTime: User-driven Selection of Salient Time Steps for Large-Scale Geospatial Data Visualization
by: Chen, Juntong, et al.
Published: (2024)
by: Chen, Juntong, et al.
Published: (2024)
Similar Items
-
MiLo: Efficient Quantized MoE Inference with Mixture of Low-Rank Compensators
by: Huang, Beichen, et al.
Published: (2025) -
PuzzleMoE: Efficient Compression of Large Mixture-of-Experts Models via Sparse Expert Merging and Bit-packed inference
by: Zhao, Yushu, et al.
Published: (2025) -
Revisiting the Effectiveness of LLM Pruning for Test-Time Scaling
by: Monjur, Ocean, et al.
Published: (2026) -
Inference-Time Chain-of-Thought Pruning with Latent Informativeness Signals
by: Li, Sophie, et al.
Published: (2025) -
Learning to (Learn at Test Time): RNNs with Expressive Hidden States
by: Sun, Yu, et al.
Published: (2024)