Hidden States as Early Signals: Step-level Trace Evaluation and Pruning for Efficient Test-Time Scaling
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Liang, Zhixiang, Huang, Beichen, Wang, Zheng, Zhang, Minjia |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
MiLo: Efficient Quantized MoE Inference with Mixture of Low-Rank Compensators
von: Huang, Beichen, et al.
Veröffentlicht: (2025)
von: Huang, Beichen, et al.
Veröffentlicht: (2025)
PuzzleMoE: Efficient Compression of Large Mixture-of-Experts Models via Sparse Expert Merging and Bit-packed inference
von: Zhao, Yushu, et al.
Veröffentlicht: (2025)
von: Zhao, Yushu, et al.
Veröffentlicht: (2025)
Revisiting the Effectiveness of LLM Pruning for Test-Time Scaling
von: Monjur, Ocean, et al.
Veröffentlicht: (2026)
von: Monjur, Ocean, et al.
Veröffentlicht: (2026)
Inference-Time Chain-of-Thought Pruning with Latent Informativeness Signals
von: Li, Sophie, et al.
Veröffentlicht: (2025)
von: Li, Sophie, et al.
Veröffentlicht: (2025)
Learning to (Learn at Test Time): RNNs with Expressive Hidden States
von: Sun, Yu, et al.
Veröffentlicht: (2024)
von: Sun, Yu, et al.
Veröffentlicht: (2024)
Downstream-Pretext Domain Knowledge Traceback for Active Learning
von: Zhang, Beichen, et al.
Veröffentlicht: (2024)
von: Zhang, Beichen, et al.
Veröffentlicht: (2024)
Efficient Test-Time Scaling via Self-Calibration
von: Huang, Chengsong, et al.
Veröffentlicht: (2025)
von: Huang, Chengsong, et al.
Veröffentlicht: (2025)
Towards Efficient Automatic Self-Pruning of Large Language Models
von: Huang, Weizhong, et al.
Veröffentlicht: (2025)
von: Huang, Weizhong, et al.
Veröffentlicht: (2025)
Test-Time Scaling in Diffusion LLMs via Hidden Semi-Autoregressive Experts
von: Lee, Jihoon, et al.
Veröffentlicht: (2025)
von: Lee, Jihoon, et al.
Veröffentlicht: (2025)
Critique to Verify: Accurate and Honest Test-Time Scaling with RL-Trained Verifiers
von: Yang, Zhicheng, et al.
Veröffentlicht: (2025)
von: Yang, Zhicheng, et al.
Veröffentlicht: (2025)
When Fewer Layers Break More Chains: Layer Pruning Harms Test-Time Scaling in LLMs
von: Wang, Keyu, et al.
Veröffentlicht: (2025)
von: Wang, Keyu, et al.
Veröffentlicht: (2025)
Evaluating Judges as Evaluators: The JETTS Benchmark of LLM-as-Judges as Test-Time Scaling Evaluators
von: Zhou, Yilun, et al.
Veröffentlicht: (2025)
von: Zhou, Yilun, et al.
Veröffentlicht: (2025)
VoltanaLLM: Feedback-Driven Frequency Control and State-Space Routing for Energy-Efficient LLM Serving
von: Yu, Jiahuan, et al.
Veröffentlicht: (2025)
von: Yu, Jiahuan, et al.
Veröffentlicht: (2025)
FLOP-Efficient Training: Early Stopping Based on Test-Time Compute Awareness
von: Amer, Hossam, et al.
Veröffentlicht: (2026)
von: Amer, Hossam, et al.
Veröffentlicht: (2026)
EfficientLLM: Scalable Pruning-Aware Pretraining for Architecture-Agnostic Edge Language Models
von: Xing, Xingrun, et al.
Veröffentlicht: (2025)
von: Xing, Xingrun, et al.
Veröffentlicht: (2025)
Cut Your Losses! Learning to Prune Paths Early for Efficient Parallel Reasoning
von: Bi, Jiaxi, et al.
Veröffentlicht: (2026)
von: Bi, Jiaxi, et al.
Veröffentlicht: (2026)
State Tuning: State-based Test-Time Scaling on RWKV-7
von: Xiao, Liu, et al.
Veröffentlicht: (2025)
von: Xiao, Liu, et al.
Veröffentlicht: (2025)
Beyond What Seems Necessary: Hidden Gains from Scaling Training-Time Reasoning Length under Outcome Supervision
von: Xue, Yihao, et al.
Veröffentlicht: (2026)
von: Xue, Yihao, et al.
Veröffentlicht: (2026)
Plan and Budget: Effective and Efficient Test-Time Scaling on Reasoning Large Language Models
von: Lin, Junhong, et al.
Veröffentlicht: (2025)
von: Lin, Junhong, et al.
Veröffentlicht: (2025)
Early Detection of Misinformation for Infodemic Management: A Domain Adaptation Approach
von: Mao, Minjia, et al.
Veröffentlicht: (2024)
von: Mao, Minjia, et al.
Veröffentlicht: (2024)
Every Rollout Counts: Optimal Resource Allocation for Efficient Test-Time Scaling
von: Wang, Xinglin, et al.
Veröffentlicht: (2025)
von: Wang, Xinglin, et al.
Veröffentlicht: (2025)
PrunePEFT: Iterative Hybrid Pruning for Parameter-Efficient Fine-tuning of LLMs
von: Yu, Tongzhou, et al.
Veröffentlicht: (2025)
von: Yu, Tongzhou, et al.
Veröffentlicht: (2025)
Do Not Waste Your Rollouts: Recycling Search Experience for Efficient Test-Time Scaling
von: Wang, Xinglin, et al.
Veröffentlicht: (2026)
von: Wang, Xinglin, et al.
Veröffentlicht: (2026)
SparseSSM: Efficient Selective Structured State Space Models Can Be Pruned in One-Shot
von: Tuo, Kaiwen, et al.
Veröffentlicht: (2025)
von: Tuo, Kaiwen, et al.
Veröffentlicht: (2025)
S2O: Early Stopping for Sparse Attention via Online Permutation
von: Zhang, Yu, et al.
Veröffentlicht: (2026)
von: Zhang, Yu, et al.
Veröffentlicht: (2026)
A Fine Evaluation Method for Cube Copying Test for Early Detection of Alzheimer's Disease
von: Jiang, Xinyu, et al.
Veröffentlicht: (2025)
von: Jiang, Xinyu, et al.
Veröffentlicht: (2025)
Test-Time Scaling with Reflective Generative Model
von: Wang, Zixiao, et al.
Veröffentlicht: (2025)
von: Wang, Zixiao, et al.
Veröffentlicht: (2025)
SWE-Replay: Efficient Test-Time Scaling for Software Engineering Agents
von: Ding, Yifeng, et al.
Veröffentlicht: (2026)
von: Ding, Yifeng, et al.
Veröffentlicht: (2026)
Beyond the Frontier: Stochastic Backtracking for Efficient Test-Time Scaling
von: Tran, Dao, et al.
Veröffentlicht: (2026)
von: Tran, Dao, et al.
Veröffentlicht: (2026)
MetaScale: Test-Time Scaling with Evolving Meta-Thoughts
von: Liu, Qin, et al.
Veröffentlicht: (2025)
von: Liu, Qin, et al.
Veröffentlicht: (2025)
SuperOffload: Unleashing the Power of Large-Scale LLM Training on Superchips
von: Lian, Xinyu, et al.
Veröffentlicht: (2025)
von: Lian, Xinyu, et al.
Veröffentlicht: (2025)
FastForward Pruning: Efficient LLM Pruning via Single-Step Reinforcement Learning
von: Yuan, Xin, et al.
Veröffentlicht: (2025)
von: Yuan, Xin, et al.
Veröffentlicht: (2025)
Kinetics: Rethinking Test-Time Scaling Laws
von: Sadhukhan, Ranajoy, et al.
Veröffentlicht: (2025)
von: Sadhukhan, Ranajoy, et al.
Veröffentlicht: (2025)
CipherPrune: Efficient and Scalable Private Transformer Inference
von: Zhang, Yancheng, et al.
Veröffentlicht: (2025)
von: Zhang, Yancheng, et al.
Veröffentlicht: (2025)
PaCoRe: Learning to Scale Test-Time Compute with Parallel Coordinated Reasoning
von: Hu, Jingcheng, et al.
Veröffentlicht: (2026)
von: Hu, Jingcheng, et al.
Veröffentlicht: (2026)
TraceNAS: Zero-shot LLM Pruning via Gradient Trace Correlation
von: Malettira, Prajna G., et al.
Veröffentlicht: (2026)
von: Malettira, Prajna G., et al.
Veröffentlicht: (2026)
EPSD: Early Pruning with Self-Distillation for Efficient Model Compression
von: Chen, Dong, et al.
Veröffentlicht: (2024)
von: Chen, Dong, et al.
Veröffentlicht: (2024)
DART-ing Through the Drift: Dynamic Tracing of Knowledge Neurons for Adaptive Inference-Time Pruning
von: Tyagi, Abhishek, et al.
Veröffentlicht: (2026)
von: Tyagi, Abhishek, et al.
Veröffentlicht: (2026)
Pyramidal Hidden Markov Model For Multivariate Time Series Forecasting
von: Huang, YeXin
Veröffentlicht: (2023)
von: Huang, YeXin
Veröffentlicht: (2023)
SalienTime: User-driven Selection of Salient Time Steps for Large-Scale Geospatial Data Visualization
von: Chen, Juntong, et al.
Veröffentlicht: (2024)
von: Chen, Juntong, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
MiLo: Efficient Quantized MoE Inference with Mixture of Low-Rank Compensators
von: Huang, Beichen, et al.
Veröffentlicht: (2025) -
PuzzleMoE: Efficient Compression of Large Mixture-of-Experts Models via Sparse Expert Merging and Bit-packed inference
von: Zhao, Yushu, et al.
Veröffentlicht: (2025) -
Revisiting the Effectiveness of LLM Pruning for Test-Time Scaling
von: Monjur, Ocean, et al.
Veröffentlicht: (2026) -
Inference-Time Chain-of-Thought Pruning with Latent Informativeness Signals
von: Li, Sophie, et al.
Veröffentlicht: (2025) -
Learning to (Learn at Test Time): RNNs with Expressive Hidden States
von: Sun, Yu, et al.
Veröffentlicht: (2024)