LiteStage: Latency-aware Layer Skipping for Multi-stage Reasoning
Fuente:
arXiv
Saved in:
| Main Authors: | Kang, Beomseok, Song, Jiwon, Kim, Jae-Joon |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
CompactAttention: Accelerating Chunked Prefill with Block-Union KV Selection
by: Song, Jiwon, et al.
Published: (2026)
by: Song, Jiwon, et al.
Published: (2026)
Rotation-Aligned Key Channel Pruning for Efficient Vision-Language Model Inference
by: Kang, Beomseok, et al.
Published: (2026)
by: Kang, Beomseok, et al.
Published: (2026)
Token Sparse Attention: Efficient Long-Context Inference with Interleaved Token Selection
by: Jo, Dongwon, et al.
Published: (2026)
by: Jo, Dongwon, et al.
Published: (2026)
RelayGen: Intra-Generation Model Switching for Efficient Reasoning
by: Song, Jiwon, et al.
Published: (2026)
by: Song, Jiwon, et al.
Published: (2026)
Hop, Skip, and Overthink: Diagnosing Why Reasoning Models Fumble during Multi-Hop Analysis
by: Yadav, Anushka, et al.
Published: (2025)
by: Yadav, Anushka, et al.
Published: (2025)
What Layers When: Learning to Skip Compute in LLMs with Residual Gates
by: Laitenberger, Filipe, et al.
Published: (2025)
by: Laitenberger, Filipe, et al.
Published: (2025)
Reasoning Path Compression: Compressing Generation Trajectories for Efficient LLM Reasoning
by: Song, Jiwon, et al.
Published: (2025)
by: Song, Jiwon, et al.
Published: (2025)
MSCoRe: A Benchmark for Multi-Stage Collaborative Reasoning in LLM Agents
by: Lei, Yuzhen, et al.
Published: (2025)
by: Lei, Yuzhen, et al.
Published: (2025)
AdaSkip: Adaptive Sublayer Skipping for Accelerating Long-Context LLM Inference
by: He, Zhuomin, et al.
Published: (2025)
by: He, Zhuomin, et al.
Published: (2025)
Large Language Models are Clinical Reasoners: Reasoning-Aware Diagnosis Framework with Prompt-Generated Rationales
by: Kwon, Taeyoon, et al.
Published: (2023)
by: Kwon, Taeyoon, et al.
Published: (2023)
Skip-It? Theoretical Conditions for Layer Skipping in Vision-Language Models
by: Hartman, Max, et al.
Published: (2025)
by: Hartman, Max, et al.
Published: (2025)
ASKD-Whisper: Adaptive Self-knowledge Distillation for Efficient and Low-Latency Automatic Speech Recognition
by: Lee, Junseok, et al.
Published: (2026)
by: Lee, Junseok, et al.
Published: (2026)
Correct, Concise and Complete: Multi-stage Training For Adaptive Reasoning
by: Rakotonirina, Nathanaël Carraz, et al.
Published: (2026)
by: Rakotonirina, Nathanaël Carraz, et al.
Published: (2026)
What Defines Good Reasoning in LLMs? Dissecting Reasoning Steps with Multi-Aspect Evaluation
by: Do, Heejin, et al.
Published: (2025)
by: Do, Heejin, et al.
Published: (2025)
LayerSkip: Enabling Early Exit Inference and Self-Speculative Decoding
by: Elhoushi, Mostafa, et al.
Published: (2024)
by: Elhoushi, Mostafa, et al.
Published: (2024)
Adaptive GoGI-Skip: Coupling Goal-Gradient Importance with Dynamic Uncertainty for Efficient Reasoning
by: Zhuang, Ren
Published: (2025)
by: Zhuang, Ren
Published: (2025)
Sandwich Reasoning: An Answer-Reasoning-Answer Approach for Low-Latency Query Correction
by: Zhang, Chen, et al.
Published: (2026)
by: Zhang, Chen, et al.
Published: (2026)
Geometry-Lite: Interpretable Safety Probing via Layer-Wise Margin Geometry
by: Sim, Woo Seob, et al.
Published: (2026)
by: Sim, Woo Seob, et al.
Published: (2026)
Avoiding Knowledge Edit Skipping in Multi-hop Question Answering with Guided Decomposition
by: Liu, Yi, et al.
Published: (2025)
by: Liu, Yi, et al.
Published: (2025)
SHAPE: Stage-aware Hierarchical Advantage via Potential Estimation for LLM Reasoning
by: Ai, Zhengyang, et al.
Published: (2026)
by: Ai, Zhengyang, et al.
Published: (2026)
MapCoder-Lite: Distilling Multi-Agent Coding into a Single Small LLM
by: Lee, Woongkyu, et al.
Published: (2025)
by: Lee, Woongkyu, et al.
Published: (2025)
SkipCat: Rank-Maximized Low-Rank Compression of Large Language Models via Shared Projection and Block Skipping
by: Lu, Yu-Chen, et al.
Published: (2025)
by: Lu, Yu-Chen, et al.
Published: (2025)
Federated Learning with Layer Skipping: Efficient Training of Large Language Models for Healthcare NLP
by: Zhang, Lihong, et al.
Published: (2025)
by: Zhang, Lihong, et al.
Published: (2025)
HUMORCHAIN: Theory-Guided Multi-Stage Reasoning for Interpretable Multimodal Humor Generation
by: Zhang, Jiajun, et al.
Published: (2025)
by: Zhang, Jiajun, et al.
Published: (2025)
LiteSearch: Efficacious Tree Search for LLM
by: Wang, Ante, et al.
Published: (2024)
by: Wang, Ante, et al.
Published: (2024)
OPSD Compresses What RLVR Teaches: A Post-RL Compaction Stage for Reasoning Models
by: Kim, Jaehoon, et al.
Published: (2026)
by: Kim, Jaehoon, et al.
Published: (2026)
Multi-stage Prompt Refinement for Mitigating Hallucinations in Large Language Models
by: Shim, Jung-Woo, et al.
Published: (2025)
by: Shim, Jung-Woo, et al.
Published: (2025)
A Multi-Stage Workflow for the Review of Marketing Content with Reasoning Large Language Models
by: Purpura, Alberto, et al.
Published: (2025)
by: Purpura, Alberto, et al.
Published: (2025)
TokenSkip: Controllable Chain-of-Thought Compression in LLMs
by: Xia, Heming, et al.
Published: (2025)
by: Xia, Heming, et al.
Published: (2025)
Hierarchical Skip Decoding for Efficient Autoregressive Text Generation
by: Zhu, Yunqi, et al.
Published: (2024)
by: Zhu, Yunqi, et al.
Published: (2024)
QWHA: Quantization-Aware Walsh-Hadamard Adaptation for Parameter-Efficient Fine-Tuning on Large Language Models
by: Jeon, Hyesung, et al.
Published: (2025)
by: Jeon, Hyesung, et al.
Published: (2025)
Arena-Lite: Efficient and Reliable Large Language Model Evaluation via Tournament-Based Direct Comparisons
by: Son, Seonil, et al.
Published: (2024)
by: Son, Seonil, et al.
Published: (2024)
Aligning Reasoning LLMs for Materials Discovery with Physics-aware Rejection Sampling
by: Hyun, Lee, et al.
Published: (2025)
by: Hyun, Lee, et al.
Published: (2025)
Bridging Symbolic Control and Neural Reasoning in LLM Agents: Structured Cognitive Loop with a Governance Layer
by: Kim, Myung Ho
Published: (2025)
by: Kim, Myung Ho
Published: (2025)
Learning When to Translate for Multilingual Reasoning
by: Kang, Deokhyung, et al.
Published: (2026)
by: Kang, Deokhyung, et al.
Published: (2026)
Large Language Models Are Better Logical Fallacy Reasoners with Counterargument, Explanation, and Goal-Aware Prompt Formulation
by: Jeong, Jiwon, et al.
Published: (2025)
by: Jeong, Jiwon, et al.
Published: (2025)
LiteLong: Resource-Efficient Long-Context Data Synthesis for LLMs
by: Jia, Junlong, et al.
Published: (2025)
by: Jia, Junlong, et al.
Published: (2025)
Context-aware Inductive Knowledge Graph Completion with Latent Type Constraints and Subgraph Reasoning
by: Li, Muzhi, et al.
Published: (2024)
by: Li, Muzhi, et al.
Published: (2024)
StoryCoder: Narrative Reformulation for Structured Reasoning in LLM Code Generation
by: Jang, Geonhui, et al.
Published: (2026)
by: Jang, Geonhui, et al.
Published: (2026)
A2R: An Asymmetric Two-Stage Reasoning Framework for Parallel Reasoning
by: Wang, Ziqi, et al.
Published: (2025)
by: Wang, Ziqi, et al.
Published: (2025)
Similar Items
-
CompactAttention: Accelerating Chunked Prefill with Block-Union KV Selection
by: Song, Jiwon, et al.
Published: (2026) -
Rotation-Aligned Key Channel Pruning for Efficient Vision-Language Model Inference
by: Kang, Beomseok, et al.
Published: (2026) -
Token Sparse Attention: Efficient Long-Context Inference with Interleaved Token Selection
by: Jo, Dongwon, et al.
Published: (2026) -
RelayGen: Intra-Generation Model Switching for Efficient Reasoning
by: Song, Jiwon, et al.
Published: (2026) -
Hop, Skip, and Overthink: Diagnosing Why Reasoning Models Fumble during Multi-Hop Analysis
by: Yadav, Anushka, et al.
Published: (2025)