RelayGen: Intra-Generation Model Switching for Efficient Reasoning
Fuente:
arXiv
Saved in:
| Main Authors: | Song, Jiwon, Kim, Yoongon, Kim, Jae-Joon |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Reasoning Path Compression: Compressing Generation Trajectories for Efficient LLM Reasoning
by: Song, Jiwon, et al.
Published: (2025)
by: Song, Jiwon, et al.
Published: (2025)
LiteStage: Latency-aware Layer Skipping for Multi-stage Reasoning
by: Kang, Beomseok, et al.
Published: (2025)
by: Kang, Beomseok, et al.
Published: (2025)
Token Sparse Attention: Efficient Long-Context Inference with Interleaved Token Selection
by: Jo, Dongwon, et al.
Published: (2026)
by: Jo, Dongwon, et al.
Published: (2026)
FastKV: Decoupling of Context Reduction and KV Cache Compression for Prefill-Decoding Acceleration
by: Jo, Dongwon, et al.
Published: (2025)
by: Jo, Dongwon, et al.
Published: (2025)
CompactAttention: Accelerating Chunked Prefill with Block-Union KV Selection
by: Song, Jiwon, et al.
Published: (2026)
by: Song, Jiwon, et al.
Published: (2026)
SLEB: Streamlining LLMs through Redundancy Verification and Elimination of Transformer Blocks
by: Song, Jiwon, et al.
Published: (2024)
by: Song, Jiwon, et al.
Published: (2024)
QWHA: Quantization-Aware Walsh-Hadamard Adaptation for Parameter-Efficient Fine-Tuning on Large Language Models
by: Jeon, Hyesung, et al.
Published: (2025)
by: Jeon, Hyesung, et al.
Published: (2025)
MedRep: Medical Concept Representation for General Electronic Health Record Foundation Models
by: Kim, Junmo, et al.
Published: (2025)
by: Kim, Junmo, et al.
Published: (2025)
Rotation-Aligned Key Channel Pruning for Efficient Vision-Language Model Inference
by: Kang, Beomseok, et al.
Published: (2026)
by: Kang, Beomseok, et al.
Published: (2026)
Rebellious Student: Reversing Teacher Signals for Reasoning Exploration with Self-Distilled RLVR
by: Kim, Jeonghye, et al.
Published: (2026)
by: Kim, Jeonghye, et al.
Published: (2026)
L4Q: Parameter Efficient Quantization-Aware Fine-Tuning on Large Language Models
by: Jeon, Hyesung, et al.
Published: (2024)
by: Jeon, Hyesung, et al.
Published: (2024)
RelayLLM: Efficient Reasoning via Collaborative Decoding
by: Huang, Chengsong, et al.
Published: (2026)
by: Huang, Chengsong, et al.
Published: (2026)
Exploration of COVID-19 Discourse on Twitter: American Politician Edition
by: Kim, Cindy, et al.
Published: (2025)
by: Kim, Cindy, et al.
Published: (2025)
Context-Aware LLM Translation System Using Conversation Summarization and Dialogue History
by: Sung, Mingi, et al.
Published: (2024)
by: Sung, Mingi, et al.
Published: (2024)
Reasoning Abilities of Large Language Models: In-Depth Analysis on the Abstraction and Reasoning Corpus
by: Lee, Seungpil, et al.
Published: (2024)
by: Lee, Seungpil, et al.
Published: (2024)
A Dual-Layered Evaluation of Geopolitical and Cultural Bias in LLMs
by: Kim, Sean, et al.
Published: (2025)
by: Kim, Sean, et al.
Published: (2025)
Why Does Self-Distillation (Sometimes) Degrade the Reasoning Capability of LLMs?
by: Kim, Jeonghye, et al.
Published: (2026)
by: Kim, Jeonghye, et al.
Published: (2026)
Can Code-Switched Texts Activate a Knowledge Switch in LLMs? A Case Study on English-Korean Code-Switching
by: Kim, Seoyeon, et al.
Published: (2024)
by: Kim, Seoyeon, et al.
Published: (2024)
DocHop-QA: Towards Multi-Hop Reasoning over Multimodal Document Collections
by: Park, Jiwon, et al.
Published: (2025)
by: Park, Jiwon, et al.
Published: (2025)
AVCap: Leveraging Audio-Visual Features as Text Tokens for Captioning
by: Kim, Jongsuk, et al.
Published: (2024)
by: Kim, Jongsuk, et al.
Published: (2024)
UniGen: Universal Domain Generalization for Sentiment Classification via Zero-shot Dataset Generation
by: Choi, Juhwan, et al.
Published: (2024)
by: Choi, Juhwan, et al.
Published: (2024)
RelayAttention for Efficient Large Language Model Serving with Long System Prompts
by: Zhu, Lei, et al.
Published: (2024)
by: Zhu, Lei, et al.
Published: (2024)
Small Language Models are Equation Reasoners
by: Kim, Bumjun, et al.
Published: (2024)
by: Kim, Bumjun, et al.
Published: (2024)
Confidence-Guided Stepwise Model Routing for Cost-Efficient Reasoning
by: Lee, Sangmook, et al.
Published: (2025)
by: Lee, Sangmook, et al.
Published: (2025)
Guiding Reasoning in Small Language Models with LLM Assistance
by: Kim, Yujin, et al.
Published: (2025)
by: Kim, Yujin, et al.
Published: (2025)
Think Multilingual, Not Harder: A Data-Efficient Framework for Teaching Reasoning Models to Code-Switch
by: Lin, Eleanor M., et al.
Published: (2026)
by: Lin, Eleanor M., et al.
Published: (2026)
When to Speak, When to Abstain: Contrastive Decoding with Abstention
by: Kim, Hyuhng Joon, et al.
Published: (2024)
by: Kim, Hyuhng Joon, et al.
Published: (2024)
RPM: Reasoning-Level Personalization for Black-Box Large Language Models
by: Kim, Jieyong, et al.
Published: (2025)
by: Kim, Jieyong, et al.
Published: (2025)
Unplug and Play Language Models: Decomposing Experts in Language Models at Inference Time
by: Yang, Nakyeong, et al.
Published: (2024)
by: Yang, Nakyeong, et al.
Published: (2024)
GenPRM: Scaling Test-Time Compute of Process Reward Models via Generative Reasoning
by: Zhao, Jian, et al.
Published: (2025)
by: Zhao, Jian, et al.
Published: (2025)
Reasoning Models Better Express Their Confidence
by: Yoon, Dongkeun, et al.
Published: (2025)
by: Yoon, Dongkeun, et al.
Published: (2025)
Speculative Verification: Exploiting Information Gain to Refine Speculative Decoding
by: Kim, Sungkyun, et al.
Published: (2025)
by: Kim, Sungkyun, et al.
Published: (2025)
A Two-Step Approach for Data-Efficient French Pronunciation Learning
by: Lee, Hoyeon, et al.
Published: (2024)
by: Lee, Hoyeon, et al.
Published: (2024)
Aligning Language Models to Explicitly Handle Ambiguity
by: Kim, Hyuhng Joon, et al.
Published: (2024)
by: Kim, Hyuhng Joon, et al.
Published: (2024)
COMPASS: A Compiler Framework for Resource-Constrained Crossbar-Array Based In-Memory Deep Learning Accelerators
by: Park, Jihoon, et al.
Published: (2025)
by: Park, Jihoon, et al.
Published: (2025)
Investigating the Influence of Prompt-Specific Shortcuts in AI Generated Text Detection
by: Park, Choonghyun, et al.
Published: (2024)
by: Park, Choonghyun, et al.
Published: (2024)
Adaptive Contrastive Decoding in Retrieval-Augmented Generation for Handling Noisy Contexts
by: Kim, Youna, et al.
Published: (2024)
by: Kim, Youna, et al.
Published: (2024)
Reasoning Models Generate Societies of Thought
by: Kim, Junsol, et al.
Published: (2026)
by: Kim, Junsol, et al.
Published: (2026)
FCMR: Robust Evaluation of Financial Cross-Modal Multi-Hop Reasoning
by: Kim, Seunghee, et al.
Published: (2024)
by: Kim, Seunghee, et al.
Published: (2024)
UniKnow: A Unified Framework for Reliable Language Model Behavior across Parametric and External Knowledge
by: Kim, Youna, et al.
Published: (2025)
by: Kim, Youna, et al.
Published: (2025)
Similar Items
-
Reasoning Path Compression: Compressing Generation Trajectories for Efficient LLM Reasoning
by: Song, Jiwon, et al.
Published: (2025) -
LiteStage: Latency-aware Layer Skipping for Multi-stage Reasoning
by: Kang, Beomseok, et al.
Published: (2025) -
Token Sparse Attention: Efficient Long-Context Inference with Interleaved Token Selection
by: Jo, Dongwon, et al.
Published: (2026) -
FastKV: Decoupling of Context Reduction and KV Cache Compression for Prefill-Decoding Acceleration
by: Jo, Dongwon, et al.
Published: (2025) -
CompactAttention: Accelerating Chunked Prefill with Block-Union KV Selection
by: Song, Jiwon, et al.
Published: (2026)