Router-Tuning: A Simple and Effective Approach for Enabling Dynamic-Depth in Transformers
Fuente:
arXiv
Saved in:
| Main Authors: | He, Shwai, Ge, Tao, Sun, Guoheng, Tian, Bowei, Wang, Xiaoyang, Yu, Dong |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
What Matters in Transformers? Not All Attention is Needed
by: He, Shwai, et al.
Published: (2024)
by: He, Shwai, et al.
Published: (2024)
Demystifying When Pruning Works via Representation Hierarchies
by: He, Shwai, et al.
Published: (2026)
by: He, Shwai, et al.
Published: (2026)
Selective Reflection-Tuning: Student-Selected Data Recycling for LLM Instruction-Tuning
by: Li, Ming, et al.
Published: (2024)
by: Li, Ming, et al.
Published: (2024)
OpenCharacter: Training Customizable Role-Playing LLMs with Large-Scale Synthetic Personas
by: Wang, Xiaoyang, et al.
Published: (2025)
by: Wang, Xiaoyang, et al.
Published: (2025)
Superfiltering: Weak-to-Strong Data Filtering for Fast Instruction-Tuning
by: Li, Ming, et al.
Published: (2024)
by: Li, Ming, et al.
Published: (2024)
WebRouter: Query-specific Router via Variational Information Bottleneck for Cost-sensitive Web Agent
by: Li, Tao, et al.
Published: (2025)
by: Li, Tao, et al.
Published: (2025)
98$\times$ Faster LLM Routing Without a Dedicated GPU: Flash Attention, Prompt Compression, and Near-Streaming for the vLLM Semantic Router
by: Liu, Xunzhuo, et al.
Published: (2026)
by: Liu, Xunzhuo, et al.
Published: (2026)
Mixture of Routers
by: Zhang, Jia-Chen, et al.
Published: (2025)
by: Zhang, Jia-Chen, et al.
Published: (2025)
SHED: Shapley-Based Automated Dataset Refinement for Instruction Fine-Tuning
by: He, Yexiao, et al.
Published: (2024)
by: He, Yexiao, et al.
Published: (2024)
Router Upcycling: Leveraging Mixture-of-Routers in Mixture-of-Experts Upcycling
by: Ran, Junfeng, et al.
Published: (2025)
by: Ran, Junfeng, et al.
Published: (2025)
CP-Router: An Uncertainty-Aware Router Between LLM and LRM
by: Su, Jiayuan, et al.
Published: (2025)
by: Su, Jiayuan, et al.
Published: (2025)
DMON: A Simple yet Effective Approach for Argument Structure Learning
by: Sun, Wei, et al.
Published: (2024)
by: Sun, Wei, et al.
Published: (2024)
Scaling Synthetic Data Creation with 1,000,000,000 Personas
by: Ge, Tao, et al.
Published: (2024)
by: Ge, Tao, et al.
Published: (2024)
Inner Thinking Transformer: Leveraging Dynamic Depth Scaling to Foster Adaptive Internal Thinking
by: Chen, Yilong, et al.
Published: (2025)
by: Chen, Yilong, et al.
Published: (2025)
Making Large Language Models Efficient Dense Retrievers
by: Lei, Yibin, et al.
Published: (2025)
by: Lei, Yibin, et al.
Published: (2025)
Language Bias in LVLMs: From In-Depth Analysis to Simple and Effective Mitigation
by: Chen, Yangneng, et al.
Published: (2026)
by: Chen, Yangneng, et al.
Published: (2026)
GMTRouter: Personalized LLM Router over Multi-turn User Interactions
by: Xie, Encheng, et al.
Published: (2025)
by: Xie, Encheng, et al.
Published: (2025)
A Simple and Effective Pruning Approach for Large Language Models
by: Sun, Mingjie, et al.
Published: (2023)
by: Sun, Mingjie, et al.
Published: (2023)
OrcaRouter: A Production-Oriented LLM Router with Hybrid Offline-Online Learning
by: Bao, Zhenghua, et al.
Published: (2026)
by: Bao, Zhenghua, et al.
Published: (2026)
AgentRouter: A Knowledge-Graph-Guided LLM Router for Collaborative Multi-Agent Question Answering
by: Zhang, Zheyuan, et al.
Published: (2025)
by: Zhang, Zheyuan, et al.
Published: (2025)
Arctic-Text2SQL-R1: Simple Rewards, Strong Reasoning in Text-to-SQL
by: Yao, Zhewei, et al.
Published: (2025)
by: Yao, Zhewei, et al.
Published: (2025)
LoRA-Squeeze: Simple and Effective Post-Tuning and In-Tuning Compression of LoRA Modules
by: Vulić, Ivan, et al.
Published: (2026)
by: Vulić, Ivan, et al.
Published: (2026)
Fair Diagnosis: Leveraging Causal Modeling to Mitigate Medical Bias
by: Tian, Bowei, et al.
Published: (2024)
by: Tian, Bowei, et al.
Published: (2024)
FiRST: Finetuning Router-Selective Transformers for Input-Adaptive Latency Reduction
by: Jain, Akriti, et al.
Published: (2024)
by: Jain, Akriti, et al.
Published: (2024)
Mario at EXIST 2025: A Simple Gateway to Effective Multilingual Sexism Detection
by: Tian, Lin, et al.
Published: (2025)
by: Tian, Lin, et al.
Published: (2025)
Instruction Matters: A Simple yet Effective Task Selection for Optimized Instruction Tuning of Specific Tasks
by: Lee, Changho, et al.
Published: (2024)
by: Lee, Changho, et al.
Published: (2024)
SymRTLO: Enhancing RTL Code Optimization with LLMs and Neuron-Inspired Symbolic Reasoning
by: Wang, Yiting, et al.
Published: (2025)
by: Wang, Yiting, et al.
Published: (2025)
Capacity-Aware Inference: Mitigating the Straggler Effect in Mixture of Experts
by: He, Shwai, et al.
Published: (2025)
by: He, Shwai, et al.
Published: (2025)
Yuan 2.0-M32: Mixture of Experts with Attention Router
by: Wu, Shaohua, et al.
Published: (2024)
by: Wu, Shaohua, et al.
Published: (2024)
Layerwise Recurrent Router for Mixture-of-Experts
by: Qiu, Zihan, et al.
Published: (2024)
by: Qiu, Zihan, et al.
Published: (2024)
LVPruning: An Effective yet Simple Language-Guided Vision Token Pruning Approach for Multi-modal Large Language Models
by: Sun, Yizheng, et al.
Published: (2025)
by: Sun, Yizheng, et al.
Published: (2025)
The First Few Tokens Are All You Need: An Efficient and Effective Unsupervised Prefix Fine-Tuning Method for Reasoning Models
by: Ji, Ke, et al.
Published: (2025)
by: Ji, Ke, et al.
Published: (2025)
Towards counterfactual fairness through auxiliary variables
by: Tian, Bowei, et al.
Published: (2024)
by: Tian, Bowei, et al.
Published: (2024)
XFormParser: A Simple and Effective Multimodal Multilingual Semi-structured Form Parser
by: Cheng, Xianfu, et al.
Published: (2024)
by: Cheng, Xianfu, et al.
Published: (2024)
NaturalConv: A Chinese Dialogue Dataset Towards Multi-turn Topic-driven Conversation
by: Wang, Xiaoyang, et al.
Published: (2021)
by: Wang, Xiaoyang, et al.
Published: (2021)
Scaling Mobile Agent Systems: From Capability Density to Collective Intelligence
by: He, Bowei
Published: (2026)
by: He, Bowei
Published: (2026)
Simple and Effective Input Reformulations for Translation
by: Yu, Brian, et al.
Published: (2023)
by: Yu, Brian, et al.
Published: (2023)
A Simple Linear Patch Revives Layer-Pruned Large Language Models
by: Chen, Xinrui, et al.
Published: (2025)
by: Chen, Xinrui, et al.
Published: (2025)
Simple Yet Effective: Extracting Private Data Across Clients in Federated Fine-Tuning of Large Language Models
by: Hu, Yingqi, et al.
Published: (2025)
by: Hu, Yingqi, et al.
Published: (2025)
RouterDC: Query-Based Router by Dual Contrastive Learning for Assembling Large Language Models
by: Chen, Shuhao, et al.
Published: (2024)
by: Chen, Shuhao, et al.
Published: (2024)
Similar Items
-
What Matters in Transformers? Not All Attention is Needed
by: He, Shwai, et al.
Published: (2024) -
Demystifying When Pruning Works via Representation Hierarchies
by: He, Shwai, et al.
Published: (2026) -
Selective Reflection-Tuning: Student-Selected Data Recycling for LLM Instruction-Tuning
by: Li, Ming, et al.
Published: (2024) -
OpenCharacter: Training Customizable Role-Playing LLMs with Large-Scale Synthetic Personas
by: Wang, Xiaoyang, et al.
Published: (2025) -
Superfiltering: Weak-to-Strong Data Filtering for Fast Instruction-Tuning
by: Li, Ming, et al.
Published: (2024)