LLM Router: Rethinking Routing with Prefill Activations
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Varshney, Tanay, Surla, Annie, Xu, Michelle, Krishnan, Gomathy Venkata, Jeblick, Maximilian, Austin, David, Vaidya, Neal, Onofrio, Davide |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
KVzap: Fast, Adaptive, and Faithful KV Cache Pruning
von: Jegou, Simon, et al.
Veröffentlicht: (2026)
von: Jegou, Simon, et al.
Veröffentlicht: (2026)
CLAA: Cross-Layer Attention Aggregation for Accelerating LLM Prefill
von: McDanel, Bradley, et al.
Veröffentlicht: (2026)
von: McDanel, Bradley, et al.
Veröffentlicht: (2026)
Expected Attention: KV Cache Compression by Estimating Attention from Future Queries Distribution
von: Devoto, Alessio, et al.
Veröffentlicht: (2025)
von: Devoto, Alessio, et al.
Veröffentlicht: (2025)
Route to Rome Attack: Directing LLM Routers to Expensive Models via Adversarial Suffix Optimization
von: Tang, Haochun, et al.
Veröffentlicht: (2026)
von: Tang, Haochun, et al.
Veröffentlicht: (2026)
OrcaRouter: A Production-Oriented LLM Router with Hybrid Offline-Online Learning
von: Bao, Zhenghua, et al.
Veröffentlicht: (2026)
von: Bao, Zhenghua, et al.
Veröffentlicht: (2026)
ThinkRouter: Efficient Reasoning via Routing Thinking between Latent and Discrete Spaces
von: Xu, Xin, et al.
Veröffentlicht: (2026)
von: Xu, Xin, et al.
Veröffentlicht: (2026)
OmniRouter: Budget and Performance Controllable Multi-LLM Routing
von: Mei, Kai, et al.
Veröffentlicht: (2025)
von: Mei, Kai, et al.
Veröffentlicht: (2025)
VL-RouterBench: A Benchmark for Vision-Language Model Routing
von: Huang, Zhehao, et al.
Veröffentlicht: (2025)
von: Huang, Zhehao, et al.
Veröffentlicht: (2025)
HierRouter: Coordinated Routing of Specialized Large Language Models via Reinforcement Learning
von: Gupta, Nikunj, et al.
Veröffentlicht: (2025)
von: Gupta, Nikunj, et al.
Veröffentlicht: (2025)
H2O-Danube-1.8B Technical Report
von: Singer, Philipp, et al.
Veröffentlicht: (2024)
von: Singer, Philipp, et al.
Veröffentlicht: (2024)
Arch-Router: Aligning LLM Routing with Human Preferences
von: Tran, Co, et al.
Veröffentlicht: (2025)
von: Tran, Co, et al.
Veröffentlicht: (2025)
AttnCache: Accelerating Self-Attention Inference for LLM Prefill via Attention Cache
von: Song, Dinghong, et al.
Veröffentlicht: (2025)
von: Song, Dinghong, et al.
Veröffentlicht: (2025)
CritiPrefill: A Segment-wise Criticality-based Approach for Prefilling Acceleration in LLMs
von: Lv, Junlin, et al.
Veröffentlicht: (2024)
von: Lv, Junlin, et al.
Veröffentlicht: (2024)
Block-Attention for Efficient Prefilling
von: Ma, Dongyang, et al.
Veröffentlicht: (2024)
von: Ma, Dongyang, et al.
Veröffentlicht: (2024)
GMTRouter: Personalized LLM Router over Multi-turn User Interactions
von: Xie, Encheng, et al.
Veröffentlicht: (2025)
von: Xie, Encheng, et al.
Veröffentlicht: (2025)
Router-R1: Teaching LLMs Multi-Round Routing and Aggregation via Reinforcement Learning
von: Zhang, Haozhen, et al.
Veröffentlicht: (2025)
von: Zhang, Haozhen, et al.
Veröffentlicht: (2025)
RouteNator: A Router-Based Multi-Modal Architecture for Generating Synthetic Training Data for Function Calling LLMs
von: Belavadi, Vibha, et al.
Veröffentlicht: (2025)
von: Belavadi, Vibha, et al.
Veröffentlicht: (2025)
SkillRouter: Skill Routing for LLM Agents at Scale
von: Zheng, YanZhao, et al.
Veröffentlicht: (2026)
von: Zheng, YanZhao, et al.
Veröffentlicht: (2026)
R2-Router: A New Paradigm for LLM Routing with Reasoning
von: Xue, Jiaqi, et al.
Veröffentlicht: (2026)
von: Xue, Jiaqi, et al.
Veröffentlicht: (2026)
Queryable LoRA: Instruction-Regularized Routing Over Shared Low-Rank Update Atoms
von: Vaidya, Omatharv Bharat, et al.
Veröffentlicht: (2026)
von: Vaidya, Omatharv Bharat, et al.
Veröffentlicht: (2026)
Rethinking Predictive Modeling for LLM Routing: When Simple kNN Beats Complex Learned Routers
von: Li, Yang
Veröffentlicht: (2025)
von: Li, Yang
Veröffentlicht: (2025)
How Robust Are Router-LLMs? Analysis of the Fragility of LLM Routing Capabilities
von: Kassem, Aly M., et al.
Veröffentlicht: (2025)
von: Kassem, Aly M., et al.
Veröffentlicht: (2025)
Sockpuppetting: Jailbreaking LLMs by Combining Prefilling with Optimization
von: Dotsinski, Asen, et al.
Veröffentlicht: (2026)
von: Dotsinski, Asen, et al.
Veröffentlicht: (2026)
Omni-Router: Sharing Routing Decisions in Sparse Mixture-of-Experts for Speech Recognition
von: Gu, Zijin, et al.
Veröffentlicht: (2025)
von: Gu, Zijin, et al.
Veröffentlicht: (2025)
RouterDC: Query-Based Router by Dual Contrastive Learning for Assembling Large Language Models
von: Chen, Shuhao, et al.
Veröffentlicht: (2024)
von: Chen, Shuhao, et al.
Veröffentlicht: (2024)
Rethinking LLM Ensembling from the Perspective of Mixture Models
von: Fu, Jiale, et al.
Veröffentlicht: (2026)
von: Fu, Jiale, et al.
Veröffentlicht: (2026)
Glider: Global and Local Instruction-Driven Expert Router
von: Li, Pingzhi, et al.
Veröffentlicht: (2024)
von: Li, Pingzhi, et al.
Veröffentlicht: (2024)
CP-Router: An Uncertainty-Aware Router Between LLM and LRM
von: Su, Jiayuan, et al.
Veröffentlicht: (2025)
von: Su, Jiayuan, et al.
Veröffentlicht: (2025)
FastKV: Decoupling of Context Reduction and KV Cache Compression for Prefill-Decoding Acceleration
von: Jo, Dongwon, et al.
Veröffentlicht: (2025)
von: Jo, Dongwon, et al.
Veröffentlicht: (2025)
Rethinking Machine Unlearning for Large Language Models
von: Liu, Sijia, et al.
Veröffentlicht: (2024)
von: Liu, Sijia, et al.
Veröffentlicht: (2024)
PreFT: Prefill-only finetuning for efficient inference
von: Lanpouthakoun, Andrew, et al.
Veröffentlicht: (2026)
von: Lanpouthakoun, Andrew, et al.
Veröffentlicht: (2026)
Accelerating Prefilling for Long-Context LLMs via Sparse Pattern Sharing
von: Peng, Dan, et al.
Veröffentlicht: (2025)
von: Peng, Dan, et al.
Veröffentlicht: (2025)
Prefill-Guided Thinking for zero-shot detection of AI-generated images
von: Kachwala, Zoher, et al.
Veröffentlicht: (2025)
von: Kachwala, Zoher, et al.
Veröffentlicht: (2025)
FlexPrefill: A Context-Aware Sparse Attention Mechanism for Efficient Long-Sequence Inference
von: Lai, Xunhao, et al.
Veröffentlicht: (2025)
von: Lai, Xunhao, et al.
Veröffentlicht: (2025)
Coupling Experts and Routers in Mixture-of-Experts via an Auxiliary Loss
von: Lv, Ang, et al.
Veröffentlicht: (2025)
von: Lv, Ang, et al.
Veröffentlicht: (2025)
RouteLLM: Learning to Route LLMs with Preference Data
von: Ong, Isaac, et al.
Veröffentlicht: (2024)
von: Ong, Isaac, et al.
Veröffentlicht: (2024)
ICL-Router: In-Context Learned Model Representations for LLM Routing
von: Wang, Chenxu, et al.
Veröffentlicht: (2025)
von: Wang, Chenxu, et al.
Veröffentlicht: (2025)
RouterBench: A Benchmark for Multi-LLM Routing System
von: Hu, Qitian Jason, et al.
Veröffentlicht: (2024)
von: Hu, Qitian Jason, et al.
Veröffentlicht: (2024)
IntroLM: Introspective Language Models via Prefilling-Time Self-Evaluation
von: Kasnavieh, Hossein Hosseini, et al.
Veröffentlicht: (2026)
von: Kasnavieh, Hossein Hosseini, et al.
Veröffentlicht: (2026)
SwiftKV: Fast Prefill-Optimized Inference with Knowledge-Preserving Model Transformation
von: Qiao, Aurick, et al.
Veröffentlicht: (2024)
von: Qiao, Aurick, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
KVzap: Fast, Adaptive, and Faithful KV Cache Pruning
von: Jegou, Simon, et al.
Veröffentlicht: (2026) -
CLAA: Cross-Layer Attention Aggregation for Accelerating LLM Prefill
von: McDanel, Bradley, et al.
Veröffentlicht: (2026) -
Expected Attention: KV Cache Compression by Estimating Attention from Future Queries Distribution
von: Devoto, Alessio, et al.
Veröffentlicht: (2025) -
Route to Rome Attack: Directing LLM Routers to Expensive Models via Adversarial Suffix Optimization
von: Tang, Haochun, et al.
Veröffentlicht: (2026) -
OrcaRouter: A Production-Oriented LLM Router with Hybrid Offline-Online Learning
von: Bao, Zhenghua, et al.
Veröffentlicht: (2026)