Guardado en:
| Autor principal: | Ferrari, Alan |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | https://arxiv.org/abs/2605.28384 |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
TIDE: Token-Informed Depth Execution for Per-Token Early Exit in LLM Inference
por: Jaber, Jaber, et al.
Publicado: (2026)
por: Jaber, Jaber, et al.
Publicado: (2026)
Token Sparse Attention: Efficient Long-Context Inference with Interleaved Token Selection
por: Jo, Dongwon, et al.
Publicado: (2026)
por: Jo, Dongwon, et al.
Publicado: (2026)
CAST: Clustering Self-Attention using Surrogate Tokens for Efficient Transformers
por: van Engelenhoven, Adjorn, et al.
Publicado: (2024)
por: van Engelenhoven, Adjorn, et al.
Publicado: (2024)
Learning to Route: Per-Sample Adaptive Routing for Multimodal Multitask Prediction
por: Ajirak, Marzieh, et al.
Publicado: (2025)
por: Ajirak, Marzieh, et al.
Publicado: (2025)
CITER: Collaborative Inference for Efficient Large Language Model Decoding with Token-Level Routing
por: Zheng, Wenhao, et al.
Publicado: (2025)
por: Zheng, Wenhao, et al.
Publicado: (2025)
VecFormer: Towards Efficient and Generalizable Graph Transformer with Graph Token Attention
por: Zhou, Jingbo, et al.
Publicado: (2026)
por: Zhou, Jingbo, et al.
Publicado: (2026)
Transformers Can Do Bayesian Inference
por: Müller, Samuel, et al.
Publicado: (2021)
por: Müller, Samuel, et al.
Publicado: (2021)
The Bayesian Geometry of Transformer Attention
por: Agarwal, Naman, et al.
Publicado: (2025)
por: Agarwal, Naman, et al.
Publicado: (2025)
Attention Once Is All You Need: Efficient Streaming Inference with Stateful Transformers
por: Norgren, Victor
Publicado: (2026)
por: Norgren, Victor
Publicado: (2026)
Adaptive Computation Depth via Learned Token Routing in Transformers
por: Mohammed, Ahmed Abdelmuniem Abdalla
Publicado: (2026)
por: Mohammed, Ahmed Abdelmuniem Abdalla
Publicado: (2026)
Token-Efficient RL for LLM Reasoning
por: Lee, Alan, et al.
Publicado: (2025)
por: Lee, Alan, et al.
Publicado: (2025)
Flow: Per-Instance Personalized Federated Learning Through Dynamic Routing
por: Panchal, Kunjal, et al.
Publicado: (2022)
por: Panchal, Kunjal, et al.
Publicado: (2022)
Transformers with Joint Tokens and Local-Global Attention for Efficient Human Pose Estimation
por: Kinfu, Kaleab A., et al.
Publicado: (2025)
por: Kinfu, Kaleab A., et al.
Publicado: (2025)
Universal Model Routing for Efficient LLM Inference
por: Jitkrittum, Wittawat, et al.
Publicado: (2025)
por: Jitkrittum, Wittawat, et al.
Publicado: (2025)
ToFe: Lagged Token Freezing and Reusing for Efficient Vision Transformer Inference
por: Zhang, Haoyue, et al.
Publicado: (2025)
por: Zhang, Haoyue, et al.
Publicado: (2025)
DTRNet: Dynamic Token Routing Network to Reduce Quadratic Costs in Transformers
por: Sharma, Aman, et al.
Publicado: (2025)
por: Sharma, Aman, et al.
Publicado: (2025)
KV Cache is 1 Bit Per Channel: Efficient Large Language Model Inference with Coupled Quantization
por: Zhang, Tianyi, et al.
Publicado: (2024)
por: Zhang, Tianyi, et al.
Publicado: (2024)
Can Transformers Learn Full Bayesian Inference in Context?
por: Reuter, Arik, et al.
Publicado: (2025)
por: Reuter, Arik, et al.
Publicado: (2025)
xPerT: Extended Persistence Transformer
por: Kim, Sehun
Publicado: (2024)
por: Kim, Sehun
Publicado: (2024)
Bayesian Inverse Problems Meet Flow Matching: Efficient and Flexible Inference via Transformers
por: Sherki, Daniil, et al.
Publicado: (2025)
por: Sherki, Daniil, et al.
Publicado: (2025)
Jetfire: Efficient and Accurate Transformer Pretraining with INT8 Data Flow and Per-Block Quantization
por: Xi, Haocheng, et al.
Publicado: (2024)
por: Xi, Haocheng, et al.
Publicado: (2024)
Conv-Basis: A New Paradigm for Efficient Attention Inference and Gradient Computation in Transformers
por: Liang, Yingyu, et al.
Publicado: (2024)
por: Liang, Yingyu, et al.
Publicado: (2024)
STAR: Stage-Wise Attention-Guided Token Reduction for Efficient Large Vision-Language Models Inference
por: Guo, Yichen, et al.
Publicado: (2025)
por: Guo, Yichen, et al.
Publicado: (2025)
STS: Efficient Sparse Attention with Speculative Token Sparsity
por: Xu, Ceyu, et al.
Publicado: (2026)
por: Xu, Ceyu, et al.
Publicado: (2026)
Adaptive Semantic Token Communication for Transformer-based Edge Inference
por: Devoto, Alessio, et al.
Publicado: (2025)
por: Devoto, Alessio, et al.
Publicado: (2025)
Neural Bayesian Sequential Routing
por: Huang, Yongchao
Publicado: (2026)
por: Huang, Yongchao
Publicado: (2026)
In-Context Learning Is Provably Bayesian Inference: A Generalization Theory for Meta-Learning
por: Wakayama, Tomoya, et al.
Publicado: (2025)
por: Wakayama, Tomoya, et al.
Publicado: (2025)
Token Statistics Transformer: Linear-Time Attention via Variational Rate Reduction
por: Wu, Ziyang, et al.
Publicado: (2024)
por: Wu, Ziyang, et al.
Publicado: (2024)
Semiparametric Efficient Inference in Adaptive Experiments
por: Cook, Thomas, et al.
Publicado: (2023)
por: Cook, Thomas, et al.
Publicado: (2023)
IntAttention: A Fully Integer Attention Pipeline for Efficient Edge Inference
por: Zhong, Wanli, et al.
Publicado: (2025)
por: Zhong, Wanli, et al.
Publicado: (2025)
SparQ Attention: Bandwidth-Efficient LLM Inference
por: Ribar, Luka, et al.
Publicado: (2023)
por: Ribar, Luka, et al.
Publicado: (2023)
Learning to Explain: Supervised Token Attribution from Transformer Attention Patterns
por: Mihaila, George
Publicado: (2026)
por: Mihaila, George
Publicado: (2026)
Value-State Gated Attention for Mitigating Extreme-Token Phenomena in Transformers
por: Bu, Rui, et al.
Publicado: (2025)
por: Bu, Rui, et al.
Publicado: (2025)
Unsupervised Multi-Attention Meta Transformer for Rotating Machinery Fault Diagnosis
por: Wang, Hanyang, et al.
Publicado: (2025)
por: Wang, Hanyang, et al.
Publicado: (2025)
Route Experts by Sequence, not by Token
por: Wen, Tiansheng, et al.
Publicado: (2025)
por: Wen, Tiansheng, et al.
Publicado: (2025)
Variational Routing: A Scalable Bayesian Framework for Calibrated Mixture-of-Experts Transformers
por: Li, Albus Yizhuo, et al.
Publicado: (2026)
por: Li, Albus Yizhuo, et al.
Publicado: (2026)
Token Sample Complexity of Attention
por: Bohbot, Léa, et al.
Publicado: (2025)
por: Bohbot, Léa, et al.
Publicado: (2025)
Assessing Per-Sample Membership Inference Vulnerability without Retraining
por: Dorseuil, Valentin, et al.
Publicado: (2026)
por: Dorseuil, Valentin, et al.
Publicado: (2026)
Distribution Transformers: Fast Approximate Bayesian Inference With On-The-Fly Prior Adaptation
por: Whittle, George, et al.
Publicado: (2025)
por: Whittle, George, et al.
Publicado: (2025)
Improving Routing in Sparse Mixture of Experts with Graph of Tokens
por: Nguyen, Tam, et al.
Publicado: (2025)
por: Nguyen, Tam, et al.
Publicado: (2025)
Ejemplares similares
-
TIDE: Token-Informed Depth Execution for Per-Token Early Exit in LLM Inference
por: Jaber, Jaber, et al.
Publicado: (2026) -
Token Sparse Attention: Efficient Long-Context Inference with Interleaved Token Selection
por: Jo, Dongwon, et al.
Publicado: (2026) -
CAST: Clustering Self-Attention using Surrogate Tokens for Efficient Transformers
por: van Engelenhoven, Adjorn, et al.
Publicado: (2024) -
Learning to Route: Per-Sample Adaptive Routing for Multimodal Multitask Prediction
por: Ajirak, Marzieh, et al.
Publicado: (2025) -
CITER: Collaborative Inference for Efficient Large Language Model Decoding with Token-Level Routing
por: Zheng, Wenhao, et al.
Publicado: (2025)