Guardado en:
| Autor principal: | Nidhi, Amrit |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | https://arxiv.org/abs/2605.05697 |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Budgeted LoRA: Distillation as Structured Compute Allocation for Efficient Inference
por: Sabry, Mohammed, et al.
Publicado: (2026)
por: Sabry, Mohammed, et al.
Publicado: (2026)
Conformal Thinking: Risk Control for Reasoning on a Compute Budget
por: Wang, Xi, et al.
Publicado: (2026)
por: Wang, Xi, et al.
Publicado: (2026)
One Jump Is All You Need: Short-Cutting Transformers for Early Exit Prediction with One Jump to Fit All Exit Levels
por: Seshadri, Amrit Diggavi
Publicado: (2025)
por: Seshadri, Amrit Diggavi
Publicado: (2025)
DUET: Optimize Token-Budget Allocation for Reinforcement Learning with Verifiable Rewards
por: Hu, Haoyu, et al.
Publicado: (2026)
por: Hu, Haoyu, et al.
Publicado: (2026)
LAVa: Layer-wise KV Cache Eviction with Dynamic Budget Allocation
por: Shen, Yiqun, et al.
Publicado: (2025)
por: Shen, Yiqun, et al.
Publicado: (2025)
Conv-Basis: A New Paradigm for Efficient Attention Inference and Gradient Computation in Transformers
por: Liang, Yingyu, et al.
Publicado: (2024)
por: Liang, Yingyu, et al.
Publicado: (2024)
Understanding Dynamic Compute Allocation in Recurrent Transformers
por: Moosa, Ibraheem Muhammad, et al.
Publicado: (2026)
por: Moosa, Ibraheem Muhammad, et al.
Publicado: (2026)
Alloc-MoE: Budget-Aware Expert Activation Allocation for Efficient Mixture-of-Experts Inference
por: Liu, Baihui, et al.
Publicado: (2026)
por: Liu, Baihui, et al.
Publicado: (2026)
CoBA-RL: Capability-Oriented Budget Allocation for Reinforcement Learning in LLMs
por: Yao, Zhiyuan, et al.
Publicado: (2026)
por: Yao, Zhiyuan, et al.
Publicado: (2026)
GAMMA: Global Bit Allocation for Mixed-Precision Models under Arbitrary Budgets
por: Yao, Zhangyang, et al.
Publicado: (2026)
por: Yao, Zhangyang, et al.
Publicado: (2026)
Adaptive Budget Allocation for Orthogonal-Subspace Adapter Tuning in LLMs Continual Learning
por: Wan, Zhiyi, et al.
Publicado: (2025)
por: Wan, Zhiyi, et al.
Publicado: (2025)
BaKlaVa -- Budgeted Allocation of KV cache for Long-context Inference
por: Gulhan, Ahmed Burak, et al.
Publicado: (2025)
por: Gulhan, Ahmed Burak, et al.
Publicado: (2025)
BudgetThinker: Empowering Budget-aware LLM Reasoning with Control Tokens
por: Wen, Hao, et al.
Publicado: (2025)
por: Wen, Hao, et al.
Publicado: (2025)
Reasoning on a Budget: A Survey of Adaptive and Controllable Test-Time Compute in LLMs
por: Alomrani, Mohammad Ali, et al.
Publicado: (2025)
por: Alomrani, Mohammad Ali, et al.
Publicado: (2025)
Joint Optimization of Resource Allocation and Data Selection for Fast and Cost-Efficient Federated Edge Learning
por: Jia, Yunjian, et al.
Publicado: (2024)
por: Jia, Yunjian, et al.
Publicado: (2024)
Draft-Conditioned Constrained Decoding for Structured Generation in LLMs
por: Reddy, Avinash, et al.
Publicado: (2026)
por: Reddy, Avinash, et al.
Publicado: (2026)
Unveiling and Controlling Anomalous Attention Distribution in Transformers
por: Yan, Ruiqing, et al.
Publicado: (2024)
por: Yan, Ruiqing, et al.
Publicado: (2024)
ZeroS: Zero-Sum Linear Attention for Efficient Transformers
por: Lu, Jiecheng, et al.
Publicado: (2026)
por: Lu, Jiecheng, et al.
Publicado: (2026)
Knapsack RL: Unlocking Exploration of LLMs via Optimizing Budget Allocation
por: Li, Ziniu, et al.
Publicado: (2025)
por: Li, Ziniu, et al.
Publicado: (2025)
Attention Schema-based Attention Control (ASAC): A Cognitive-Inspired Approach for Attention Management in Transformers
por: Saxena, Krati, et al.
Publicado: (2025)
por: Saxena, Krati, et al.
Publicado: (2025)
LOOKAT: Lookup-Optimized Key-Attention for Memory-Efficient Transformers
por: Karmore, Aryan
Publicado: (2026)
por: Karmore, Aryan
Publicado: (2026)
Fine-Grained Iterative Adversarial Attacks with Limited Computation Budget
por: Hou, Zhichao, et al.
Publicado: (2025)
por: Hou, Zhichao, et al.
Publicado: (2025)
To Trust or Not to Trust: On Calibration in ML-based Resource Allocation for Wireless Networks
por: Raina, Rashika, et al.
Publicado: (2025)
por: Raina, Rashika, et al.
Publicado: (2025)
DARA: Few-shot Budget Allocation in Online Advertising via In-Context Decision Making with RL-Finetuned LLMs
por: Song, Mingxuan, et al.
Publicado: (2026)
por: Song, Mingxuan, et al.
Publicado: (2026)
Budget-Constrained Agentic Large Language Models: Intention-Based Planning for Costly Tool Use
por: Liu, Hanbing, et al.
Publicado: (2026)
por: Liu, Hanbing, et al.
Publicado: (2026)
Confidence-Controlled Exploration: Efficient Sparse-Reward Policy Learning for Robot Navigation
por: Patel, Bhrij, et al.
Publicado: (2023)
por: Patel, Bhrij, et al.
Publicado: (2023)
Attention Needs to Focus: A Unified Perspective on Attention Allocation
por: Fu, Zichuan, et al.
Publicado: (2026)
por: Fu, Zichuan, et al.
Publicado: (2026)
ORBIT: On-policy Exploration-Exploitation for Controllable Multi-Budget Reasoning
por: Liang, Kun, et al.
Publicado: (2026)
por: Liang, Kun, et al.
Publicado: (2026)
Fine-Grained Graph Generation through Latent Mixture Scheduling
por: Vakil, Nidhi, et al.
Publicado: (2026)
por: Vakil, Nidhi, et al.
Publicado: (2026)
VecFormer: Towards Efficient and Generalizable Graph Transformer with Graph Token Attention
por: Zhou, Jingbo, et al.
Publicado: (2026)
por: Zhou, Jingbo, et al.
Publicado: (2026)
AtMan: Understanding Transformer Predictions Through Memory Efficient Attention Manipulation
por: Deiseroth, Björn, et al.
Publicado: (2023)
por: Deiseroth, Björn, et al.
Publicado: (2023)
Predict-then-Diffuse: Adaptive Response Length for Compute-Budgeted Inference in Diffusion LLMs
por: Rottoli, Michael, et al.
Publicado: (2026)
por: Rottoli, Michael, et al.
Publicado: (2026)
Beyond Hard Constraints: Budget-Conditioned Reachability For Safe Offline Reinforcement Learning
por: Brahmanage, Janaka Chathuranga, et al.
Publicado: (2026)
por: Brahmanage, Janaka Chathuranga, et al.
Publicado: (2026)
The Bayesian Geometry of Transformer Attention
por: Agarwal, Naman, et al.
Publicado: (2025)
por: Agarwal, Naman, et al.
Publicado: (2025)
Signature-Informed Transformer for Asset Allocation
por: Hwang, Yoontae, et al.
Publicado: (2025)
por: Hwang, Yoontae, et al.
Publicado: (2025)
Depth-Structured Music Recurrence: Budgeted Recurrent Attention for Full-Piece Symbolic Music Modeling
por: Yi, Yungang, et al.
Publicado: (2026)
por: Yi, Yungang, et al.
Publicado: (2026)
Crossfusor: A Cross-Attention Transformer Enhanced Conditional Diffusion Model for Car-Following Trajectory Prediction
por: You, Junwei, et al.
Publicado: (2024)
por: You, Junwei, et al.
Publicado: (2024)
Do Efficient Transformers Really Save Computation?
por: Yang, Kai, et al.
Publicado: (2024)
por: Yang, Kai, et al.
Publicado: (2024)
Not All Turns Are Equally Hard: Adaptive Thinking Budgets For Efficient Multi-Turn Reasoning
por: Jali, Neharika, et al.
Publicado: (2026)
por: Jali, Neharika, et al.
Publicado: (2026)
Geometric Attention: A Regime-Explicit Operator Semantics for Transformer Attention
por: Freytes, Luis Rosario
Publicado: (2026)
por: Freytes, Luis Rosario
Publicado: (2026)
Ejemplares similares
-
Budgeted LoRA: Distillation as Structured Compute Allocation for Efficient Inference
por: Sabry, Mohammed, et al.
Publicado: (2026) -
Conformal Thinking: Risk Control for Reasoning on a Compute Budget
por: Wang, Xi, et al.
Publicado: (2026) -
One Jump Is All You Need: Short-Cutting Transformers for Early Exit Prediction with One Jump to Fit All Exit Levels
por: Seshadri, Amrit Diggavi
Publicado: (2025) -
DUET: Optimize Token-Budget Allocation for Reinforcement Learning with Verifiable Rewards
por: Hu, Haoyu, et al.
Publicado: (2026) -
LAVa: Layer-wise KV Cache Eviction with Dynamic Budget Allocation
por: Shen, Yiqun, et al.
Publicado: (2025)