Self-Tuning Sparse Attention: Multi-Fidelity Hyperparameter Optimization for Transformer Acceleration
Fuente:
arXiv
Guardado en:
| Autores principales: | Dev, Arundhathi, Zhan, Justin |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Efficient Domain Adaptation for Text Line Recognition via Decoupled Language Models
por: Dev, Arundhathi, et al.
Publicado: (2026)
por: Dev, Arundhathi, et al.
Publicado: (2026)
Bayesian Optimization for Hyperparameters Tuning in Neural Networks
por: Onorato, Gabriele
Publicado: (2024)
por: Onorato, Gabriele
Publicado: (2024)
OptiMindTune: A Multi-Agent Framework for Intelligent Hyperparameter Optimization
por: Madiraju, Meher Bhaskar, et al.
Publicado: (2025)
por: Madiraju, Meher Bhaskar, et al.
Publicado: (2025)
Interactive Hyperparameter Optimization in Multi-Objective Problems via Preference Learning
por: Giovanelli, Joseph, et al.
Publicado: (2023)
por: Giovanelli, Joseph, et al.
Publicado: (2023)
A Unified Hyperparameter Optimization Pipeline for Transformer-Based Time Series Forecasting Models
por: Xu, Jingjing, et al.
Publicado: (2025)
por: Xu, Jingjing, et al.
Publicado: (2025)
ORTHOBO: Orthogonal Bayesian Hyperparameter Optimization
por: Schröder, Maresa, et al.
Publicado: (2026)
por: Schröder, Maresa, et al.
Publicado: (2026)
HSR-Enhanced Sparse Attention Acceleration
por: Chen, Bo, et al.
Publicado: (2024)
por: Chen, Bo, et al.
Publicado: (2024)
SEA: State-Exchange Attention for High-Fidelity Physics Based Transformers
por: Esmati, Parsa, et al.
Publicado: (2024)
por: Esmati, Parsa, et al.
Publicado: (2024)
Fine-Tuning Adaptive Stochastic Optimizers: Determining the Optimal Hyperparameter $ε$ via Gradient Magnitude Histogram Analysis
por: Silva, Gustavo, et al.
Publicado: (2023)
por: Silva, Gustavo, et al.
Publicado: (2023)
Sparse-vDiT: Unleashing the Power of Sparse Attention to Accelerate Video Diffusion Transformers
por: Chen, Pengtao, et al.
Publicado: (2025)
por: Chen, Pengtao, et al.
Publicado: (2025)
Using Large Language Models for Hyperparameter Optimization
por: Zhang, Michael R., et al.
Publicado: (2023)
por: Zhang, Michael R., et al.
Publicado: (2023)
Hyperparameter Optimization via Interacting with Probabilistic Circuits
por: Seng, Jonas, et al.
Publicado: (2025)
por: Seng, Jonas, et al.
Publicado: (2025)
Sequential Policy Gradient for Adaptive Hyperparameter Optimization
por: Li, Zheng, et al.
Publicado: (2025)
por: Li, Zheng, et al.
Publicado: (2025)
MultiMax: Sparse and Multi-Modal Attention Learning
por: Zhou, Yuxuan, et al.
Publicado: (2024)
por: Zhou, Yuxuan, et al.
Publicado: (2024)
Multi-Objective Multi-Fidelity Bayesian Optimization with Causal Priors
por: Hossen, Md Abir, et al.
Publicado: (2026)
por: Hossen, Md Abir, et al.
Publicado: (2026)
Large Language Model Enhanced Particle Swarm Optimization for Hyperparameter Tuning for Deep Learning Models
por: Hameed, Saad, et al.
Publicado: (2025)
por: Hameed, Saad, et al.
Publicado: (2025)
Scaling Graph Transformers: A Comparative Study of Sparse and Dense Attention
por: Dimitrov, Leon
Publicado: (2025)
por: Dimitrov, Leon
Publicado: (2025)
Hyperparameter Importance Analysis for Multi-Objective AutoML
por: Theodorakopoulos, Daphne, et al.
Publicado: (2024)
por: Theodorakopoulos, Daphne, et al.
Publicado: (2024)
A Unified Gaussian Process for Branching and Nested Hyperparameter Optimization
por: Zhang, Jiazhao, et al.
Publicado: (2024)
por: Zhang, Jiazhao, et al.
Publicado: (2024)
HyperSHAP: Shapley Values and Interactions for Explaining Hyperparameter Optimization
por: Wever, Marcel, et al.
Publicado: (2025)
por: Wever, Marcel, et al.
Publicado: (2025)
Frozen Layers: Memory-efficient Many-fidelity Hyperparameter Optimization
por: Carstensen, Timur, et al.
Publicado: (2025)
por: Carstensen, Timur, et al.
Publicado: (2025)
Self-Indexing KVCache: Predicting Sparse Attention from Compressed Keys
por: Yang, Xu, et al.
Publicado: (2026)
por: Yang, Xu, et al.
Publicado: (2026)
Graph Convolutions Enrich the Self-Attention in Transformers!
por: Choi, Jeongwhan, et al.
Publicado: (2023)
por: Choi, Jeongwhan, et al.
Publicado: (2023)
vAttention: Verified Sparse Attention
por: Desai, Aditya, et al.
Publicado: (2025)
por: Desai, Aditya, et al.
Publicado: (2025)
Stay Tuned: An Empirical Study of the Impact of Hyperparameters on LLM Tuning in Real-World Applications
por: Halfon, Alon, et al.
Publicado: (2024)
por: Halfon, Alon, et al.
Publicado: (2024)
On the Learn-to-Optimize Capabilities of Transformers in In-Context Sparse Recovery
por: Liu, Renpu, et al.
Publicado: (2024)
por: Liu, Renpu, et al.
Publicado: (2024)
Accelerating Large-Scale Reasoning Model Inference with Sparse Self-Speculative Decoding
por: Zhao, Yilong, et al.
Publicado: (2025)
por: Zhao, Yilong, et al.
Publicado: (2025)
ECG-NAT: A Self-supervised Neighborhood Attention Transformer for Multi-lead Electrocardiogram Classification
por: Gazeran, Mahsa, et al.
Publicado: (2026)
por: Gazeran, Mahsa, et al.
Publicado: (2026)
MoESD: Unveil Speculative Decoding's Potential for Accelerating Sparse MoE
por: Huang, Zongle, et al.
Publicado: (2025)
por: Huang, Zongle, et al.
Publicado: (2025)
Speeding Up Multi-Objective Hyperparameter Optimization by Task Similarity-Based Meta-Learning for the Tree-Structured Parzen Estimator
por: Watanabe, Shuhei, et al.
Publicado: (2022)
por: Watanabe, Shuhei, et al.
Publicado: (2022)
Orion-MSP: Multi-Scale Sparse Attention for Tabular In-Context Learning
por: Bouadi, Mohamed, et al.
Publicado: (2025)
por: Bouadi, Mohamed, et al.
Publicado: (2025)
Sparse Low-Ranked Self-Attention Transformer for Remaining Useful Lifetime Prediction of Optical Fiber Amplifiers
por: Schneider, Dominic, et al.
Publicado: (2024)
por: Schneider, Dominic, et al.
Publicado: (2024)
Fast Benchmarking of Asynchronous Multi-Fidelity Optimization on Zero-Cost Benchmarks
por: Watanabe, Shuhei, et al.
Publicado: (2024)
por: Watanabe, Shuhei, et al.
Publicado: (2024)
Time-Efficient Hybrid Hyperparameter Tuning Approach for Cardiovascular Disease Classification
por: Pathak, Abhay Kumar, et al.
Publicado: (2024)
por: Pathak, Abhay Kumar, et al.
Publicado: (2024)
FlashOmni: A Unified Sparse Attention Engine for Diffusion Transformers
por: Qiao, Liang, et al.
Publicado: (2025)
por: Qiao, Liang, et al.
Publicado: (2025)
Default Machine Learning Hyperparameters Do Not Provide Informative Initialization for Bayesian Optimization
por: Prieto, Nicolás Villagrán, et al.
Publicado: (2026)
por: Prieto, Nicolás Villagrán, et al.
Publicado: (2026)
ULTHO: Ultra-Lightweight yet Efficient Hyperparameter Optimization in Deep Reinforcement Learning
por: Yuan, Mingqi, et al.
Publicado: (2025)
por: Yuan, Mingqi, et al.
Publicado: (2025)
Cross-Entropy Optimization for Hyperparameter Optimization in Stochastic Gradient-based Approaches to Train Deep Neural Networks
por: Li, Kevin, et al.
Publicado: (2024)
por: Li, Kevin, et al.
Publicado: (2024)
SampleAttention: Near-Lossless Acceleration of Long Context LLM Inference with Adaptive Structured Sparse Attention
por: Zhu, Qianchao, et al.
Publicado: (2024)
por: Zhu, Qianchao, et al.
Publicado: (2024)
Cost-Sensitive Multi-Fidelity Bayesian Optimization with Transfer of Learning Curve Extrapolation
por: Lee, Dong Bok, et al.
Publicado: (2024)
por: Lee, Dong Bok, et al.
Publicado: (2024)
Ejemplares similares
-
Efficient Domain Adaptation for Text Line Recognition via Decoupled Language Models
por: Dev, Arundhathi, et al.
Publicado: (2026) -
Bayesian Optimization for Hyperparameters Tuning in Neural Networks
por: Onorato, Gabriele
Publicado: (2024) -
OptiMindTune: A Multi-Agent Framework for Intelligent Hyperparameter Optimization
por: Madiraju, Meher Bhaskar, et al.
Publicado: (2025) -
Interactive Hyperparameter Optimization in Multi-Objective Problems via Preference Learning
por: Giovanelli, Joseph, et al.
Publicado: (2023) -
A Unified Hyperparameter Optimization Pipeline for Transformer-Based Time Series Forecasting Models
por: Xu, Jingjing, et al.
Publicado: (2025)