MonarchAttention: Zero-Shot Conversion to Fast, Hardware-Aware Structured Attention
Fuente:
arXiv
Saved in:
| Main Authors: | Yaras, Can, Xu, Alec S., Abillama, Pierre, Lee, Changwoo, Balzano, Laura |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Emergent Low-Rank Training Dynamics in MLPs with Smooth Activations
by: Xu, Alec S., et al.
Published: (2026)
by: Xu, Alec S., et al.
Published: (2026)
Out-of-Distribution Generalization of In-Context Learning: A Low-Dimensional Subspace Perspective
by: Kwon, Soo Min, et al.
Published: (2025)
by: Kwon, Soo Min, et al.
Published: (2025)
Compressible Dynamics in Deep Overparameterized Low-Rank Learning & Adaptation
by: Yaras, Can, et al.
Published: (2024)
by: Yaras, Can, et al.
Published: (2024)
Linearly Separable Features in Shallow Nonlinear Networks: Width Scales Polynomially with Intrinsic Data Dimension
by: Xu, Alec S., et al.
Published: (2025)
by: Xu, Alec S., et al.
Published: (2025)
Memory-Efficient Acceleration of Block Low-Rank Foundation Models on Resource Constrained GPUs
by: Abillama, Pierre, et al.
Published: (2025)
by: Abillama, Pierre, et al.
Published: (2025)
Understanding Deep Representation Learning via Layerwise Feature Compression and Discrimination
by: Wang, Peng, et al.
Published: (2023)
by: Wang, Peng, et al.
Published: (2023)
Hardware-Efficient Attention for Fast Decoding
by: Zadouri, Ted, et al.
Published: (2025)
by: Zadouri, Ted, et al.
Published: (2025)
An Overview of Low-Rank Structures in the Training and Adaptation of Large Models
by: Balzano, Laura, et al.
Published: (2025)
by: Balzano, Laura, et al.
Published: (2025)
MonarchRT: Efficient Attention for Real-Time Video Generation
by: Agarwal, Krish, et al.
Published: (2026)
by: Agarwal, Krish, et al.
Published: (2026)
Stochastic Sparse Attention for Memory-Bound Inference
by: Lee, Kyle, et al.
Published: (2026)
by: Lee, Kyle, et al.
Published: (2026)
Modality-Aware Zero-Shot Pruning and Sparse Attention for Efficient Multimodal Edge Inference
by: Sui, Yueyuan, et al.
Published: (2026)
by: Sui, Yueyuan, et al.
Published: (2026)
Cross-RAG: Zero-Shot Retrieval-Augmented Time Series Forecasting via Cross-Attention
by: Lee, Seunghan, et al.
Published: (2026)
by: Lee, Seunghan, et al.
Published: (2026)
Delta Attention: Fast and Accurate Sparse Attention Inference by Delta Correction
by: Willette, Jeffrey, et al.
Published: (2025)
by: Willette, Jeffrey, et al.
Published: (2025)
Lean Attention: Hardware-Aware Scalable Attention Mechanism for the Decode-Phase of Transformers
by: Sanovar, Rya, et al.
Published: (2024)
by: Sanovar, Rya, et al.
Published: (2024)
Multi-label Zero-Shot Audio Classification with Temporal Attention
by: Dogan, Duygu, et al.
Published: (2024)
by: Dogan, Duygu, et al.
Published: (2024)
Masked Extended Attention for Zero-Shot Virtual Try-On In The Wild
by: Orzech, Nadav, et al.
Published: (2024)
by: Orzech, Nadav, et al.
Published: (2024)
Fair Community Detection and Structure Learning in Heterogeneous Graphical Models
by: Tarzanagh, Davoud Ataee, et al.
Published: (2021)
by: Tarzanagh, Davoud Ataee, et al.
Published: (2021)
Explaining and Mitigating the Modality Gap in Contrastive Multimodal Learning
by: Yaras, Can, et al.
Published: (2024)
by: Yaras, Can, et al.
Published: (2024)
Structure-Aware Set Transformers: Temporal and Variable-Type Attention Biases for Asynchronous Clinical Time Series
by: Lee, Joohyung, et al.
Published: (2026)
by: Lee, Joohyung, et al.
Published: (2026)
Scaling Probabilistic Circuits via Monarch Matrices
by: Zhang, Honghua, et al.
Published: (2025)
by: Zhang, Honghua, et al.
Published: (2025)
ZeroS: Zero-Sum Linear Attention for Efficient Transformers
by: Lu, Jiecheng, et al.
Published: (2026)
by: Lu, Jiecheng, et al.
Published: (2026)
AgilePruner: An Empirical Study of Attention and Diversity for Adaptive Visual Token Pruning in Large Vision-Language Models
by: Baek, Changwoo, et al.
Published: (2026)
by: Baek, Changwoo, et al.
Published: (2026)
Native Sparse Attention: Hardware-Aligned and Natively Trainable Sparse Attention
by: Yuan, Jingyang, et al.
Published: (2025)
by: Yuan, Jingyang, et al.
Published: (2025)
HATA: Trainable and Hardware-Efficient Hash-Aware Top-k Attention for Scalable Large Model Inference
by: Gong, Ping, et al.
Published: (2025)
by: Gong, Ping, et al.
Published: (2025)
A Spectral Framework for Tracking Communities in Evolving Networks
by: Hume, Jacob, et al.
Published: (2024)
by: Hume, Jacob, et al.
Published: (2024)
AnchorAttention: Difference-Aware Sparse Attention with Stripe Granularity
by: Zhang, Yu, et al.
Published: (2025)
by: Zhang, Yu, et al.
Published: (2025)
Attention Based Simple Primitives for Open World Compositional Zero-Shot Learning
by: Munir, Ans, et al.
Published: (2024)
by: Munir, Ans, et al.
Published: (2024)
MARché: Fast Masked Autoregressive Image Generation with Cache-Aware Attention
by: Jiang, Chaoyi, et al.
Published: (2025)
by: Jiang, Chaoyi, et al.
Published: (2025)
DAFMSVC: One-Shot Singing Voice Conversion with Dual Attention Mechanism and Flow Matching
by: Chen, Wei, et al.
Published: (2025)
by: Chen, Wei, et al.
Published: (2025)
FastAttention: Extend FlashAttention2 to NPUs and Low-resource GPUs
by: Lin, Haoran, et al.
Published: (2024)
by: Lin, Haoran, et al.
Published: (2024)
OTSeg: Multi-prompt Sinkhorn Attention for Zero-Shot Semantic Segmentation
by: Kim, Kwanyoung, et al.
Published: (2024)
by: Kim, Kwanyoung, et al.
Published: (2024)
Exact Conversion of In-Context Learning to Model Weights in Linearized-Attention Transformers
by: Chen, Brian K, et al.
Published: (2024)
by: Chen, Brian K, et al.
Published: (2024)
AttentionEngine: A Versatile Framework for Efficient Attention Mechanisms on Diverse Hardware Platforms
by: Chen, Feiyang, et al.
Published: (2025)
by: Chen, Feiyang, et al.
Published: (2025)
How Smooth Is Attention?
by: Castin, Valérie, et al.
Published: (2023)
by: Castin, Valérie, et al.
Published: (2023)
ZEST: Attention-based Zero-Shot Learning for Unseen IoT Device Classification
by: Wu, Binghui, et al.
Published: (2023)
by: Wu, Binghui, et al.
Published: (2023)
Gated Linear Attention Transformers with Hardware-Efficient Training
by: Yang, Songlin, et al.
Published: (2023)
by: Yang, Songlin, et al.
Published: (2023)
Attention-based clustering
by: Maulen-Soto, Rodrigo, et al.
Published: (2025)
by: Maulen-Soto, Rodrigo, et al.
Published: (2025)
FlashAttention-3: Fast and Accurate Attention with Asynchrony and Low-precision
by: Shah, Jay, et al.
Published: (2024)
by: Shah, Jay, et al.
Published: (2024)
FlashBias: Fast Computation of Attention with Bias
by: Wu, Haixu, et al.
Published: (2025)
by: Wu, Haixu, et al.
Published: (2025)
Fast KV Compaction via Attention Matching
by: Zweiger, Adam, et al.
Published: (2026)
by: Zweiger, Adam, et al.
Published: (2026)
Similar Items
-
Emergent Low-Rank Training Dynamics in MLPs with Smooth Activations
by: Xu, Alec S., et al.
Published: (2026) -
Out-of-Distribution Generalization of In-Context Learning: A Low-Dimensional Subspace Perspective
by: Kwon, Soo Min, et al.
Published: (2025) -
Compressible Dynamics in Deep Overparameterized Low-Rank Learning & Adaptation
by: Yaras, Can, et al.
Published: (2024) -
Linearly Separable Features in Shallow Nonlinear Networks: Width Scales Polynomially with Intrinsic Data Dimension
by: Xu, Alec S., et al.
Published: (2025) -
Memory-Efficient Acceleration of Block Low-Rank Foundation Models on Resource Constrained GPUs
by: Abillama, Pierre, et al.
Published: (2025)