SequentialAttention++ for Block Sparsification: Differentiable Pruning Meets Combinatorial Optimization
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Yasuda, Taisuke, Axiotis, Kyriakos, Fu, Gang, Bateni, MohammadHossein, Mirrokni, Vahab |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Sequential Attention for Feature Selection
von: Yasuda, Taisuke, et al.
Veröffentlicht: (2022)
von: Yasuda, Taisuke, et al.
Veröffentlicht: (2022)
DeepCrossAttention: Supercharging Transformer Residual Connections
von: Heddes, Mike, et al.
Veröffentlicht: (2025)
von: Heddes, Mike, et al.
Veröffentlicht: (2025)
Budget Allocation for Unknown Value Functions in a Lipschitz Space
von: Bateni, MohammadHossein, et al.
Veröffentlicht: (2025)
von: Bateni, MohammadHossein, et al.
Veröffentlicht: (2025)
Data-Efficient Learning via Clustering-Based Sensitivity Sampling: Foundation Models and Beyond
von: Axiotis, Kyriakos, et al.
Veröffentlicht: (2024)
von: Axiotis, Kyriakos, et al.
Veröffentlicht: (2024)
Replicable Composition
von: Banihashem, Kiarash, et al.
Veröffentlicht: (2026)
von: Banihashem, Kiarash, et al.
Veröffentlicht: (2026)
It's All Connected: A Journey Through Test-Time Memorization, Attentional Bias, Retention, and Online Optimization
von: Behrouz, Ali, et al.
Veröffentlicht: (2025)
von: Behrouz, Ali, et al.
Veröffentlicht: (2025)
Approximately Optimal Core Shapes for Tensor Decompositions
von: Ghadiri, Mehrdad, et al.
Veröffentlicht: (2023)
von: Ghadiri, Mehrdad, et al.
Veröffentlicht: (2023)
A Scalable Algorithm for Individually Fair K-means Clustering
von: Bateni, MohammadHossein, et al.
Veröffentlicht: (2024)
von: Bateni, MohammadHossein, et al.
Veröffentlicht: (2024)
Synthetic Text Generation for Training Large Language Models via Gradient Matching
von: Nguyen, Dang, et al.
Veröffentlicht: (2025)
von: Nguyen, Dang, et al.
Veröffentlicht: (2025)
Improving the Variance of Differentially Private Randomized Experiments through Clustering
von: Javanmard, Adel, et al.
Veröffentlicht: (2023)
von: Javanmard, Adel, et al.
Veröffentlicht: (2023)
Data Selection for Fine-tuning Vision Language Models via Cross Modal Alignment Trajectories
von: Naharas, Nilay, et al.
Veröffentlicht: (2025)
von: Naharas, Nilay, et al.
Veröffentlicht: (2025)
Trellis: Learning to Compress Key-Value Memory in Attention Models
von: Karami, Mahdi, et al.
Veröffentlicht: (2025)
von: Karami, Mahdi, et al.
Veröffentlicht: (2025)
Smooth Anonymity for Sparse Graphs
von: Epasto, Alessandro, et al.
Veröffentlicht: (2022)
von: Epasto, Alessandro, et al.
Veröffentlicht: (2022)
PolySketchFormer: Fast Transformers via Sketching Polynomial Kernels
von: Kacham, Praneeth, et al.
Veröffentlicht: (2023)
von: Kacham, Praneeth, et al.
Veröffentlicht: (2023)
Learning from Aggregate responses: Instance Level versus Bag Level Loss Functions
von: Javanmard, Adel, et al.
Veröffentlicht: (2024)
von: Javanmard, Adel, et al.
Veröffentlicht: (2024)
DiSK: Differentially Private Optimizer with Simplified Kalman Filter for Noise Reduction
von: Zhang, Xinwei, et al.
Veröffentlicht: (2024)
von: Zhang, Xinwei, et al.
Veröffentlicht: (2024)
Networked Information Aggregation for Binary Classification
von: Bateni, MohammadHossein, et al.
Veröffentlicht: (2026)
von: Bateni, MohammadHossein, et al.
Veröffentlicht: (2026)
Optimal Approximation -- Smoothness Tradeoffs for Soft-Max Functions
von: Epasto, Alessandro, et al.
Veröffentlicht: (2020)
von: Epasto, Alessandro, et al.
Veröffentlicht: (2020)
Lattice: Learning to Efficiently Compress the Memory
von: Karami, Mahdi, et al.
Veröffentlicht: (2025)
von: Karami, Mahdi, et al.
Veröffentlicht: (2025)
Understanding the Role of Training Data in Test-Time Scaling
von: Javanmard, Adel, et al.
Veröffentlicht: (2025)
von: Javanmard, Adel, et al.
Veröffentlicht: (2025)
Theoretical Perspectives on Data Quality and Synergistic Effects in Pre- and Post-Training Reasoning Models
von: Javanmard, Adel, et al.
Veröffentlicht: (2026)
von: Javanmard, Adel, et al.
Veröffentlicht: (2026)
SYNAPSE-G: Bridging Large Language Models and Graph Learning for Rare Event Classification
von: Tavakkol, Sasan, et al.
Veröffentlicht: (2025)
von: Tavakkol, Sasan, et al.
Veröffentlicht: (2025)
Replicable Clustering
von: Esfandiari, Hossein, et al.
Veröffentlicht: (2023)
von: Esfandiari, Hossein, et al.
Veröffentlicht: (2023)
PriorBoost: An Adaptive Algorithm for Learning from Aggregate Responses
von: Javanmard, Adel, et al.
Veröffentlicht: (2024)
von: Javanmard, Adel, et al.
Veröffentlicht: (2024)
Titans: Learning to Memorize at Test Time
von: Behrouz, Ali, et al.
Veröffentlicht: (2024)
von: Behrouz, Ali, et al.
Veröffentlicht: (2024)
Differentially Private Graph Learning via Sensitivity-Bounded Personalized PageRank
von: Epasto, Alessandro, et al.
Veröffentlicht: (2022)
von: Epasto, Alessandro, et al.
Veröffentlicht: (2022)
Optimistic Rates for Learning from Label Proportions
von: Li, Gene, et al.
Veröffentlicht: (2024)
von: Li, Gene, et al.
Veröffentlicht: (2024)
Less is More: Convergence Benefits of Fewer Data Weight Updates over Longer Horizon
von: Das, Rudrajit, et al.
Veröffentlicht: (2026)
von: Das, Rudrajit, et al.
Veröffentlicht: (2026)
Self-Boost via Optimal Retraining: An Analysis via Approximate Message Passing
von: Javanmard, Adel, et al.
Veröffentlicht: (2025)
von: Javanmard, Adel, et al.
Veröffentlicht: (2025)
Learning Rate Schedules in the Presence of Distribution Shift
von: Fahrbach, Matthew, et al.
Veröffentlicht: (2023)
von: Fahrbach, Matthew, et al.
Veröffentlicht: (2023)
Sampling and Loss Weights in Multi-Domain Training
von: Salmani, Mahdi, et al.
Veröffentlicht: (2025)
von: Salmani, Mahdi, et al.
Veröffentlicht: (2025)
Nested Learning: The Illusion of Deep Learning Architectures
von: Behrouz, Ali, et al.
Veröffentlicht: (2025)
von: Behrouz, Ali, et al.
Veröffentlicht: (2025)
PiKE: Adaptive Data Mixing for Large-Scale Multi-Task Learning Under Low Gradient Conflicts
von: Li, Zeman, et al.
Veröffentlicht: (2025)
von: Li, Zeman, et al.
Veröffentlicht: (2025)
MS-SSM: A Multi-Scale State Space Model for Efficient Sequence Modeling
von: Karami, Mahdi, et al.
Veröffentlicht: (2025)
von: Karami, Mahdi, et al.
Veröffentlicht: (2025)
ECO: Quantized Training without Full-Precision Master Weights
von: Nikdan, Mahdi, et al.
Veröffentlicht: (2026)
von: Nikdan, Mahdi, et al.
Veröffentlicht: (2026)
Efficient Data Selection at Scale via Influence Distillation
von: Nikdan, Mahdi, et al.
Veröffentlicht: (2025)
von: Nikdan, Mahdi, et al.
Veröffentlicht: (2025)
SubGen: Token Generation in Sublinear Time and Memory
von: Zandieh, Amir, et al.
Veröffentlicht: (2024)
von: Zandieh, Amir, et al.
Veröffentlicht: (2024)
Secure Aggregation Meets Sparsification in Decentralized Learning
von: Biswas, Sayan, et al.
Veröffentlicht: (2024)
von: Biswas, Sayan, et al.
Veröffentlicht: (2024)
Perturb-and-Project: Differentially Private Similarities and Marginals
von: Cohen-Addad, Vincent, et al.
Veröffentlicht: (2024)
von: Cohen-Addad, Vincent, et al.
Veröffentlicht: (2024)
TurboQuant: Online Vector Quantization with Near-optimal Distortion Rate
von: Zandieh, Amir, et al.
Veröffentlicht: (2025)
von: Zandieh, Amir, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Sequential Attention for Feature Selection
von: Yasuda, Taisuke, et al.
Veröffentlicht: (2022) -
DeepCrossAttention: Supercharging Transformer Residual Connections
von: Heddes, Mike, et al.
Veröffentlicht: (2025) -
Budget Allocation for Unknown Value Functions in a Lipschitz Space
von: Bateni, MohammadHossein, et al.
Veröffentlicht: (2025) -
Data-Efficient Learning via Clustering-Based Sensitivity Sampling: Foundation Models and Beyond
von: Axiotis, Kyriakos, et al.
Veröffentlicht: (2024) -
Replicable Composition
von: Banihashem, Kiarash, et al.
Veröffentlicht: (2026)