Even Sparser Graph Transformers
Fuente:
arXiv
Saved in:
| Main Authors: | Shirzad, Hamed, Lin, Honghao, Venkatachalam, Balaji, Velingker, Ameya, Woodruff, David, Sutherland, Danica |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
A Theory for Compressibility of Graph Transformers for Transductive Learning
by: Shirzad, Hamed, et al.
Published: (2024)
by: Shirzad, Hamed, et al.
Published: (2024)
SeedER: Seed-and-Expand Retrieval from Knowledge Graphs
by: Shirzad, Hamed, et al.
Published: (2026)
by: Shirzad, Hamed, et al.
Published: (2026)
Locality-Aware Graph-Rewiring in GNNs
by: Barbero, Federico, et al.
Published: (2023)
by: Barbero, Federico, et al.
Published: (2023)
Optimal Sketching for Residual Error Estimation for Matrix and Vector Norms
by: Li, Yi, et al.
Published: (2024)
by: Li, Yi, et al.
Published: (2024)
Sparser, Faster, Lighter Transformer Language Models
by: Cetin, Edoardo, et al.
Published: (2026)
by: Cetin, Edoardo, et al.
Published: (2026)
Weisfeiler-Leman at the margin: When more expressivity matters
by: Franks, Billy J., et al.
Published: (2024)
by: Franks, Billy J., et al.
Published: (2024)
Sparser is Faster and Less is More: Efficient Sparse Attention for Long-Range Transformers
by: Lou, Chao, et al.
Published: (2024)
by: Lou, Chao, et al.
Published: (2024)
Learning the Positions in CountSketch
by: Li, Yi, et al.
Published: (2023)
by: Li, Yi, et al.
Published: (2023)
Understanding Simplicity Bias towards Compositional Mappings via Learning Dynamics
by: Ren, Yi, et al.
Published: (2024)
by: Ren, Yi, et al.
Published: (2024)
Generalized Coverage for More Robust Low-Budget Active Learning
by: Bae, Wonho, et al.
Published: (2024)
by: Bae, Wonho, et al.
Published: (2024)
Learning Representations for Independence Testing
by: Xu, Nathaniel, et al.
Published: (2024)
by: Xu, Nathaniel, et al.
Published: (2024)
Efficient kernelized bandit algorithms via exploration distributions
by: Hu, Bingshan, et al.
Published: (2025)
by: Hu, Bingshan, et al.
Published: (2025)
Exploring Active Learning in Meta-Learning: Enhancing Context Set Labeling
by: Bae, Wonho, et al.
Published: (2023)
by: Bae, Wonho, et al.
Published: (2023)
Learning Dynamics of LLM Finetuning
by: Ren, Yi, et al.
Published: (2024)
by: Ren, Yi, et al.
Published: (2024)
In Transformer We Trust? A Perspective on Transformer Architecture Failure Modes
by: Mondal, Trishit, et al.
Published: (2026)
by: Mondal, Trishit, et al.
Published: (2026)
Uncertainty Herding: One Active Learning Method for All Label Budgets
by: Bae, Wonho, et al.
Published: (2024)
by: Bae, Wonho, et al.
Published: (2024)
Sparser, Better, Faster, Stronger: Sparsity Detection for Efficient Automatic Differentiation
by: Hill, Adrian, et al.
Published: (2025)
by: Hill, Adrian, et al.
Published: (2025)
Efficient Attention via Pre-Scoring: Prioritizing Informative Keys in Transformers
by: Li, Zhexiang, et al.
Published: (2025)
by: Li, Zhexiang, et al.
Published: (2025)
Integrating AI and Ensemble Forecasting: Explainable Materials Planning with Scorecards and Trend Insights for a Large-Scale Manufacturer
by: Venkatachalam, Saravanan
Published: (2025)
by: Venkatachalam, Saravanan
Published: (2025)
Maximum Mean Discrepancy with Unequal Sample Sizes via Generalized U-Statistics
by: Wei, Aaron, et al.
Published: (2025)
by: Wei, Aaron, et al.
Published: (2025)
Projection-Free CNN Pruning via Frank-Wolfe with Momentum: Sparser Models with Less Pretraining
by: Shili, Hamza ElMokhtar, et al.
Published: (2025)
by: Shili, Hamza ElMokhtar, et al.
Published: (2025)
Sparser, Better, Deeper, Stronger: Improving Sparse Training with Exact Orthogonal Initialization
by: Nowak, Aleksandra Irena, et al.
Published: (2024)
by: Nowak, Aleksandra Irena, et al.
Published: (2024)
Decentralized Personalized Federated Learning based on a Conditional Sparse-to-Sparser Scheme
by: Long, Qianyu, et al.
Published: (2024)
by: Long, Qianyu, et al.
Published: (2024)
Why Do You Grok? A Theoretical Analysis of Grokking Modular Addition
by: Mohamadi, Mohamad Amin, et al.
Published: (2024)
by: Mohamadi, Mohamad Amin, et al.
Published: (2024)
Practical Kernel Tests of Conditional Independence
by: Pogodin, Roman, et al.
Published: (2024)
by: Pogodin, Roman, et al.
Published: (2024)
The Road to Generalizable Neuro-Symbolic Learning Should be Paved with Foundation Models
by: Stein, Adam, et al.
Published: (2025)
by: Stein, Adam, et al.
Published: (2025)
MAPPING: Debiasing Graph Neural Networks for Fair Node Classification with Limited Sensitive Information Leakage
by: Song, Ying, et al.
Published: (2024)
by: Song, Ying, et al.
Published: (2024)
Once Upon an Input: Reasoning via Per-Instance Program Synthesis
by: Stein, Adam, et al.
Published: (2025)
by: Stein, Adam, et al.
Published: (2025)
GraphToxin: Reconstructing Full Unlearned Graphs from Graph Unlearning
by: Song, Ying, et al.
Published: (2025)
by: Song, Ying, et al.
Published: (2025)
Search-contempt: a hybrid MCTS algorithm for training AlphaZero-like engines with better computational efficiency
by: Joshi, Ameya
Published: (2025)
by: Joshi, Ameya
Published: (2025)
Can Transformers Smell Like Humans?
by: Taleb, Farzaneh, et al.
Published: (2024)
by: Taleb, Farzaneh, et al.
Published: (2024)
Query-Efficient Locally Private Hypothesis Selection via the Scheffe Graph
by: Kamath, Gautam, et al.
Published: (2025)
by: Kamath, Gautam, et al.
Published: (2025)
LLM Unlearning using Gradient Ratio-Based Influence Estimation and Noise Injection
by: Anjarlekar, Ameya, et al.
Published: (2025)
by: Anjarlekar, Ameya, et al.
Published: (2025)
Ridge Leverage Score Sampling for $\ell_p$ Subspace Approximation
by: Woodruff, David P., et al.
Published: (2024)
by: Woodruff, David P., et al.
Published: (2024)
Reweighted Solutions for Weighted Low Rank Approximation
by: Woodruff, David P., et al.
Published: (2024)
by: Woodruff, David P., et al.
Published: (2024)
Coresets for Multiple $\ell_p$ Regression
by: Woodruff, David P., et al.
Published: (2024)
by: Woodruff, David P., et al.
Published: (2024)
Better Bounds for the Distributed Experts Problem
by: Woodruff, David P., et al.
Published: (2026)
by: Woodruff, David P., et al.
Published: (2026)
Sharper Bounds for $\ell_p$ Sensitivity Sampling
by: Woodruff, David P., et al.
Published: (2023)
by: Woodruff, David P., et al.
Published: (2023)
John Ellipsoids via Lazy Updates
by: Woodruff, David P., et al.
Published: (2025)
by: Woodruff, David P., et al.
Published: (2025)
On the Hardness of Conditional Independence Testing In Practice
by: He, Zheng, et al.
Published: (2025)
by: He, Zheng, et al.
Published: (2025)
Similar Items
-
A Theory for Compressibility of Graph Transformers for Transductive Learning
by: Shirzad, Hamed, et al.
Published: (2024) -
SeedER: Seed-and-Expand Retrieval from Knowledge Graphs
by: Shirzad, Hamed, et al.
Published: (2026) -
Locality-Aware Graph-Rewiring in GNNs
by: Barbero, Federico, et al.
Published: (2023) -
Optimal Sketching for Residual Error Estimation for Matrix and Vector Norms
by: Li, Yi, et al.
Published: (2024) -
Sparser, Faster, Lighter Transformer Language Models
by: Cetin, Edoardo, et al.
Published: (2026)