Pruning is Optimal for Learning Sparse Features in High-Dimensions
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Vural, Nuri Mert, Erdogdu, Murat A. |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Learning quadratic neural networks in high dimensions: SGD dynamics and scaling laws
von: Arous, Gérard Ben, et al.
Veröffentlicht: (2025)
von: Arous, Gérard Ben, et al.
Veröffentlicht: (2025)
Robust Feature Learning for Multi-Index Models in High Dimensions
von: Mousavi-Hosseini, Alireza, et al.
Veröffentlicht: (2024)
von: Mousavi-Hosseini, Alireza, et al.
Veröffentlicht: (2024)
Learning to Recall with Transformers Beyond Orthogonal Embeddings
von: Vural, Nuri Mert, et al.
Veröffentlicht: (2026)
von: Vural, Nuri Mert, et al.
Veröffentlicht: (2026)
Post-Training with Policy Gradients: Optimality and the Base Model Barrier
von: Mousavi-Hosseini, Alireza, et al.
Veröffentlicht: (2026)
von: Mousavi-Hosseini, Alireza, et al.
Veröffentlicht: (2026)
On the Efficiency of ERM in Feature Learning
von: Hanchi, Ayoub El, et al.
Veröffentlicht: (2024)
von: Hanchi, Ayoub El, et al.
Veröffentlicht: (2024)
Optimal Excess Risk Bounds for Empirical Risk Minimization on $p$-Norm Linear Regression
von: Hanchi, Ayoub El, et al.
Veröffentlicht: (2023)
von: Hanchi, Ayoub El, et al.
Veröffentlicht: (2023)
Learning Multi-Index Models with Neural Networks via Mean-Field Langevin Dynamics
von: Mousavi-Hosseini, Alireza, et al.
Veröffentlicht: (2024)
von: Mousavi-Hosseini, Alireza, et al.
Veröffentlicht: (2024)
From Information to Generative Exponent: Learning Rate Induces Phase Transitions in SGD
von: Tsiolis, Konstantinos Christopher, et al.
Veröffentlicht: (2025)
von: Tsiolis, Konstantinos Christopher, et al.
Veröffentlicht: (2025)
A Geometric Analysis of PCA
von: Hanchi, Ayoub El, et al.
Veröffentlicht: (2025)
von: Hanchi, Ayoub El, et al.
Veröffentlicht: (2025)
When Do Transformers Outperform Feedforward and Recurrent Networks? A Statistical Perspective
von: Mousavi-Hosseini, Alireza, et al.
Veröffentlicht: (2025)
von: Mousavi-Hosseini, Alireza, et al.
Veröffentlicht: (2025)
Minimax Linear Regression under the Quantile Risk
von: Hanchi, Ayoub El, et al.
Veröffentlicht: (2024)
von: Hanchi, Ayoub El, et al.
Veröffentlicht: (2024)
Beyond Labeling Oracles: What does it mean to steal ML models?
von: Shafran, Avital, et al.
Veröffentlicht: (2023)
von: Shafran, Avital, et al.
Veröffentlicht: (2023)
A Separation in Heavy-Tailed Sampling: Gaussian vs. Stable Oracles for Proximal Samplers
von: He, Ye, et al.
Veröffentlicht: (2024)
von: He, Ye, et al.
Veröffentlicht: (2024)
Sampling from the Mean-Field Stationary Distribution
von: Kook, Yunbum, et al.
Veröffentlicht: (2024)
von: Kook, Yunbum, et al.
Veröffentlicht: (2024)
Analysis of Langevin Monte Carlo from Poincaré to Log-Sobolev
von: Chewi, Sinho, et al.
Veröffentlicht: (2021)
von: Chewi, Sinho, et al.
Veröffentlicht: (2021)
Negotiated Representations to Prevent Overfitting in Machine Learning Applications
von: Korhan, Nuri, et al.
Veröffentlicht: (2023)
von: Korhan, Nuri, et al.
Veröffentlicht: (2023)
Unveiling Hidden Convexity in Deep Learning: a Sparse Signal Processing Perspective
von: Zeger, Emi, et al.
Veröffentlicht: (2026)
von: Zeger, Emi, et al.
Veröffentlicht: (2026)
Optimal Embedding Dimension for Sparse Subspace Embeddings
von: Chenakkod, Shabarish, et al.
Veröffentlicht: (2023)
von: Chenakkod, Shabarish, et al.
Veröffentlicht: (2023)
Sparse Additive Model Pruning for Order-Based Causal Structure Learning
von: Kanamori, Kentaro, et al.
Veröffentlicht: (2026)
von: Kanamori, Kentaro, et al.
Veröffentlicht: (2026)
Dynamics of Stochastic Momentum with Sparse Updates in High Dimensions
von: Everett, Katie, et al.
Veröffentlicht: (2026)
von: Everett, Katie, et al.
Veröffentlicht: (2026)
ALPS: Improved Optimization for Highly Sparse One-Shot Pruning for Large Language Models
von: Meng, Xiang, et al.
Veröffentlicht: (2024)
von: Meng, Xiang, et al.
Veröffentlicht: (2024)
RoBERTurk: Adjusting RoBERTa for Turkish
von: Tas, Nuri
Veröffentlicht: (2024)
von: Tas, Nuri
Veröffentlicht: (2024)
Optimal Sets and Solution Paths of ReLU Networks
von: Mishkin, Aaron, et al.
Veröffentlicht: (2023)
von: Mishkin, Aaron, et al.
Veröffentlicht: (2023)
Learning Flock: Enhancing Sets of Particles for Multi~Sub-State Particle Filtering with Neural Augmentation
von: Nuri, Itai, et al.
Veröffentlicht: (2024)
von: Nuri, Itai, et al.
Veröffentlicht: (2024)
A Novel Approach for Intrinsic Dimension Estimation
von: Özçoban, Kadir, et al.
Veröffentlicht: (2025)
von: Özçoban, Kadir, et al.
Veröffentlicht: (2025)
A Unified Analysis of Generalization and Sample Complexity for Semi-Supervised Domain Adaptation
von: Vural, Elif, et al.
Veröffentlicht: (2025)
von: Vural, Elif, et al.
Veröffentlicht: (2025)
Locally Stationary Graph Processes
von: Canbolat, Abdullah, et al.
Veröffentlicht: (2023)
von: Canbolat, Abdullah, et al.
Veröffentlicht: (2023)
A Recovery Guarantee for Sparse Neural Networks
von: Fridovich-Keil, Sara, et al.
Veröffentlicht: (2025)
von: Fridovich-Keil, Sara, et al.
Veröffentlicht: (2025)
Graph Signal Inference by Learning Narrowband Spectral Kernels
von: Kar, Osman Furkan, et al.
Veröffentlicht: (2025)
von: Kar, Osman Furkan, et al.
Veröffentlicht: (2025)
Optimal Attention Temperature Improves the Robustness of In-Context Learning under Distribution Shift in High Dimensions
von: Demir, Samet, et al.
Veröffentlicht: (2025)
von: Demir, Samet, et al.
Veröffentlicht: (2025)
Distributional Training Data Attribution: What do Influence Functions Sample?
von: Mlodozeniec, Bruno, et al.
Veröffentlicht: (2025)
von: Mlodozeniec, Bruno, et al.
Veröffentlicht: (2025)
Sparse Modelling for Feature Learning in High Dimensional Data
von: Neelam, Harish, et al.
Veröffentlicht: (2024)
von: Neelam, Harish, et al.
Veröffentlicht: (2024)
Pruning for Sparse Diffusion Models based on Gradient Flow
von: Wan, Ben, et al.
Veröffentlicht: (2025)
von: Wan, Ben, et al.
Veröffentlicht: (2025)
SLOPE: Search with Learned Optimal Pruning-based Expansion
von: Bokan, Davor, et al.
Veröffentlicht: (2024)
von: Bokan, Davor, et al.
Veröffentlicht: (2024)
Semantic Optimal Transport for Sparse Autoencoder Feature Matching and Circuit Compression
von: Cao, Tue M., et al.
Veröffentlicht: (2026)
von: Cao, Tue M., et al.
Veröffentlicht: (2026)
Pruned Adaptation Modules: A Simple yet Strong Baseline for Continual Foundation Models
von: Yildirim, Elif Ceren Gok, et al.
Veröffentlicht: (2026)
von: Yildirim, Elif Ceren Gok, et al.
Veröffentlicht: (2026)
How Pruning Reshapes Features: Sparse Autoencoder Analysis of Weight-Pruned Language Models
von: Borobia, Hector, et al.
Veröffentlicht: (2026)
von: Borobia, Hector, et al.
Veröffentlicht: (2026)
Optimal Scaling for the Proximal Langevin Algorithm in High Dimensions
von: Pillai, Natesh S.
Veröffentlicht: (2022)
von: Pillai, Natesh S.
Veröffentlicht: (2022)
EvoESAP: Non-Uniform Expert Pruning for Sparse MoE
von: Liu, Zongfang, et al.
Veröffentlicht: (2026)
von: Liu, Zongfang, et al.
Veröffentlicht: (2026)
Optimal Parameter and Neuron Pruning for Out-of-Distribution Detection
von: Chen, Chao, et al.
Veröffentlicht: (2024)
von: Chen, Chao, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Learning quadratic neural networks in high dimensions: SGD dynamics and scaling laws
von: Arous, Gérard Ben, et al.
Veröffentlicht: (2025) -
Robust Feature Learning for Multi-Index Models in High Dimensions
von: Mousavi-Hosseini, Alireza, et al.
Veröffentlicht: (2024) -
Learning to Recall with Transformers Beyond Orthogonal Embeddings
von: Vural, Nuri Mert, et al.
Veröffentlicht: (2026) -
Post-Training with Policy Gradients: Optimality and the Base Model Barrier
von: Mousavi-Hosseini, Alireza, et al.
Veröffentlicht: (2026) -
On the Efficiency of ERM in Feature Learning
von: Hanchi, Ayoub El, et al.
Veröffentlicht: (2024)