A Three-regime Model of Network Pruning
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Zhou, Yefan, Yang, Yaoqing, Chang, Arin, Mahoney, Michael W. |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2023
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
AlphaPruning: Using Heavy-Tailed Self Regularization Theory for Improved Layer-wise Pruning of Large Language Models
von: Lu, Haiquan, et al.
Veröffentlicht: (2024)
von: Lu, Haiquan, et al.
Veröffentlicht: (2024)
A Model Zoo on Phase Transitions in Neural Networks
von: Schürholt, Konstantin, et al.
Veröffentlicht: (2025)
von: Schürholt, Konstantin, et al.
Veröffentlicht: (2025)
Sharpness-diversity tradeoff: improving flat ensembles with SharpBalance
von: Lu, Haiquan, et al.
Veröffentlicht: (2024)
von: Lu, Haiquan, et al.
Veröffentlicht: (2024)
MD tree: a model-diagnostic tree grown on loss landscape
von: Zhou, Yefan, et al.
Veröffentlicht: (2024)
von: Zhou, Yefan, et al.
Veröffentlicht: (2024)
Model Balancing Helps Low-data Training and Fine-tuning
von: Liu, Zihang, et al.
Veröffentlicht: (2024)
von: Liu, Zihang, et al.
Veröffentlicht: (2024)
Learning to Discover Iterative Spectral Algorithms
von: Liu, Zihang, et al.
Veröffentlicht: (2026)
von: Liu, Zihang, et al.
Veröffentlicht: (2026)
RL4RLA: Teaching ML to Discover Randomized Linear Algebra Algorithms Through Curriculum Design and Graph-Based Search
von: Xiong, Jinglong, et al.
Veröffentlicht: (2026)
von: Xiong, Jinglong, et al.
Veröffentlicht: (2026)
Unveiling Multi-regime Patterns in SciML: Distinct Failure Modes and Regime-specific Optimization
von: Wang, Yuxin, et al.
Veröffentlicht: (2026)
von: Wang, Yuxin, et al.
Veröffentlicht: (2026)
Mitigating Memorization In Language Models
von: Sakarvadia, Mansi, et al.
Veröffentlicht: (2024)
von: Sakarvadia, Mansi, et al.
Veröffentlicht: (2024)
Random Matrix Theory for Deep Learning: Beyond Eigenvalues of Linear Models
von: Liao, Zhenyu, et al.
Veröffentlicht: (2025)
von: Liao, Zhenyu, et al.
Veröffentlicht: (2025)
The False Promise of Zero-Shot Super-Resolution in Machine-Learned Operators
von: Sakarvadia, Mansi, et al.
Veröffentlicht: (2025)
von: Sakarvadia, Mansi, et al.
Veröffentlicht: (2025)
Evaluating Loss Landscapes from a Topology Perspective
von: Xie, Tiankai, et al.
Veröffentlicht: (2024)
von: Xie, Tiankai, et al.
Veröffentlicht: (2024)
Eigenspectrum Analysis of Neural Networks without Aspect Ratio Bias
von: Hu, Yuanzhe, et al.
Veröffentlicht: (2025)
von: Hu, Yuanzhe, et al.
Veröffentlicht: (2025)
Balancing Learning Rates Across Layers: Exact Two-Step Dynamics and Optimal Scaling in Linear Neural Networks
von: Pang, Tianyu, et al.
Veröffentlicht: (2026)
von: Pang, Tianyu, et al.
Veröffentlicht: (2026)
KCES: Training-Free Defense for Robust Graph Neural Networks via Kernel Complexity
von: Jia, Yaning, et al.
Veröffentlicht: (2025)
von: Jia, Yaning, et al.
Veröffentlicht: (2025)
Models of Heavy-Tailed Mechanistic Universality
von: Hodgkinson, Liam, et al.
Veröffentlicht: (2025)
von: Hodgkinson, Liam, et al.
Veröffentlicht: (2025)
Spectral Insights into Data-Oblivious Critical Layers in Large Language Models
von: Liu, Xuyuan, et al.
Veröffentlicht: (2025)
von: Liu, Xuyuan, et al.
Veröffentlicht: (2025)
HOPE for a Robust Parameterization of Long-memory State Space Models
von: Yu, Annan, et al.
Veröffentlicht: (2024)
von: Yu, Annan, et al.
Veröffentlicht: (2024)
Visualizing Loss Functions as Topological Landscape Profiles
von: Geniesse, Caleb, et al.
Veröffentlicht: (2024)
von: Geniesse, Caleb, et al.
Veröffentlicht: (2024)
LossLens: Diagnostics for Machine Learning through Loss Landscape Visual Analytics
von: Xie, Tiankai, et al.
Veröffentlicht: (2024)
von: Xie, Tiankai, et al.
Veröffentlicht: (2024)
Recent and Upcoming Developments in Randomized Numerical Linear Algebra for Machine Learning
von: Dereziński, Michał, et al.
Veröffentlicht: (2024)
von: Dereziński, Michał, et al.
Veröffentlicht: (2024)
Statistical Inference of Constrained Stochastic Optimization via Sketched Sequential Quadratic Programming
von: Na, Sen, et al.
Veröffentlicht: (2022)
von: Na, Sen, et al.
Veröffentlicht: (2022)
A PRISMA Driven Systematic Review of Publicly Available Datasets for Benchmark and Model Developments for Industrial Defect Detection
von: Akbas, Can, et al.
Veröffentlicht: (2024)
von: Akbas, Can, et al.
Veröffentlicht: (2024)
From Spikes to Heavy Tails: Unveiling the Spectral Evolution of Neural Networks
von: Kothapalli, Vignesh, et al.
Veröffentlicht: (2024)
von: Kothapalli, Vignesh, et al.
Veröffentlicht: (2024)
Towards Scalable and Versatile Weight Space Learning
von: Schürholt, Konstantin, et al.
Veröffentlicht: (2024)
von: Schürholt, Konstantin, et al.
Veröffentlicht: (2024)
Gated Recurrent Neural Networks with Weighted Time-Delay Feedback
von: Erichson, N. Benjamin, et al.
Veröffentlicht: (2022)
von: Erichson, N. Benjamin, et al.
Veröffentlicht: (2022)
Depth, Not Data: An Analysis of Hessian Spectral Bifurcation
von: Deng, Shenyang, et al.
Veröffentlicht: (2026)
von: Deng, Shenyang, et al.
Veröffentlicht: (2026)
Fast and Compact Tsetlin Machine Inference on CPUs Using Instruction-Level Optimization
von: Zeng, Yefan, et al.
Veröffentlicht: (2025)
von: Zeng, Yefan, et al.
Veröffentlicht: (2025)
Variation in Verification: Understanding Verification Dynamics in Large Language Models
von: Zhou, Yefan, et al.
Veröffentlicht: (2025)
von: Zhou, Yefan, et al.
Veröffentlicht: (2025)
Landscaper: Understanding Loss Landscapes Through Multi-Dimensional Topological Analysis
von: Chen, Jiaqing, et al.
Veröffentlicht: (2026)
von: Chen, Jiaqing, et al.
Veröffentlicht: (2026)
Suspicious Alignment of SGD: A Fine-Grained Step Size Condition Analysis
von: Deng, Shenyang, et al.
Veröffentlicht: (2026)
von: Deng, Shenyang, et al.
Veröffentlicht: (2026)
Iterative Refinement Neural Operators are Learned Fixed-Point Solvers: A Principled Approach to Spectral Bias Mitigation
von: Liu, Xiaotian, et al.
Veröffentlicht: (2026)
von: Liu, Xiaotian, et al.
Veröffentlicht: (2026)
Balanced Edge Pruning for Graph Anomaly Detection with Noisy Labels
von: Wang, Zhu, et al.
Veröffentlicht: (2024)
von: Wang, Zhu, et al.
Veröffentlicht: (2024)
SwiftPrune: Hessian-Free Weight Pruning for Large Language Models
von: Kang, Yuhan, et al.
Veröffentlicht: (2025)
von: Kang, Yuhan, et al.
Veröffentlicht: (2025)
Generative Modeling of Regular and Irregular Time Series Data via Koopman VAEs
von: Naiman, Ilan, et al.
Veröffentlicht: (2023)
von: Naiman, Ilan, et al.
Veröffentlicht: (2023)
The Interpolating Information Criterion for Overparameterized Models
von: Hodgkinson, Liam, et al.
Veröffentlicht: (2023)
von: Hodgkinson, Liam, et al.
Veröffentlicht: (2023)
Spectral Pruning for Recurrent Neural Networks
von: Furuya, Takashi, et al.
Veröffentlicht: (2021)
von: Furuya, Takashi, et al.
Veröffentlicht: (2021)
Tuning Frequency Bias of State Space Models
von: Yu, Annan, et al.
Veröffentlicht: (2024)
von: Yu, Annan, et al.
Veröffentlicht: (2024)
QPruner: Probabilistic Decision Quantization for Structured Pruning in Large Language Models
von: Zhou, Changhai, et al.
Veröffentlicht: (2024)
von: Zhou, Changhai, et al.
Veröffentlicht: (2024)
How many classifiers do we need?
von: Kim, Hyunsuk, et al.
Veröffentlicht: (2024)
von: Kim, Hyunsuk, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
AlphaPruning: Using Heavy-Tailed Self Regularization Theory for Improved Layer-wise Pruning of Large Language Models
von: Lu, Haiquan, et al.
Veröffentlicht: (2024) -
A Model Zoo on Phase Transitions in Neural Networks
von: Schürholt, Konstantin, et al.
Veröffentlicht: (2025) -
Sharpness-diversity tradeoff: improving flat ensembles with SharpBalance
von: Lu, Haiquan, et al.
Veröffentlicht: (2024) -
MD tree: a model-diagnostic tree grown on loss landscape
von: Zhou, Yefan, et al.
Veröffentlicht: (2024) -
Model Balancing Helps Low-data Training and Fine-tuning
von: Liu, Zihang, et al.
Veröffentlicht: (2024)