The Unseen Frontier: Pushing the Limits of LLM Sparsity with Surrogate-Free ADMM
Fuente:
arXiv
Salvato in:
| Autori principali: | Lee, Kwanhee, Jang, Hyeondo, Lee, Dongyeop, Alistarh, Dan, Lee, Namhoon |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
SAFE: Finding Sparse and Flat Minima to Improve Pruning
di: Lee, Dongyeop, et al.
Pubblicazione: (2025)
di: Lee, Dongyeop, et al.
Pubblicazione: (2025)
SASSHA: Sharpness-aware Adaptive Second-order Optimization with Stable Hessian Approximation
di: Shin, Dahun, et al.
Pubblicazione: (2025)
di: Shin, Dahun, et al.
Pubblicazione: (2025)
Wasserstein Distances, Neuronal Entanglement, and Sparsity
di: Sawmya, Shashata, et al.
Pubblicazione: (2024)
di: Sawmya, Shashata, et al.
Pubblicazione: (2024)
PLADIS: Pushing the Limits of Attention in Diffusion Models at Inference Time by Leveraging Sparsity
di: Kim, Kwanyoung, et al.
Pubblicazione: (2025)
di: Kim, Kwanyoung, et al.
Pubblicazione: (2025)
An Analysis of Concept Bottleneck Models: Measuring, Understanding, and Mitigating the Impact of Noisy Annotations
di: Park, Seonghwan, et al.
Pubblicazione: (2025)
di: Park, Seonghwan, et al.
Pubblicazione: (2025)
Fast Forward: Accelerating LLM Prefill with Predictive FFN Sparsity
di: Gautam, Aayush, et al.
Pubblicazione: (2026)
di: Gautam, Aayush, et al.
Pubblicazione: (2026)
Sparsity and Out-of-Distribution Generalization
di: Aaronson, Scott, et al.
Pubblicazione: (2026)
di: Aaronson, Scott, et al.
Pubblicazione: (2026)
Mitigating Staleness in Asynchronous Pipeline Parallelism via Basis Rotation
di: Jung, Hyunji, et al.
Pubblicazione: (2026)
di: Jung, Hyunji, et al.
Pubblicazione: (2026)
Critical Influence of Overparameterization on Sharpness-aware Minimization
di: Shin, Sungbin, et al.
Pubblicazione: (2023)
di: Shin, Sungbin, et al.
Pubblicazione: (2023)
MatryoshkaLoRA: Learning Accurate Hierarchical Low-Rank Representations for LLM Fine-Tuning
di: Modoranu, Ionut-Vlad, et al.
Pubblicazione: (2026)
di: Modoranu, Ionut-Vlad, et al.
Pubblicazione: (2026)
"Give Me BF16 or Give Me Death"? Accuracy-Performance Trade-Offs in LLM Quantization
di: Kurtic, Eldar, et al.
Pubblicazione: (2024)
di: Kurtic, Eldar, et al.
Pubblicazione: (2024)
Generalized Gaussian Temporal Difference Error for Uncertainty-aware Reinforcement Learning
di: Kim, Seyeon, et al.
Pubblicazione: (2024)
di: Kim, Seyeon, et al.
Pubblicazione: (2024)
Seen-to-Scene: Keep the Seen, Generate the Unseen for Video Outpainting
di: Jeon, Inseok, et al.
Pubblicazione: (2026)
di: Jeon, Inseok, et al.
Pubblicazione: (2026)
Conformal Constrained Policy Optimization for Cost-Effective LLM Agents
di: Si, Wenwen, et al.
Pubblicazione: (2025)
di: Si, Wenwen, et al.
Pubblicazione: (2025)
Weight Concentration Regularization for Improving Pruning Robustness Under High Sparsity
di: Yun, Vincent-Daniel, et al.
Pubblicazione: (2025)
di: Yun, Vincent-Daniel, et al.
Pubblicazione: (2025)
ChemFixer: Correcting Invalid Molecules to Unlock Previously Unseen Chemical Space
di: Park, Jun-Hyoung, et al.
Pubblicazione: (2025)
di: Park, Jun-Hyoung, et al.
Pubblicazione: (2025)
Toward Efficient Permutation for Hierarchical N:M Sparsity on GPUs
di: Yu, Seungmin, et al.
Pubblicazione: (2024)
di: Yu, Seungmin, et al.
Pubblicazione: (2024)
Pushing Forward Pareto Frontiers of Proactive Agents with Behavioral Agentic Optimization
di: Yao, Yihang, et al.
Pubblicazione: (2026)
di: Yao, Yihang, et al.
Pubblicazione: (2026)
BiLLM: Pushing the Limit of Post-Training Quantization for LLMs
di: Huang, Wei, et al.
Pubblicazione: (2024)
di: Huang, Wei, et al.
Pubblicazione: (2024)
Meta-Controller: Few-Shot Imitation of Unseen Embodiments and Tasks in Continuous Control
di: Cho, Seongwoong, et al.
Pubblicazione: (2024)
di: Cho, Seongwoong, et al.
Pubblicazione: (2024)
Is Prompt Selection Necessary for Task-Free Online Continual Learning?
di: Park, Seoyoung, et al.
Pubblicazione: (2026)
di: Park, Seoyoung, et al.
Pubblicazione: (2026)
To Compress or Not? Pushing the Frontier of Lossless GenAI Model Weights Compression with Exponent Concentration
di: Yang, Zeyu, et al.
Pubblicazione: (2025)
di: Yang, Zeyu, et al.
Pubblicazione: (2025)
Pushing the Limits of Block Rotations in Post-Training Quantization
di: Sanjeet, Sai, et al.
Pubblicazione: (2026)
di: Sanjeet, Sai, et al.
Pubblicazione: (2026)
Polar Sparsity: High Throughput Batched LLM Inferencing with Scalable Contextual Sparsity
di: Shrestha, Susav, et al.
Pubblicazione: (2025)
di: Shrestha, Susav, et al.
Pubblicazione: (2025)
DIVER-1: Scaling Intracranial EEG Foundation Models for Transferable Representations
di: Han, Danny Dongyeop, et al.
Pubblicazione: (2025)
di: Han, Danny Dongyeop, et al.
Pubblicazione: (2025)
Generative AI Policies under the Microscope: How CS Conferences Are Navigating the New Frontier in Scholarly Writing
di: Nahar, Mahjabin, et al.
Pubblicazione: (2024)
di: Nahar, Mahjabin, et al.
Pubblicazione: (2024)
The Amazing Agent Race: Strong Tool Users, Weak Navigators
di: Kim, Zae Myung, et al.
Pubblicazione: (2026)
di: Kim, Zae Myung, et al.
Pubblicazione: (2026)
Accelerating LMO-Based Optimization via Implicit Gradient Transport
di: Jang, Won-Jun, et al.
Pubblicazione: (2026)
di: Jang, Won-Jun, et al.
Pubblicazione: (2026)
A Simple and Scalable Representation for Graph Generation
di: Jang, Yunhui, et al.
Pubblicazione: (2023)
di: Jang, Yunhui, et al.
Pubblicazione: (2023)
Token-Efficient RL for LLM Reasoning
di: Lee, Alan, et al.
Pubblicazione: (2025)
di: Lee, Alan, et al.
Pubblicazione: (2025)
MegaScience: Pushing the Frontiers of Post-Training Datasets for Science Reasoning
di: Fan, Run-Ze, et al.
Pubblicazione: (2025)
di: Fan, Run-Ze, et al.
Pubblicazione: (2025)
Pushing the Limits of Low-Bit Optimizers: A Focus on EMA Dynamics
di: Xu, Cong, et al.
Pubblicazione: (2025)
di: Xu, Cong, et al.
Pubblicazione: (2025)
ECO: Quantized Training without Full-Precision Master Weights
di: Nikdan, Mahdi, et al.
Pubblicazione: (2026)
di: Nikdan, Mahdi, et al.
Pubblicazione: (2026)
RoSA: Accurate Parameter-Efficient Fine-Tuning via Robust Adaptation
di: Nikdan, Mahdi, et al.
Pubblicazione: (2024)
di: Nikdan, Mahdi, et al.
Pubblicazione: (2024)
Connecting Federated ADMM to Bayes
di: Swaroop, Siddharth, et al.
Pubblicazione: (2025)
di: Swaroop, Siddharth, et al.
Pubblicazione: (2025)
DIVER-0 : A Fully Channel Equivariant EEG Foundation Model
di: Han, Danny Dongyeop, et al.
Pubblicazione: (2025)
di: Han, Danny Dongyeop, et al.
Pubblicazione: (2025)
DAFA: Distance-Aware Fair Adversarial Training
di: Lee, Hyungyu, et al.
Pubblicazione: (2024)
di: Lee, Hyungyu, et al.
Pubblicazione: (2024)
Stochastic Parameter Decomposition
di: Bushnaq, Lucius, et al.
Pubblicazione: (2025)
di: Bushnaq, Lucius, et al.
Pubblicazione: (2025)
RaaS: Reasoning-Aware Attention Sparsity for Efficient LLM Reasoning
di: Hu, Junhao, et al.
Pubblicazione: (2025)
di: Hu, Junhao, et al.
Pubblicazione: (2025)
To FP8 and Back Again: Quantifying Reduced Precision Effects on LLM Training Stability
di: Lee, Joonhyung, et al.
Pubblicazione: (2024)
di: Lee, Joonhyung, et al.
Pubblicazione: (2024)
Documenti analoghi
-
SAFE: Finding Sparse and Flat Minima to Improve Pruning
di: Lee, Dongyeop, et al.
Pubblicazione: (2025) -
SASSHA: Sharpness-aware Adaptive Second-order Optimization with Stable Hessian Approximation
di: Shin, Dahun, et al.
Pubblicazione: (2025) -
Wasserstein Distances, Neuronal Entanglement, and Sparsity
di: Sawmya, Shashata, et al.
Pubblicazione: (2024) -
PLADIS: Pushing the Limits of Attention in Diffusion Models at Inference Time by Leveraging Sparsity
di: Kim, Kwanyoung, et al.
Pubblicazione: (2025) -
An Analysis of Concept Bottleneck Models: Measuring, Understanding, and Mitigating the Impact of Noisy Annotations
di: Park, Seonghwan, et al.
Pubblicazione: (2025)