Salvato in:
| Autori principali: | Yin, Lu, Jaiswal, Ajay, Liu, Shiwei, Kundu, Souvik, Wang, Zhangyang |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2023
|
| Soggetti: | |
| Accesso online: | https://arxiv.org/abs/2310.02277 |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
AlphaDecay: Module-wise Weight Decay for Heavy-Tailed Balancing in LLMs
di: He, Di, et al.
Pubblicazione: (2025)
di: He, Di, et al.
Pubblicazione: (2025)
From Low Rank Gradient Subspace Stabilization to Low-Rank Weights: Observations, Theories, and Applications
di: Jaiswal, Ajay, et al.
Pubblicazione: (2024)
di: Jaiswal, Ajay, et al.
Pubblicazione: (2024)
Q-GaLore: Quantized GaLore with INT4 Projection and Layer-Adaptive Low-Rank Gradients
di: Zhang, Zhenyu, et al.
Pubblicazione: (2024)
di: Zhang, Zhenyu, et al.
Pubblicazione: (2024)
Outlier Weighed Layerwise Sparsity (OWL): A Missing Secret Sauce for Pruning LLMs to High Sparsity
di: Yin, Lu, et al.
Pubblicazione: (2023)
di: Yin, Lu, et al.
Pubblicazione: (2023)
Is C4 Dataset Optimal for Pruning? An Investigation of Calibration Data for LLM Pruning
di: Bandari, Abhinav, et al.
Pubblicazione: (2024)
di: Bandari, Abhinav, et al.
Pubblicazione: (2024)
Condense, Don't Just Prune: Enhancing Efficiency and Performance in MoE Layer Pruning
di: Cao, Mingyu, et al.
Pubblicazione: (2024)
di: Cao, Mingyu, et al.
Pubblicazione: (2024)
On the Hardness of Junking LLMs
di: Rando, Marco, et al.
Pubblicazione: (2026)
di: Rando, Marco, et al.
Pubblicazione: (2026)
AlphaPruning: Using Heavy-Tailed Self Regularization Theory for Improved Layer-wise Pruning of Large Language Models
di: Lu, Haiquan, et al.
Pubblicazione: (2024)
di: Lu, Haiquan, et al.
Pubblicazione: (2024)
Compressing LLMs: The Truth is Rarely Pure and Never Simple
di: Jaiswal, Ajay, et al.
Pubblicazione: (2023)
di: Jaiswal, Ajay, et al.
Pubblicazione: (2023)
FFN-SkipLLM: A Hidden Gem for Autoregressive Decoding with Adaptive Feed Forward Skipping
di: Jaiswal, Ajay, et al.
Pubblicazione: (2024)
di: Jaiswal, Ajay, et al.
Pubblicazione: (2024)
Assortment of Attention Heads: Accelerating Federated PEFT with Head Pruning and Strategic Client Selection
di: Venkatesha, Yeshwanth, et al.
Pubblicazione: (2025)
di: Venkatesha, Yeshwanth, et al.
Pubblicazione: (2025)
Rethinking the Starting Point: Collaborative Pre-Training for Federated Downstream Tasks
di: Chu, Yun-Wei, et al.
Pubblicazione: (2024)
di: Chu, Yun-Wei, et al.
Pubblicazione: (2024)
Generalized Graph Prompt: Toward a Unification of Pre-Training and Downstream Tasks on Graphs
di: Yu, Xingtong, et al.
Pubblicazione: (2023)
di: Yu, Xingtong, et al.
Pubblicazione: (2023)
On-the-Fly Adaptive Distillation of Transformer to Dual-State Linear Attention
di: Ro, Yeonju, et al.
Pubblicazione: (2025)
di: Ro, Yeonju, et al.
Pubblicazione: (2025)
SEAL: Steerable Reasoning Calibration of Large Language Models for Free
di: Chen, Runjin, et al.
Pubblicazione: (2025)
di: Chen, Runjin, et al.
Pubblicazione: (2025)
LLaGA: Large Language and Graph Assistant
di: Chen, Runjin, et al.
Pubblicazione: (2024)
di: Chen, Runjin, et al.
Pubblicazione: (2024)
SPAM: Spike-Aware Adam with Momentum Reset for Stable LLM Training
di: Huang, Tianjin, et al.
Pubblicazione: (2025)
di: Huang, Tianjin, et al.
Pubblicazione: (2025)
Censorship and Junk Food Journalism.
di: Jensen, Carl
Pubblicazione: (1984)
di: Jensen, Carl
Pubblicazione: (1984)
RedVTP: Training-Free Acceleration of Diffusion Vision-Language Models Inference via Masked Token-Guided Visual Token Pruning
di: Xu, Jingqi, et al.
Pubblicazione: (2025)
di: Xu, Jingqi, et al.
Pubblicazione: (2025)
Pushing the Limits of Sparsity: A Bag of Tricks for Extreme Pruning
di: Li, Andy, et al.
Pubblicazione: (2024)
di: Li, Andy, et al.
Pubblicazione: (2024)
You Can Have Better Graph Neural Networks by Not Training Weights at All: Finding Untrained GNNs Tickets
di: Huang, Tianjin, et al.
Pubblicazione: (2022)
di: Huang, Tianjin, et al.
Pubblicazione: (2022)
Improving the Throughput of Diffusion-based Large Language Models via a Training-Free Confidence-Aware Calibration
di: Shen, Jucheng, et al.
Pubblicazione: (2025)
di: Shen, Jucheng, et al.
Pubblicazione: (2025)
Mix-LN: Unleashing the Power of Deeper Layers by Combining Pre-LN and Post-LN
di: Li, Pengxiang, et al.
Pubblicazione: (2024)
di: Li, Pengxiang, et al.
Pubblicazione: (2024)
A Study of Student Dependency on Artificial Intelligence Applications in their Education: With Reference to Indore City
di: Ajay Jaiswal
Pubblicazione: (2025)
di: Ajay Jaiswal
Pubblicazione: (2025)
Full-Rank No More: Low-Rank Weight Training for Modern Speech Recognition Models
di: Fernandez-Lopez, Adriana, et al.
Pubblicazione: (2024)
di: Fernandez-Lopez, Adriana, et al.
Pubblicazione: (2024)
Pruning Weights but Not Truth: Safeguarding Truthfulness While Pruning LLMs
di: Fu, Yao, et al.
Pubblicazione: (2025)
di: Fu, Yao, et al.
Pubblicazione: (2025)
Value-Based Pre-Training with Downstream Feedback
di: Ke, Shuqi, et al.
Pubblicazione: (2026)
di: Ke, Shuqi, et al.
Pubblicazione: (2026)
When Fewer Layers Break More Chains: Layer Pruning Harms Test-Time Scaling in LLMs
di: Wang, Keyu, et al.
Pubblicazione: (2025)
di: Wang, Keyu, et al.
Pubblicazione: (2025)
From Junk DNA to Genomic Treasure: Impacts of Transposable Element DNA, RNA, and Protein in Mammalian Development and Disease
di: Ten D. Li, et al.
Pubblicazione: (2025)
di: Ten D. Li, et al.
Pubblicazione: (2025)
Difficult Task Yes but Simple Task No: Unveiling the Laziness in Multimodal LLMs
di: Zhao, Sihang, et al.
Pubblicazione: (2024)
di: Zhao, Sihang, et al.
Pubblicazione: (2024)
IDEA Prune: An Integrated Enlarge-and-Prune Pipeline in Generative Language Model Pretraining
di: Li, Yixiao, et al.
Pubblicazione: (2025)
di: Li, Yixiao, et al.
Pubblicazione: (2025)
Pre-Trained Model Recommendation for Downstream Fine-tuning
di: Bai, Jiameng, et al.
Pubblicazione: (2024)
di: Bai, Jiameng, et al.
Pubblicazione: (2024)
CLAMP-ViT: Contrastive Data-Free Learning for Adaptive Post-Training Quantization of ViTs
di: Ramachandran, Akshat, et al.
Pubblicazione: (2024)
di: Ramachandran, Akshat, et al.
Pubblicazione: (2024)
Fast and Accurate Probing of In-Training LLMs' Downstream Performances
di: Liu, Zhichen, et al.
Pubblicazione: (2026)
di: Liu, Zhichen, et al.
Pubblicazione: (2026)
Understanding and Mitigating the Label Noise in Pre-training on Downstream Tasks
di: Chen, Hao, et al.
Pubblicazione: (2023)
di: Chen, Hao, et al.
Pubblicazione: (2023)
Task-tailored Pre-processing: Fair Downstream Supervised Learning
di: Sohn, Jinwon, et al.
Pubblicazione: (2026)
di: Sohn, Jinwon, et al.
Pubblicazione: (2026)
Learning to Rewrite Prompts for Bootstrapping LLMs on Downstream Tasks
di: Zhou, Qinhao, et al.
Pubblicazione: (2025)
di: Zhou, Qinhao, et al.
Pubblicazione: (2025)
GPTailor: Large Language Model Pruning Through Layer Cutting and Stitching
di: Su, Guinan, et al.
Pubblicazione: (2025)
di: Su, Guinan, et al.
Pubblicazione: (2025)
Dr. Space Junk vs the Universe: Archaeology and the Future
di: Ilse María Sosa Ehnis
Pubblicazione: (2023)
di: Ilse María Sosa Ehnis
Pubblicazione: (2023)
Exploring Learning Complexity for Efficient Downstream Dataset Pruning
di: Jiang, Wenyu, et al.
Pubblicazione: (2024)
di: Jiang, Wenyu, et al.
Pubblicazione: (2024)
Documenti analoghi
-
AlphaDecay: Module-wise Weight Decay for Heavy-Tailed Balancing in LLMs
di: He, Di, et al.
Pubblicazione: (2025) -
From Low Rank Gradient Subspace Stabilization to Low-Rank Weights: Observations, Theories, and Applications
di: Jaiswal, Ajay, et al.
Pubblicazione: (2024) -
Q-GaLore: Quantized GaLore with INT4 Projection and Layer-Adaptive Low-Rank Gradients
di: Zhang, Zhenyu, et al.
Pubblicazione: (2024) -
Outlier Weighed Layerwise Sparsity (OWL): A Missing Secret Sauce for Pruning LLMs to High Sparsity
di: Yin, Lu, et al.
Pubblicazione: (2023) -
Is C4 Dataset Optimal for Pruning? An Investigation of Calibration Data for LLM Pruning
di: Bandari, Abhinav, et al.
Pubblicazione: (2024)