Junk DNA Hypothesis: Pruning Small Pre-Trained Weights Irreversibly and Monotonically Impairs "Difficult" Downstream Tasks in LLMs
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Yin, Lu, Jaiswal, Ajay, Liu, Shiwei, Kundu, Souvik, Wang, Zhangyang |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2023
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
AlphaDecay: Module-wise Weight Decay for Heavy-Tailed Balancing in LLMs
von: He, Di, et al.
Veröffentlicht: (2025)
von: He, Di, et al.
Veröffentlicht: (2025)
Value-Based Pre-Training with Downstream Feedback
von: Ke, Shuqi, et al.
Veröffentlicht: (2026)
von: Ke, Shuqi, et al.
Veröffentlicht: (2026)
Assortment of Attention Heads: Accelerating Federated PEFT with Head Pruning and Strategic Client Selection
von: Venkatesha, Yeshwanth, et al.
Veröffentlicht: (2025)
von: Venkatesha, Yeshwanth, et al.
Veröffentlicht: (2025)
From Low Rank Gradient Subspace Stabilization to Low-Rank Weights: Observations, Theories, and Applications
von: Jaiswal, Ajay, et al.
Veröffentlicht: (2024)
von: Jaiswal, Ajay, et al.
Veröffentlicht: (2024)
SEAL: Steerable Reasoning Calibration of Large Language Models for Free
von: Chen, Runjin, et al.
Veröffentlicht: (2025)
von: Chen, Runjin, et al.
Veröffentlicht: (2025)
LLaGA: Large Language and Graph Assistant
von: Chen, Runjin, et al.
Veröffentlicht: (2024)
von: Chen, Runjin, et al.
Veröffentlicht: (2024)
Mix-LN: Unleashing the Power of Deeper Layers by Combining Pre-LN and Post-LN
von: Li, Pengxiang, et al.
Veröffentlicht: (2024)
von: Li, Pengxiang, et al.
Veröffentlicht: (2024)
Difficult Task Yes but Simple Task No: Unveiling the Laziness in Multimodal LLMs
von: Zhao, Sihang, et al.
Veröffentlicht: (2024)
von: Zhao, Sihang, et al.
Veröffentlicht: (2024)
SPAM: Spike-Aware Adam with Momentum Reset for Stable LLM Training
von: Huang, Tianjin, et al.
Veröffentlicht: (2025)
von: Huang, Tianjin, et al.
Veröffentlicht: (2025)
Q-GaLore: Quantized GaLore with INT4 Projection and Layer-Adaptive Low-Rank Gradients
von: Zhang, Zhenyu, et al.
Veröffentlicht: (2024)
von: Zhang, Zhenyu, et al.
Veröffentlicht: (2024)
When Fewer Layers Break More Chains: Layer Pruning Harms Test-Time Scaling in LLMs
von: Wang, Keyu, et al.
Veröffentlicht: (2025)
von: Wang, Keyu, et al.
Veröffentlicht: (2025)
Fast and Accurate Probing of In-Training LLMs' Downstream Performances
von: Liu, Zhichen, et al.
Veröffentlicht: (2026)
von: Liu, Zhichen, et al.
Veröffentlicht: (2026)
Outlier Weighed Layerwise Sparsity (OWL): A Missing Secret Sauce for Pruning LLMs to High Sparsity
von: Yin, Lu, et al.
Veröffentlicht: (2023)
von: Yin, Lu, et al.
Veröffentlicht: (2023)
On the Hardness of Junking LLMs
von: Rando, Marco, et al.
Veröffentlicht: (2026)
von: Rando, Marco, et al.
Veröffentlicht: (2026)
Is C4 Dataset Optimal for Pruning? An Investigation of Calibration Data for LLM Pruning
von: Bandari, Abhinav, et al.
Veröffentlicht: (2024)
von: Bandari, Abhinav, et al.
Veröffentlicht: (2024)
Pre-Trained Model Recommendation for Downstream Fine-tuning
von: Bai, Jiameng, et al.
Veröffentlicht: (2024)
von: Bai, Jiameng, et al.
Veröffentlicht: (2024)
Condense, Don't Just Prune: Enhancing Efficiency and Performance in MoE Layer Pruning
von: Cao, Mingyu, et al.
Veröffentlicht: (2024)
von: Cao, Mingyu, et al.
Veröffentlicht: (2024)
IDEA Prune: An Integrated Enlarge-and-Prune Pipeline in Generative Language Model Pretraining
von: Li, Yixiao, et al.
Veröffentlicht: (2025)
von: Li, Yixiao, et al.
Veröffentlicht: (2025)
Understanding and Mitigating the Label Noise in Pre-training on Downstream Tasks
von: Chen, Hao, et al.
Veröffentlicht: (2023)
von: Chen, Hao, et al.
Veröffentlicht: (2023)
Full-Rank No More: Low-Rank Weight Training for Modern Speech Recognition Models
von: Fernandez-Lopez, Adriana, et al.
Veröffentlicht: (2024)
von: Fernandez-Lopez, Adriana, et al.
Veröffentlicht: (2024)
Learning to Rewrite Prompts for Bootstrapping LLMs on Downstream Tasks
von: Zhou, Qinhao, et al.
Veröffentlicht: (2025)
von: Zhou, Qinhao, et al.
Veröffentlicht: (2025)
Multi-domain Knowledge Graph Collaborative Pre-training and Prompt Tuning for Diverse Downstream Tasks
von: Zhang, Yichi, et al.
Veröffentlicht: (2024)
von: Zhang, Yichi, et al.
Veröffentlicht: (2024)
CLAMP-ViT: Contrastive Data-Free Learning for Adaptive Post-Training Quantization of ViTs
von: Ramachandran, Akshat, et al.
Veröffentlicht: (2024)
von: Ramachandran, Akshat, et al.
Veröffentlicht: (2024)
ELAS: Efficient Pre-Training of Low-Rank Large Language Models via 2:4 Activation Sparsity
von: Li, Jiaxi, et al.
Veröffentlicht: (2026)
von: Li, Jiaxi, et al.
Veröffentlicht: (2026)
Optimising Language Models for Downstream Tasks: A Post-Training Perspective
von: Shi, Zhengyan
Veröffentlicht: (2025)
von: Shi, Zhengyan
Veröffentlicht: (2025)
Thermodynamic Irreversibility of Training Algorithms
von: Ziyin, Liu, et al.
Veröffentlicht: (2026)
von: Ziyin, Liu, et al.
Veröffentlicht: (2026)
The Neural Pruning Law Hypothesis
von: Barbulescu, Eugen, et al.
Veröffentlicht: (2025)
von: Barbulescu, Eugen, et al.
Veröffentlicht: (2025)
Protoknowledge Shapes Behaviour of LLMs in Downstream Tasks: Memorization and Generalization with Knowledge Graphs
von: Ranaldi, Federico, et al.
Veröffentlicht: (2025)
von: Ranaldi, Federico, et al.
Veröffentlicht: (2025)
VocabTailor: Dynamic Vocabulary Selection for Downstream Tasks in Small Language Models
von: Zhang, Hanling, et al.
Veröffentlicht: (2025)
von: Zhang, Hanling, et al.
Veröffentlicht: (2025)
GenBFA: An Evolutionary Optimization Approach to Bit-Flip Attacks on LLMs
von: Das, Sanjay, et al.
Veröffentlicht: (2024)
von: Das, Sanjay, et al.
Veröffentlicht: (2024)
One LR Doesn't Fit All: Heavy-Tail Guided Layerwise Learning Rates for LLMs
von: He, Di, et al.
Veröffentlicht: (2026)
von: He, Di, et al.
Veröffentlicht: (2026)
Downstream Transfer Attack: Adversarial Attacks on Downstream Models with Pre-trained Vision Transformers
von: Zheng, Weijie, et al.
Veröffentlicht: (2024)
von: Zheng, Weijie, et al.
Veröffentlicht: (2024)
APTBench: Benchmarking Agentic Potential of Base LLMs During Pre-Training
von: Qin, Jiarui, et al.
Veröffentlicht: (2025)
von: Qin, Jiarui, et al.
Veröffentlicht: (2025)
Effective Learning for Small Reasoning Models: An Empirical Study on 0.5B Reasoning LLMs
von: Zhuang, Xialie, et al.
Veröffentlicht: (2025)
von: Zhuang, Xialie, et al.
Veröffentlicht: (2025)
Routing-Aligned Fine-Tuning for Multilingual Downstream Tasks in Mixture-of-Experts Models
von: Deng, Guanzhi, et al.
Veröffentlicht: (2026)
von: Deng, Guanzhi, et al.
Veröffentlicht: (2026)
Downstream Task-Oriented Generative Model Selections on Synthetic Data Training for Fraud Detection Models
von: Cheng, Yinan, et al.
Veröffentlicht: (2024)
von: Cheng, Yinan, et al.
Veröffentlicht: (2024)
Decoupling Weighing and Selecting for Integrating Multiple Graph Pre-training Tasks
von: Fan, Tianyu, et al.
Veröffentlicht: (2024)
von: Fan, Tianyu, et al.
Veröffentlicht: (2024)
Hypothesis Testing the Circuit Hypothesis in LLMs
von: Shi, Claudia, et al.
Veröffentlicht: (2024)
von: Shi, Claudia, et al.
Veröffentlicht: (2024)
AlphaPruning: Using Heavy-Tailed Self Regularization Theory for Improved Layer-wise Pruning of Large Language Models
von: Lu, Haiquan, et al.
Veröffentlicht: (2024)
von: Lu, Haiquan, et al.
Veröffentlicht: (2024)
Extracting Small Translation Specialists from LLMs by Aggressively Pruning Experts
von: Martin, Liu O., et al.
Veröffentlicht: (2026)
von: Martin, Liu O., et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
AlphaDecay: Module-wise Weight Decay for Heavy-Tailed Balancing in LLMs
von: He, Di, et al.
Veröffentlicht: (2025) -
Value-Based Pre-Training with Downstream Feedback
von: Ke, Shuqi, et al.
Veröffentlicht: (2026) -
Assortment of Attention Heads: Accelerating Federated PEFT with Head Pruning and Strategic Client Selection
von: Venkatesha, Yeshwanth, et al.
Veröffentlicht: (2025) -
From Low Rank Gradient Subspace Stabilization to Low-Rank Weights: Observations, Theories, and Applications
von: Jaiswal, Ajay, et al.
Veröffentlicht: (2024) -
SEAL: Steerable Reasoning Calibration of Large Language Models for Free
von: Chen, Runjin, et al.
Veröffentlicht: (2025)