Text Quality-Based Pruning for Efficient Training of Language Models
Fuente:
arXiv
Salvato in:
| Autori principali: | Sharma, Vasu, Padthe, Karthik, Ardalani, Newsha, Tirumala, Kushal, Howes, Russell, Xu, Hu, Huang, Po-Yao, Li, Shang-Wen, Aghajanyan, Armen, Ghosh, Gargi, Zettlemoyer, Luke |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
When Worse is Better: Navigating the compression-generation tradeoff in visual tokenization
di: Ramanujan, Vivek, et al.
Pubblicazione: (2024)
di: Ramanujan, Vivek, et al.
Pubblicazione: (2024)
Demystifying CLIP Data
di: Xu, Hu, et al.
Pubblicazione: (2023)
di: Xu, Hu, et al.
Pubblicazione: (2023)
Quagmires in SFT-RL Post-Training: When High SFT Scores Mislead and What to Use Instead
di: Kang, Feiyang, et al.
Pubblicazione: (2025)
di: Kang, Feiyang, et al.
Pubblicazione: (2025)
MoMa: Efficient Early-Fusion Pre-training with Mixture of Modality-Aware Experts
di: Lin, Xi Victoria, et al.
Pubblicazione: (2024)
di: Lin, Xi Victoria, et al.
Pubblicazione: (2024)
Improving MoE Compute Efficiency by Composing Weight and Data Sparsity
di: Kilian, Maciej, et al.
Pubblicazione: (2026)
di: Kilian, Maciej, et al.
Pubblicazione: (2026)
Improving Factuality with Explicit Working Memory
di: Chen, Mingda, et al.
Pubblicazione: (2024)
di: Chen, Mingda, et al.
Pubblicazione: (2024)
Sieve: Multimodal Dataset Pruning Using Image Captioning Models
di: Mahmoud, Anas, et al.
Pubblicazione: (2023)
di: Mahmoud, Anas, et al.
Pubblicazione: (2023)
CAT: Content-Adaptive Image Tokenization
di: Shen, Junhong, et al.
Pubblicazione: (2025)
di: Shen, Junhong, et al.
Pubblicazione: (2025)
To 2:4 Sparsity and Beyond: Neuron-level Activation Function to Accelerate LLM Pre-Training
di: Madhyastha, Meghana, et al.
Pubblicazione: (2026)
di: Madhyastha, Meghana, et al.
Pubblicazione: (2026)
Brevity is the soul of wit: Pruning long files for code generation
di: Singh, Aaditya K., et al.
Pubblicazione: (2024)
di: Singh, Aaditya K., et al.
Pubblicazione: (2024)
SkillRater: Untangling Capabilities in Multimodal Data
di: Sahi, Naveen, et al.
Pubblicazione: (2026)
di: Sahi, Naveen, et al.
Pubblicazione: (2026)
Memory Layers at Scale
di: Berges, Vincent-Pierre, et al.
Pubblicazione: (2024)
di: Berges, Vincent-Pierre, et al.
Pubblicazione: (2024)
Continual Learning via Sparse Memory Finetuning
di: Lin, Jessy, et al.
Pubblicazione: (2025)
di: Lin, Jessy, et al.
Pubblicazione: (2025)
Transfusion: Predict the Next Token and Diffuse Images with One Multi-Modal Model
di: Zhou, Chunting, et al.
Pubblicazione: (2024)
di: Zhou, Chunting, et al.
Pubblicazione: (2024)
Annotating the Pangenome Reveals the Diversity in the Genetic Basis for Metabolic Enzymes
di: Ardalani, Omid
Pubblicazione: (2025)
di: Ardalani, Omid
Pubblicazione: (2025)
Latent Speech-Text Transformer
di: Lu, Yen-Ju, et al.
Pubblicazione: (2025)
di: Lu, Yen-Ju, et al.
Pubblicazione: (2025)
Altogether: Image Captioning via Re-aligning Alt-text
di: Xu, Hu, et al.
Pubblicazione: (2024)
di: Xu, Hu, et al.
Pubblicazione: (2024)
The weighted Bergman spaces and complex reflection groups
di: Ghosh, Gargi
Pubblicazione: (2021)
di: Ghosh, Gargi
Pubblicazione: (2021)
Small Molecule Optimization with Large Language Models
di: Guevorguian, Philipp, et al.
Pubblicazione: (2024)
di: Guevorguian, Philipp, et al.
Pubblicazione: (2024)
Demystifying Synthetic Data in LLM Pre-training: A Systematic Study of Scaling Laws, Benefits, and Pitfalls
di: Kang, Feiyang, et al.
Pubblicazione: (2025)
di: Kang, Feiyang, et al.
Pubblicazione: (2025)
Fast Byte Latent Transformer
di: Kallini, Julie, et al.
Pubblicazione: (2026)
di: Kallini, Julie, et al.
Pubblicazione: (2026)
MoDE: CLIP Data Experts via Clustering
di: Ma, Jiawei, et al.
Pubblicazione: (2024)
di: Ma, Jiawei, et al.
Pubblicazione: (2024)
ROAST: Risk-aware Outlier-exposure for Adversarial Selective Training of Anomaly Detectors Against Evasion Attacks
di: Elnawawy, Mohammed, et al.
Pubblicazione: (2026)
di: Elnawawy, Mohammed, et al.
Pubblicazione: (2026)
Compute Optimal Tokenization
di: Limisiewicz, Tomasz, et al.
Pubblicazione: (2026)
di: Limisiewicz, Tomasz, et al.
Pubblicazione: (2026)
Pruning for Performance: Efficient Idiom and Metaphor Classification in Low-Resource Konkani Using mBERT
di: Do, Timothy, et al.
Pubblicazione: (2025)
di: Do, Timothy, et al.
Pubblicazione: (2025)
Eigenvalue bounds for the Steklov problem on differential forms in warped product manifolds
di: Chakradhar, Tirumala
Pubblicazione: (2024)
di: Chakradhar, Tirumala
Pubblicazione: (2024)
The Unreasonable Ineffectiveness of the Deeper Layers
di: Gromov, Andrey, et al.
Pubblicazione: (2024)
di: Gromov, Andrey, et al.
Pubblicazione: (2024)
MAD Max Beyond Single-Node: Enabling Large Machine Learning Model Acceleration on Distributed Systems
di: Hsia, Samuel, et al.
Pubblicazione: (2023)
di: Hsia, Samuel, et al.
Pubblicazione: (2023)
CATransformers: Carbon Aware Transformers Through Joint Model-Hardware Optimization
di: Wang, Irene, et al.
Pubblicazione: (2025)
di: Wang, Irene, et al.
Pubblicazione: (2025)
$2$-proper holomorphic images of classical Cartan domains
di: Ghosh, Gargi, et al.
Pubblicazione: (2023)
di: Ghosh, Gargi, et al.
Pubblicazione: (2023)
Holomorphic retracts in the Lie ball and the tetrablock
di: Ghosh, Gargi, et al.
Pubblicazione: (2024)
di: Ghosh, Gargi, et al.
Pubblicazione: (2024)
VoiceCraft: Zero-Shot Speech Editing and Text-to-Speech in the Wild
di: Peng, Puyuan, et al.
Pubblicazione: (2024)
di: Peng, Puyuan, et al.
Pubblicazione: (2024)
Effective pruning of web-scale datasets based on complexity of concept clusters
di: Abbas, Amro, et al.
Pubblicazione: (2024)
di: Abbas, Amro, et al.
Pubblicazione: (2024)
Folding Tensor and Sequence Parallelism for Memory-Efficient Transformer Training & Inference
di: Shyam, Vasu, et al.
Pubblicazione: (2026)
di: Shyam, Vasu, et al.
Pubblicazione: (2026)
Mixture-of-Transformers: A Sparse and Scalable Architecture for Multi-Modal Foundation Models
di: Liang, Weixin, et al.
Pubblicazione: (2024)
di: Liang, Weixin, et al.
Pubblicazione: (2024)
Train for Truth, Keep the Skills: Binary Retrieval-Augmented Reward Mitigates Hallucinations
di: Chen, Tong, et al.
Pubblicazione: (2025)
di: Chen, Tong, et al.
Pubblicazione: (2025)
DeepCausalMMM: A Deep Learning Framework for Marketing Mix Modeling with Causal Structure Learning
di: Tirumala, Aditya Puttaparthi
Pubblicazione: (2025)
di: Tirumala, Aditya Puttaparthi
Pubblicazione: (2025)
Computational Tradeoffs in Image Synthesis: Diffusion, Masked-Token, and Next-Token Prediction
di: Kilian, Maciej, et al.
Pubblicazione: (2024)
di: Kilian, Maciej, et al.
Pubblicazione: (2024)
Comparing Hallucination Detection Metrics for Multilingual Generation
di: Kang, Haoqiang, et al.
Pubblicazione: (2024)
di: Kang, Haoqiang, et al.
Pubblicazione: (2024)
Multimodal RewardBench: Holistic Evaluation of Reward Models for Vision Language Models
di: Yasunaga, Michihiro, et al.
Pubblicazione: (2025)
di: Yasunaga, Michihiro, et al.
Pubblicazione: (2025)
Documenti analoghi
-
When Worse is Better: Navigating the compression-generation tradeoff in visual tokenization
di: Ramanujan, Vivek, et al.
Pubblicazione: (2024) -
Demystifying CLIP Data
di: Xu, Hu, et al.
Pubblicazione: (2023) -
Quagmires in SFT-RL Post-Training: When High SFT Scores Mislead and What to Use Instead
di: Kang, Feiyang, et al.
Pubblicazione: (2025) -
MoMa: Efficient Early-Fusion Pre-training with Mixture of Modality-Aware Experts
di: Lin, Xi Victoria, et al.
Pubblicazione: (2024) -
Improving MoE Compute Efficiency by Composing Weight and Data Sparsity
di: Kilian, Maciej, et al.
Pubblicazione: (2026)