Text Quality-Based Pruning for Efficient Training of Language Models
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Sharma, Vasu, Padthe, Karthik, Ardalani, Newsha, Tirumala, Kushal, Howes, Russell, Xu, Hu, Huang, Po-Yao, Li, Shang-Wen, Aghajanyan, Armen, Ghosh, Gargi, Zettlemoyer, Luke |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
When Worse is Better: Navigating the compression-generation tradeoff in visual tokenization
von: Ramanujan, Vivek, et al.
Veröffentlicht: (2024)
von: Ramanujan, Vivek, et al.
Veröffentlicht: (2024)
Demystifying CLIP Data
von: Xu, Hu, et al.
Veröffentlicht: (2023)
von: Xu, Hu, et al.
Veröffentlicht: (2023)
Quagmires in SFT-RL Post-Training: When High SFT Scores Mislead and What to Use Instead
von: Kang, Feiyang, et al.
Veröffentlicht: (2025)
von: Kang, Feiyang, et al.
Veröffentlicht: (2025)
MoMa: Efficient Early-Fusion Pre-training with Mixture of Modality-Aware Experts
von: Lin, Xi Victoria, et al.
Veröffentlicht: (2024)
von: Lin, Xi Victoria, et al.
Veröffentlicht: (2024)
Improving MoE Compute Efficiency by Composing Weight and Data Sparsity
von: Kilian, Maciej, et al.
Veröffentlicht: (2026)
von: Kilian, Maciej, et al.
Veröffentlicht: (2026)
Improving Factuality with Explicit Working Memory
von: Chen, Mingda, et al.
Veröffentlicht: (2024)
von: Chen, Mingda, et al.
Veröffentlicht: (2024)
Sieve: Multimodal Dataset Pruning Using Image Captioning Models
von: Mahmoud, Anas, et al.
Veröffentlicht: (2023)
von: Mahmoud, Anas, et al.
Veröffentlicht: (2023)
CAT: Content-Adaptive Image Tokenization
von: Shen, Junhong, et al.
Veröffentlicht: (2025)
von: Shen, Junhong, et al.
Veröffentlicht: (2025)
To 2:4 Sparsity and Beyond: Neuron-level Activation Function to Accelerate LLM Pre-Training
von: Madhyastha, Meghana, et al.
Veröffentlicht: (2026)
von: Madhyastha, Meghana, et al.
Veröffentlicht: (2026)
Brevity is the soul of wit: Pruning long files for code generation
von: Singh, Aaditya K., et al.
Veröffentlicht: (2024)
von: Singh, Aaditya K., et al.
Veröffentlicht: (2024)
SkillRater: Untangling Capabilities in Multimodal Data
von: Sahi, Naveen, et al.
Veröffentlicht: (2026)
von: Sahi, Naveen, et al.
Veröffentlicht: (2026)
Memory Layers at Scale
von: Berges, Vincent-Pierre, et al.
Veröffentlicht: (2024)
von: Berges, Vincent-Pierre, et al.
Veröffentlicht: (2024)
Continual Learning via Sparse Memory Finetuning
von: Lin, Jessy, et al.
Veröffentlicht: (2025)
von: Lin, Jessy, et al.
Veröffentlicht: (2025)
Transfusion: Predict the Next Token and Diffuse Images with One Multi-Modal Model
von: Zhou, Chunting, et al.
Veröffentlicht: (2024)
von: Zhou, Chunting, et al.
Veröffentlicht: (2024)
Annotating the Pangenome Reveals the Diversity in the Genetic Basis for Metabolic Enzymes
von: Ardalani, Omid
Veröffentlicht: (2025)
von: Ardalani, Omid
Veröffentlicht: (2025)
Latent Speech-Text Transformer
von: Lu, Yen-Ju, et al.
Veröffentlicht: (2025)
von: Lu, Yen-Ju, et al.
Veröffentlicht: (2025)
Altogether: Image Captioning via Re-aligning Alt-text
von: Xu, Hu, et al.
Veröffentlicht: (2024)
von: Xu, Hu, et al.
Veröffentlicht: (2024)
The weighted Bergman spaces and complex reflection groups
von: Ghosh, Gargi
Veröffentlicht: (2021)
von: Ghosh, Gargi
Veröffentlicht: (2021)
Small Molecule Optimization with Large Language Models
von: Guevorguian, Philipp, et al.
Veröffentlicht: (2024)
von: Guevorguian, Philipp, et al.
Veröffentlicht: (2024)
Demystifying Synthetic Data in LLM Pre-training: A Systematic Study of Scaling Laws, Benefits, and Pitfalls
von: Kang, Feiyang, et al.
Veröffentlicht: (2025)
von: Kang, Feiyang, et al.
Veröffentlicht: (2025)
Fast Byte Latent Transformer
von: Kallini, Julie, et al.
Veröffentlicht: (2026)
von: Kallini, Julie, et al.
Veröffentlicht: (2026)
MoDE: CLIP Data Experts via Clustering
von: Ma, Jiawei, et al.
Veröffentlicht: (2024)
von: Ma, Jiawei, et al.
Veröffentlicht: (2024)
ROAST: Risk-aware Outlier-exposure for Adversarial Selective Training of Anomaly Detectors Against Evasion Attacks
von: Elnawawy, Mohammed, et al.
Veröffentlicht: (2026)
von: Elnawawy, Mohammed, et al.
Veröffentlicht: (2026)
Compute Optimal Tokenization
von: Limisiewicz, Tomasz, et al.
Veröffentlicht: (2026)
von: Limisiewicz, Tomasz, et al.
Veröffentlicht: (2026)
Pruning for Performance: Efficient Idiom and Metaphor Classification in Low-Resource Konkani Using mBERT
von: Do, Timothy, et al.
Veröffentlicht: (2025)
von: Do, Timothy, et al.
Veröffentlicht: (2025)
Eigenvalue bounds for the Steklov problem on differential forms in warped product manifolds
von: Chakradhar, Tirumala
Veröffentlicht: (2024)
von: Chakradhar, Tirumala
Veröffentlicht: (2024)
The Unreasonable Ineffectiveness of the Deeper Layers
von: Gromov, Andrey, et al.
Veröffentlicht: (2024)
von: Gromov, Andrey, et al.
Veröffentlicht: (2024)
MAD Max Beyond Single-Node: Enabling Large Machine Learning Model Acceleration on Distributed Systems
von: Hsia, Samuel, et al.
Veröffentlicht: (2023)
von: Hsia, Samuel, et al.
Veröffentlicht: (2023)
CATransformers: Carbon Aware Transformers Through Joint Model-Hardware Optimization
von: Wang, Irene, et al.
Veröffentlicht: (2025)
von: Wang, Irene, et al.
Veröffentlicht: (2025)
$2$-proper holomorphic images of classical Cartan domains
von: Ghosh, Gargi, et al.
Veröffentlicht: (2023)
von: Ghosh, Gargi, et al.
Veröffentlicht: (2023)
Holomorphic retracts in the Lie ball and the tetrablock
von: Ghosh, Gargi, et al.
Veröffentlicht: (2024)
von: Ghosh, Gargi, et al.
Veröffentlicht: (2024)
VoiceCraft: Zero-Shot Speech Editing and Text-to-Speech in the Wild
von: Peng, Puyuan, et al.
Veröffentlicht: (2024)
von: Peng, Puyuan, et al.
Veröffentlicht: (2024)
Effective pruning of web-scale datasets based on complexity of concept clusters
von: Abbas, Amro, et al.
Veröffentlicht: (2024)
von: Abbas, Amro, et al.
Veröffentlicht: (2024)
Folding Tensor and Sequence Parallelism for Memory-Efficient Transformer Training & Inference
von: Shyam, Vasu, et al.
Veröffentlicht: (2026)
von: Shyam, Vasu, et al.
Veröffentlicht: (2026)
Mixture-of-Transformers: A Sparse and Scalable Architecture for Multi-Modal Foundation Models
von: Liang, Weixin, et al.
Veröffentlicht: (2024)
von: Liang, Weixin, et al.
Veröffentlicht: (2024)
Train for Truth, Keep the Skills: Binary Retrieval-Augmented Reward Mitigates Hallucinations
von: Chen, Tong, et al.
Veröffentlicht: (2025)
von: Chen, Tong, et al.
Veröffentlicht: (2025)
DeepCausalMMM: A Deep Learning Framework for Marketing Mix Modeling with Causal Structure Learning
von: Tirumala, Aditya Puttaparthi
Veröffentlicht: (2025)
von: Tirumala, Aditya Puttaparthi
Veröffentlicht: (2025)
Computational Tradeoffs in Image Synthesis: Diffusion, Masked-Token, and Next-Token Prediction
von: Kilian, Maciej, et al.
Veröffentlicht: (2024)
von: Kilian, Maciej, et al.
Veröffentlicht: (2024)
Comparing Hallucination Detection Metrics for Multilingual Generation
von: Kang, Haoqiang, et al.
Veröffentlicht: (2024)
von: Kang, Haoqiang, et al.
Veröffentlicht: (2024)
Multimodal RewardBench: Holistic Evaluation of Reward Models for Vision Language Models
von: Yasunaga, Michihiro, et al.
Veröffentlicht: (2025)
von: Yasunaga, Michihiro, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
When Worse is Better: Navigating the compression-generation tradeoff in visual tokenization
von: Ramanujan, Vivek, et al.
Veröffentlicht: (2024) -
Demystifying CLIP Data
von: Xu, Hu, et al.
Veröffentlicht: (2023) -
Quagmires in SFT-RL Post-Training: When High SFT Scores Mislead and What to Use Instead
von: Kang, Feiyang, et al.
Veröffentlicht: (2025) -
MoMa: Efficient Early-Fusion Pre-training with Mixture of Modality-Aware Experts
von: Lin, Xi Victoria, et al.
Veröffentlicht: (2024) -
Improving MoE Compute Efficiency by Composing Weight and Data Sparsity
von: Kilian, Maciej, et al.
Veröffentlicht: (2026)