Compressing LLMs: The Truth is Rarely Pure and Never Simple
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Jaiswal, Ajay, Gan, Zhe, Du, Xianzhi, Zhang, Bowen, Wang, Zhangyang, Yang, Yinfei |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2023
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
MOFI: Learning Image Representations from Noisy Entity Annotated Images
von: Wu, Wentao, et al.
Veröffentlicht: (2023)
von: Wu, Wentao, et al.
Veröffentlicht: (2023)
IDEA Prune: An Integrated Enlarge-and-Prune Pipeline in Generative Language Model Pretraining
von: Li, Yixiao, et al.
Veröffentlicht: (2025)
von: Li, Yixiao, et al.
Veröffentlicht: (2025)
MANZANO: A Simple and Scalable Unified Multimodal Model with a Hybrid Vision Tokenizer
von: Li, Yanghao, et al.
Veröffentlicht: (2025)
von: Li, Yanghao, et al.
Veröffentlicht: (2025)
CLIP-UP: A Simple and Efficient Mixture-of-Experts CLIP Training Recipe with Sparse Upcycling
von: Wang, Xinze, et al.
Veröffentlicht: (2025)
von: Wang, Xinze, et al.
Veröffentlicht: (2025)
LoCoCo: Dropping In Convolutions for Long Context Compression
von: Cai, Ruisi, et al.
Veröffentlicht: (2024)
von: Cai, Ruisi, et al.
Veröffentlicht: (2024)
Pruning Weights but Not Truth: Safeguarding Truthfulness While Pruning LLMs
von: Fu, Yao, et al.
Veröffentlicht: (2025)
von: Fu, Yao, et al.
Veröffentlicht: (2025)
TruthRL: Incentivizing Truthful LLMs via Reinforcement Learning
von: Wei, Zhepei, et al.
Veröffentlicht: (2025)
von: Wei, Zhepei, et al.
Veröffentlicht: (2025)
AttentionPredictor: Temporal Patterns Matter for KV Cache Compression
von: Yang, Qingyue, et al.
Veröffentlicht: (2025)
von: Yang, Qingyue, et al.
Veröffentlicht: (2025)
AlphaDecay: Module-wise Weight Decay for Heavy-Tailed Balancing in LLMs
von: He, Di, et al.
Veröffentlicht: (2025)
von: He, Di, et al.
Veröffentlicht: (2025)
How Easy is It to Fool Your Multimodal LLMs? An Empirical Analysis on Deceptive Prompts
von: Qian, Yusu, et al.
Veröffentlicht: (2024)
von: Qian, Yusu, et al.
Veröffentlicht: (2024)
How Post-Training Reshapes LLMs: A Mechanistic View on Knowledge, Truthfulness, Refusal, and Confidence
von: Du, Hongzhe, et al.
Veröffentlicht: (2025)
von: Du, Hongzhe, et al.
Veröffentlicht: (2025)
Pico-Banana-400K: A Large-Scale Dataset for Text-Guided Image Editing
von: Qian, Yusu, et al.
Veröffentlicht: (2025)
von: Qian, Yusu, et al.
Veröffentlicht: (2025)
Training Dynamics of Transformers to Recognize Word Co-occurrence via Gradient Flow Analysis
von: Yang, Hongru, et al.
Veröffentlicht: (2024)
von: Yang, Hongru, et al.
Veröffentlicht: (2024)
MM1.5: Methods, Analysis & Insights from Multimodal LLM Fine-tuning
von: Zhang, Haotian, et al.
Veröffentlicht: (2024)
von: Zhang, Haotian, et al.
Veröffentlicht: (2024)
TruthX: Alleviating Hallucinations by Editing Large Language Models in Truthful Space
von: Zhang, Shaolei, et al.
Veröffentlicht: (2024)
von: Zhang, Shaolei, et al.
Veröffentlicht: (2024)
Read-ME: Refactorizing LLMs as Router-Decoupled Mixture of Experts with System Co-Design
von: Cai, Ruisi, et al.
Veröffentlicht: (2024)
von: Cai, Ruisi, et al.
Veröffentlicht: (2024)
FoldGPT: Simple and Effective Large Language Model Compression Scheme
von: Liu, Songwei, et al.
Veröffentlicht: (2024)
von: Liu, Songwei, et al.
Veröffentlicht: (2024)
Rewarding the Rare: Uniqueness-Aware RL for Creative Problem Solving in LLMs
von: Hu, Zhiyuan, et al.
Veröffentlicht: (2026)
von: Hu, Zhiyuan, et al.
Veröffentlicht: (2026)
Ferret-UI 2: Mastering Universal User Interface Understanding Across Platforms
von: Li, Zhangheng, et al.
Veröffentlicht: (2024)
von: Li, Zhangheng, et al.
Veröffentlicht: (2024)
FFN-SkipLLM: A Hidden Gem for Autoregressive Decoding with Adaptive Feed Forward Skipping
von: Jaiswal, Ajay, et al.
Veröffentlicht: (2024)
von: Jaiswal, Ajay, et al.
Veröffentlicht: (2024)
Can LLMs Follow Simple Rules?
von: Mu, Norman, et al.
Veröffentlicht: (2023)
von: Mu, Norman, et al.
Veröffentlicht: (2023)
Elephants Never Forget: Testing Language Models for Memorization of Tabular Data
von: Bordt, Sebastian, et al.
Veröffentlicht: (2024)
von: Bordt, Sebastian, et al.
Veröffentlicht: (2024)
Training LLMs over Neurally Compressed Text
von: Lester, Brian, et al.
Veröffentlicht: (2024)
von: Lester, Brian, et al.
Veröffentlicht: (2024)
The Truth Lies Somewhere in the Middle (of the Generated Tokens)
von: Wang, Sophie L., et al.
Veröffentlicht: (2026)
von: Wang, Sophie L., et al.
Veröffentlicht: (2026)
Weight-Inherited Distillation for Task-Agnostic BERT Compression
von: Wu, Taiqiang, et al.
Veröffentlicht: (2023)
von: Wu, Taiqiang, et al.
Veröffentlicht: (2023)
When Bias Pretends to Be Truth: How Spurious Correlations Undermine Hallucination Detection in LLMs
von: Wang, Shaowen, et al.
Veröffentlicht: (2025)
von: Wang, Shaowen, et al.
Veröffentlicht: (2025)
RLVR Training of LLMs Does Not Improve Thinking Ability for General QA: Evaluation Method and a Simple Solution
von: Li, Kaiyuan, et al.
Veröffentlicht: (2026)
von: Li, Kaiyuan, et al.
Veröffentlicht: (2026)
Condense, Don't Just Prune: Enhancing Efficiency and Performance in MoE Layer Pruning
von: Cao, Mingyu, et al.
Veröffentlicht: (2024)
von: Cao, Mingyu, et al.
Veröffentlicht: (2024)
Adapting LLMs for Efficient Context Processing through Soft Prompt Compression
von: Wang, Cangqing, et al.
Veröffentlicht: (2024)
von: Wang, Cangqing, et al.
Veröffentlicht: (2024)
No Mean Feat: Simple, Strong Baselines for Context Compression
von: Feldman, Yair, et al.
Veröffentlicht: (2025)
von: Feldman, Yair, et al.
Veröffentlicht: (2025)
Junk DNA Hypothesis: Pruning Small Pre-Trained Weights Irreversibly and Monotonically Impairs "Difficult" Downstream Tasks in LLMs
von: Yin, Lu, et al.
Veröffentlicht: (2023)
von: Yin, Lu, et al.
Veröffentlicht: (2023)
TruthFlow: Truthful LLM Generation via Representation Flow Correction
von: Wang, Hanyu, et al.
Veröffentlicht: (2025)
von: Wang, Hanyu, et al.
Veröffentlicht: (2025)
TrimR: Verifier-based Training-Free Thinking Compression for Efficient Test-Time Scaling
von: Lin, Weizhe, et al.
Veröffentlicht: (2025)
von: Lin, Weizhe, et al.
Veröffentlicht: (2025)
LongLLMLingua: Accelerating and Enhancing LLMs in Long Context Scenarios via Prompt Compression
von: Jiang, Huiqiang, et al.
Veröffentlicht: (2023)
von: Jiang, Huiqiang, et al.
Veröffentlicht: (2023)
CRVQ: Channel-Relaxed Vector Quantization for Extreme Compression of LLMs
von: Xu, Yuzhuang, et al.
Veröffentlicht: (2024)
von: Xu, Yuzhuang, et al.
Veröffentlicht: (2024)
ProcrustesGPT: Compressing LLMs with Structured Matrices and Orthogonal Transformations
von: Grishina, Ekaterina, et al.
Veröffentlicht: (2025)
von: Grishina, Ekaterina, et al.
Veröffentlicht: (2025)
MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training
von: McKinzie, Brandon, et al.
Veröffentlicht: (2024)
von: McKinzie, Brandon, et al.
Veröffentlicht: (2024)
Finding Fantastic Experts in MoEs: A Unified Study for Expert Dropping Strategies and Observations
von: Jaiswal, Ajay, et al.
Veröffentlicht: (2025)
von: Jaiswal, Ajay, et al.
Veröffentlicht: (2025)
VeCLIP: Improving CLIP Training via Visual-enriched Captions
von: Lai, Zhengfeng, et al.
Veröffentlicht: (2023)
von: Lai, Zhengfeng, et al.
Veröffentlicht: (2023)
MiniDisc: Minimal Distillation Schedule for Language Model Compression
von: Zhang, Chen, et al.
Veröffentlicht: (2022)
von: Zhang, Chen, et al.
Veröffentlicht: (2022)
Ähnliche Einträge
-
MOFI: Learning Image Representations from Noisy Entity Annotated Images
von: Wu, Wentao, et al.
Veröffentlicht: (2023) -
IDEA Prune: An Integrated Enlarge-and-Prune Pipeline in Generative Language Model Pretraining
von: Li, Yixiao, et al.
Veröffentlicht: (2025) -
MANZANO: A Simple and Scalable Unified Multimodal Model with a Hybrid Vision Tokenizer
von: Li, Yanghao, et al.
Veröffentlicht: (2025) -
CLIP-UP: A Simple and Efficient Mixture-of-Experts CLIP Training Recipe with Sparse Upcycling
von: Wang, Xinze, et al.
Veröffentlicht: (2025) -
LoCoCo: Dropping In Convolutions for Long Context Compression
von: Cai, Ruisi, et al.
Veröffentlicht: (2024)