Is It a Free Lunch for Removing Outliers during Pretraining?
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Liao, Baohao, Monz, Christof |
|---|---|
| Format: | Preprint |
| Publié: |
2024
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
3-in-1: 2D Rotary Adaptation for Efficient Finetuning, Efficient Batching and Composability
par: Liao, Baohao, et autres
Publié: (2024)
par: Liao, Baohao, et autres
Publié: (2024)
Self-Hinting Language Models Enhance Reinforcement Learning
par: Liao, Baohao, et autres
Publié: (2026)
par: Liao, Baohao, et autres
Publié: (2026)
ClusComp: A Simple Paradigm for Model Compression and Efficient Finetuning
par: Liao, Baohao, et autres
Publié: (2025)
par: Liao, Baohao, et autres
Publié: (2025)
Reward-Guided Speculative Decoding for Efficient LLM Reasoning
par: Liao, Baohao, et autres
Publié: (2025)
par: Liao, Baohao, et autres
Publié: (2025)
Fractured Chain-of-Thought Reasoning
par: Liao, Baohao, et autres
Publié: (2025)
par: Liao, Baohao, et autres
Publié: (2025)
Reinforce-Ada: An Adaptive Sampling Framework under Non-linear RL Objectives
par: Xiong, Wei, et autres
Publié: (2025)
par: Xiong, Wei, et autres
Publié: (2025)
IKUN for WMT24 General MT Task: LLMs Are here for Multilingual Machine Translation
par: Liao, Baohao, et autres
Publié: (2024)
par: Liao, Baohao, et autres
Publié: (2024)
How to Learn in a Noisy World? Self-Correcting the Real-World Data Noise in Machine Translation
par: Meng, Yan, et autres
Publié: (2024)
par: Meng, Yan, et autres
Publié: (2024)
Analyzing the Evaluation of Cross-Lingual Knowledge Transfer in Multilingual Language Models
par: Rajaee, Sara, et autres
Publié: (2024)
par: Rajaee, Sara, et autres
Publié: (2024)
Disentangling the Roles of Target-Side Transfer and Regularization in Multilingual Machine Translation
par: Meng, Yan, et autres
Publié: (2024)
par: Meng, Yan, et autres
Publié: (2024)
Do Language Models Reason Across Languages?
par: Meng, Yan, et autres
Publié: (2026)
par: Meng, Yan, et autres
Publié: (2026)
ApiQ: Finetuning of 2-Bit Quantized Large Language Model
par: Liao, Baohao, et autres
Publié: (2024)
par: Liao, Baohao, et autres
Publié: (2024)
Communicating with Speakers and Listeners of Different Pragmatic Levels
par: Naszadi, Kata, et autres
Publié: (2024)
par: Naszadi, Kata, et autres
Publié: (2024)
On the Evaluation Practices in Multilingual NLP: Can Machine Translation Offer an Alternative to Human Translations?
par: Choenni, Rochelle, et autres
Publié: (2024)
par: Choenni, Rochelle, et autres
Publié: (2024)
Best-of-L: Cross-Lingual Reward Modeling for Mathematical Reasoning
par: Rajaee, Sara, et autres
Publié: (2025)
par: Rajaee, Sara, et autres
Publié: (2025)
Asking LLMs to Verify First is Almost Free Lunch
par: Wu, Shiguang, et autres
Publié: (2025)
par: Wu, Shiguang, et autres
Publié: (2025)
Free Lunch for Pass@$k$? Low Cost Diverse Sampling for Diffusion Language Models
par: Lamont, Sean, et autres
Publié: (2026)
par: Lamont, Sean, et autres
Publié: (2026)
Is Factuality Enhancement a Free Lunch For LLMs? Better Factuality Can Lead to Worse Context-Faithfulness
par: Bi, Baolong, et autres
Publié: (2024)
par: Bi, Baolong, et autres
Publié: (2024)
Lightning Attention-2: A Free Lunch for Handling Unlimited Sequence Lengths in Large Language Models
par: Qin, Zhen, et autres
Publié: (2024)
par: Qin, Zhen, et autres
Publié: (2024)
Control the Temperature: Selective Sampling for Diverse and High-Quality LLM Outputs
par: Troshin, Sergey, et autres
Publié: (2025)
par: Troshin, Sergey, et autres
Publié: (2025)
No Free Lunch in Active Learning: LLM Embedding Quality Dictates Query Strategy Success
par: Rauch, Lukas, et autres
Publié: (2025)
par: Rauch, Lukas, et autres
Publié: (2025)
No Free Lunch in Language Model Bias Mitigation? Targeted Bias Reduction Can Exacerbate Unmitigated LLM Biases
par: Chand, Shireen, et autres
Publié: (2025)
par: Chand, Shireen, et autres
Publié: (2025)
CoPE: Clipped RoPE as A Scalable Free Lunch for Long Context LLMs
par: Li, Haoran, et autres
Publié: (2026)
par: Li, Haoran, et autres
Publié: (2026)
The SIFo Benchmark: Investigating the Sequential Instruction Following Ability of Large Language Models
par: Chen, Xinyi, et autres
Publié: (2024)
par: Chen, Xinyi, et autres
Publié: (2024)
LLMInit: A Free Lunch from Large Language Models for Selective Initialization of Recommendation
par: Zhang, Weizhi, et autres
Publié: (2025)
par: Zhang, Weizhi, et autres
Publié: (2025)
LiveMathematicianBench: A Live Benchmark for Mathematician-Level Reasoning with Proof Sketches
par: He, Linyang, et autres
Publié: (2026)
par: He, Linyang, et autres
Publié: (2026)
Outliers Dimensions that Disrupt Transformers Are Driven by Frequency
par: Puccetti, Giovanni, et autres
Publié: (2022)
par: Puccetti, Giovanni, et autres
Publié: (2022)
Outlier Dimensions Encode Task-Specific Knowledge
par: Rudman, William, et autres
Publié: (2023)
par: Rudman, William, et autres
Publié: (2023)
ManifoldKV: Training-Free KV Cache Compression via Euclidean Outlier Detection
par: Datta, Debajyoti, et autres
Publié: (2026)
par: Datta, Debajyoti, et autres
Publié: (2026)
Unilogit: Robust Machine Unlearning for LLMs Using Uniform-Target Self-Distillation
par: Vasilev, Stefan, et autres
Publié: (2025)
par: Vasilev, Stefan, et autres
Publié: (2025)
OptRot: Mitigating Weight Outliers via Data-Free Rotations for Post-Training Quantization
par: Gadhikar, Advait, et autres
Publié: (2025)
par: Gadhikar, Advait, et autres
Publié: (2025)
Lost at the Beginning of Reasoning
par: Liao, Baohao, et autres
Publié: (2025)
par: Liao, Baohao, et autres
Publié: (2025)
Rethinking the Outlier Distribution in Large Language Models: An In-depth Study
par: Raman, Rahul, et autres
Publié: (2025)
par: Raman, Rahul, et autres
Publié: (2025)
From Noise to Signal: When Outliers Seed New Topics
par: Zve, Evangelia, et autres
Publié: (2026)
par: Zve, Evangelia, et autres
Publié: (2026)
Identifying Narrative Patterns and Outliers in Holocaust Testimonies Using Topic Modeling
par: Ifergan, Maxim, et autres
Publié: (2024)
par: Ifergan, Maxim, et autres
Publié: (2024)
Cite Pretrain: Retrieval-Free Knowledge Attribution for Large Language Models
par: Huang, Yukun, et autres
Publié: (2025)
par: Huang, Yukun, et autres
Publié: (2025)
CCI4.0: A Bilingual Pretraining Dataset for Enhancing Reasoning in Large Language Models
par: Liu, Guang, et autres
Publié: (2025)
par: Liu, Guang, et autres
Publié: (2025)
Anti-Length Shift: Dynamic Outlier Truncation for Training Efficient Reasoning Models
par: Wu, Wei, et autres
Publié: (2026)
par: Wu, Wei, et autres
Publié: (2026)
Systematic Outliers in Large Language Models
par: An, Yongqi, et autres
Publié: (2025)
par: An, Yongqi, et autres
Publié: (2025)
AI Managed Emergency Documentation with a Pretrained Model
par: Menzies, David, et autres
Publié: (2024)
par: Menzies, David, et autres
Publié: (2024)
Documents similaires
-
3-in-1: 2D Rotary Adaptation for Efficient Finetuning, Efficient Batching and Composability
par: Liao, Baohao, et autres
Publié: (2024) -
Self-Hinting Language Models Enhance Reinforcement Learning
par: Liao, Baohao, et autres
Publié: (2026) -
ClusComp: A Simple Paradigm for Model Compression and Efficient Finetuning
par: Liao, Baohao, et autres
Publié: (2025) -
Reward-Guided Speculative Decoding for Efficient LLM Reasoning
par: Liao, Baohao, et autres
Publié: (2025) -
Fractured Chain-of-Thought Reasoning
par: Liao, Baohao, et autres
Publié: (2025)