Saved in:
| Main Authors: | Yıldız, Çağatay, Ravichandran, Nishaanth Kanna, Sharma, Nitin, Bethge, Matthias, Ermis, Beyza |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2402.17400 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Adaptation Odyssey in LLMs: Why Does Additional Pretraining Sometimes Fail to Improve?
by: Öncel, Fırat, et al.
Published: (2024)
by: Öncel, Fırat, et al.
Published: (2024)
From Raw Corpora to Domain Benchmarks: Automated Evaluation of LLM Domain Expertise
by: Sharma, Nitin, et al.
Published: (2025)
by: Sharma, Nitin, et al.
Published: (2025)
From One to Many: Expanding the Scope of Toxicity Mitigation in Language Models
by: Pozzobon, Luiza, et al.
Published: (2024)
by: Pozzobon, Luiza, et al.
Published: (2024)
On The Fairness Impacts of Hardware Selection in Machine Learning
by: Nelaturu, Sree Harsha, et al.
Published: (2023)
by: Nelaturu, Sree Harsha, et al.
Published: (2023)
Question Answering under Temporal Conflict: Evaluating and Organizing Evolving Knowledge with LLMs
by: Özer, Atahan, et al.
Published: (2025)
by: Özer, Atahan, et al.
Published: (2025)
The State of Multilingual LLM Safety Research: From Measuring the Language Gap to Mitigating It
by: Yong, Zheng-Xin, et al.
Published: (2025)
by: Yong, Zheng-Xin, et al.
Published: (2025)
A Reproduction Study: The Kernel PCA Interpretation of Self-Attention Fails Under Scrutiny
by: Sarıtaş, Karahan, et al.
Published: (2025)
by: Sarıtaş, Karahan, et al.
Published: (2025)
A Practitioner's Guide to Continual Multimodal Pretraining
by: Roth, Karsten, et al.
Published: (2024)
by: Roth, Karsten, et al.
Published: (2024)
Playing repeated games with Large Language Models
by: Akata, Elif, et al.
Published: (2023)
by: Akata, Elif, et al.
Published: (2023)
No "Zero-Shot" Without Exponential Data: Pretraining Concept Frequency Determines Multimodal Model Performance
by: Udandarao, Vishaal, et al.
Published: (2024)
by: Udandarao, Vishaal, et al.
Published: (2024)
Mix Data or Merge Models? Optimizing for Diverse Multi-Task Learning
by: Aakanksha, et al.
Published: (2024)
by: Aakanksha, et al.
Published: (2024)
Multilingual Arbitrage: Optimizing Data Pools to Accelerate Multilingual Progress
by: Odumakinde, Ayomide, et al.
Published: (2024)
by: Odumakinde, Ayomide, et al.
Published: (2024)
Mapping Post-Training Forgetting in Language Models at Scale
by: Harmon, Jackson, et al.
Published: (2025)
by: Harmon, Jackson, et al.
Published: (2025)
The Bicameral Model: Bidirectional Hidden-State Coupling Between Parallel Language Models
by: Flamant, Cedric, et al.
Published: (2026)
by: Flamant, Cedric, et al.
Published: (2026)
Improving the OOD Performance of Closed-Source LLMs on NLI Through Strategic Data Selection
by: Stacey, Joe, et al.
Published: (2025)
by: Stacey, Joe, et al.
Published: (2025)
Infinite dSprites for Disentangled Continual Learning: Separating Memory Edits from Generalization
by: Dziadzio, Sebastian, et al.
Published: (2023)
by: Dziadzio, Sebastian, et al.
Published: (2023)
Object-level Self-Distillation for Vision Pretraining
by: Hızlı, Çağlar, et al.
Published: (2025)
by: Hızlı, Çağlar, et al.
Published: (2025)
Round-Trip Translation Reveals What Frontier Multilingual Benchmarks Miss
by: Skorobogat, Ronald, et al.
Published: (2026)
by: Skorobogat, Ronald, et al.
Published: (2026)
Emergent Abilities of Large Language Models under Continued Pretraining for Language Adaptation
by: Elhady, Ahmed, et al.
Published: (2025)
by: Elhady, Ahmed, et al.
Published: (2025)
CiteME: Can Language Models Accurately Cite Scientific Claims?
by: Press, Ori, et al.
Published: (2024)
by: Press, Ori, et al.
Published: (2024)
Exploring the Benefits of Domain-Pretraining of Generative Large Language Models for Chemistry
by: Acharya, Anurag, et al.
Published: (2024)
by: Acharya, Anurag, et al.
Published: (2024)
A Sober Look at Progress in Language Model Reasoning: Pitfalls and Paths to Reproducibility
by: Hochlehnert, Andreas, et al.
Published: (2025)
by: Hochlehnert, Andreas, et al.
Published: (2025)
The Multilingual Alignment Prism: Aligning Global and Local Preferences to Reduce Harm
by: Aakanksha, et al.
Published: (2024)
by: Aakanksha, et al.
Published: (2024)
Will Large Language Models Transform Clinical Prediction?
by: Yildiz, Yusuf, et al.
Published: (2025)
by: Yildiz, Yusuf, et al.
Published: (2025)
Reflecting on the State of Rehearsal-free Continual Learning with Pretrained Models
by: Thede, Lukas, et al.
Published: (2024)
by: Thede, Lukas, et al.
Published: (2024)
Integrating Text and Time-Series into (Large) Language Models to Predict Medical Outcomes
by: Larbi, Iyadh Ben Cheikh, et al.
Published: (2025)
by: Larbi, Iyadh Ben Cheikh, et al.
Published: (2025)
Identifying latent state transition in non-linear dynamical systems
by: Hızlı, Çağlar, et al.
Published: (2024)
by: Hızlı, Çağlar, et al.
Published: (2024)
The Role of Mixed-Language Documents for Multilingual Large Language Model Pretraining
by: Shao, Jiandong, et al.
Published: (2026)
by: Shao, Jiandong, et al.
Published: (2026)
Reuse, Don't Retrain: A Recipe for Continued Pretraining of Language Models
by: Parmar, Jupinder, et al.
Published: (2024)
by: Parmar, Jupinder, et al.
Published: (2024)
MuCPT: Music-related Natural Language Model Continued Pretraining
by: Tian, Kai, et al.
Published: (2025)
by: Tian, Kai, et al.
Published: (2025)
From Fact to Judgment: Investigating the Impact of Task Framing on LLM Conviction in Dialogue Systems
by: Rabbani, Parisa, et al.
Published: (2025)
by: Rabbani, Parisa, et al.
Published: (2025)
Scalable Influence and Fact Tracing for Large Language Model Pretraining
by: Chang, Tyler A., et al.
Published: (2024)
by: Chang, Tyler A., et al.
Published: (2024)
Pretraining Large Language Models with NVFP4
by: NVIDIA, et al.
Published: (2025)
by: NVIDIA, et al.
Published: (2025)
Pretraining Exposure Explains Popularity Judgments in Large Language Models
by: Mozafari, Jamshid, et al.
Published: (2026)
by: Mozafari, Jamshid, et al.
Published: (2026)
PonderLM: Pretraining Language Models to Ponder in Continuous Space
by: Zeng, Boyi, et al.
Published: (2025)
by: Zeng, Boyi, et al.
Published: (2025)
ChatVis: Automating Scientific Visualization with a Large Language Model
by: Mallick, Tanwi, et al.
Published: (2024)
by: Mallick, Tanwi, et al.
Published: (2024)
DRAssist: Dispute Resolution Assistance using Large Language Models
by: Pawar, Sachin, et al.
Published: (2025)
by: Pawar, Sachin, et al.
Published: (2025)
Pretraining Data Exposure in Large Language Models: A Survey of Membership Inference, Data Contamination, and Security Implications
by: Tong, Ziyi, et al.
Published: (2026)
by: Tong, Ziyi, et al.
Published: (2026)
MLP Memory: A Retriever-Pretrained Memory for Large Language Models
by: Wei, Rubin, et al.
Published: (2025)
by: Wei, Rubin, et al.
Published: (2025)
PretrainRL: Alleviating Factuality Hallucination of Large Language Models at the Beginning
by: Liu, Langming, et al.
Published: (2026)
by: Liu, Langming, et al.
Published: (2026)
Similar Items
-
Adaptation Odyssey in LLMs: Why Does Additional Pretraining Sometimes Fail to Improve?
by: Öncel, Fırat, et al.
Published: (2024) -
From Raw Corpora to Domain Benchmarks: Automated Evaluation of LLM Domain Expertise
by: Sharma, Nitin, et al.
Published: (2025) -
From One to Many: Expanding the Scope of Toxicity Mitigation in Language Models
by: Pozzobon, Luiza, et al.
Published: (2024) -
On The Fairness Impacts of Hardware Selection in Machine Learning
by: Nelaturu, Sree Harsha, et al.
Published: (2023) -
Question Answering under Temporal Conflict: Evaluating and Organizing Evolving Knowledge with LLMs
by: Özer, Atahan, et al.
Published: (2025)