Gradient Localization Improves Lifelong Pretraining of Language Models
Fuente:
arXiv
Saved in:
| Main Authors: | Fernandez, Jared, Bisk, Yonatan, Strubell, Emma |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Energy Considerations of Large Language Model Inference and Efficiency Optimizations
by: Fernandez, Jared, et al.
Published: (2025)
by: Fernandez, Jared, et al.
Published: (2025)
Language Models Need Inductive Biases to Count Inductively
by: Chang, Yingshan, et al.
Published: (2024)
by: Chang, Yingshan, et al.
Published: (2024)
Hardware Scaling Trends and Diminishing Returns in Large-Scale Distributed Training
by: Fernandez, Jared, et al.
Published: (2024)
by: Fernandez, Jared, et al.
Published: (2024)
AboutMe: Using Self-Descriptions in Webpages to Document the Effects of English Pretraining Data Filters
by: Lucy, Li, et al.
Published: (2024)
by: Lucy, Li, et al.
Published: (2024)
Affordance RAG: Hierarchical Multimodal Retrieval with Affordance-Aware Embodied Memory for Mobile Manipulation
by: Korekata, Ryosuke, et al.
Published: (2025)
by: Korekata, Ryosuke, et al.
Published: (2025)
Beyond Text: Characterizing Domain Expert Needs in Document Research
by: Gururaja, Sireesh, et al.
Published: (2025)
by: Gururaja, Sireesh, et al.
Published: (2025)
To Build Our Future, We Must Know Our Past: Contextualizing Paradigm Shifts in Natural Language Processing
by: Gururaja, Sireesh, et al.
Published: (2023)
by: Gururaja, Sireesh, et al.
Published: (2023)
Tools Fail: Detecting Silent Errors in Faulty Tools
by: Sun, Jimin, et al.
Published: (2024)
by: Sun, Jimin, et al.
Published: (2024)
Looking beyond the next token
by: Thankaraj, Abitha, et al.
Published: (2025)
by: Thankaraj, Abitha, et al.
Published: (2025)
SOTOPIA-$π$: Interactive Learning of Socially Intelligent Language Agents
by: Wang, Ruiyi, et al.
Published: (2024)
by: Wang, Ruiyi, et al.
Published: (2024)
FicSim: A Dataset for Multi-Faceted Semantic Similarity in Long-Form Fiction
by: Johnson, Natasha, et al.
Published: (2025)
by: Johnson, Natasha, et al.
Published: (2025)
Source-Aware Training Enables Knowledge Attribution in Language Models
by: Khalifa, Muhammad, et al.
Published: (2024)
by: Khalifa, Muhammad, et al.
Published: (2024)
Reinforced Lifelong Editing for Language Models
by: Li, Zherui, et al.
Published: (2025)
by: Li, Zherui, et al.
Published: (2025)
How to Train Your Fact Verifier: Knowledge Transfer with Multimodal Open Models
by: Lee, Jaeyoung, et al.
Published: (2024)
by: Lee, Jaeyoung, et al.
Published: (2024)
Stereotype or Personalization? User Identity Biases Chatbot Recommendations
by: Kantharuban, Anjali, et al.
Published: (2024)
by: Kantharuban, Anjali, et al.
Published: (2024)
Backward Lens: Projecting Language Model Gradients into the Vocabulary Space
by: Katz, Shahar, et al.
Published: (2024)
by: Katz, Shahar, et al.
Published: (2024)
Scalable Data Ablation Approximations for Language Models through Modular Training and Merging
by: Na, Clara, et al.
Published: (2024)
by: Na, Clara, et al.
Published: (2024)
Holistically Evaluating the Environmental Impact of Creating Language Models
by: Morrison, Jacob, et al.
Published: (2025)
by: Morrison, Jacob, et al.
Published: (2025)
Dolma: an Open Corpus of Three Trillion Tokens for Language Model Pretraining Research
by: Soldaini, Luca, et al.
Published: (2024)
by: Soldaini, Luca, et al.
Published: (2024)
Energy and Carbon Considerations of Fine-Tuning BERT
by: Wang, Xiaorong, et al.
Published: (2023)
by: Wang, Xiaorong, et al.
Published: (2023)
Improving Estonian Text Simplification through Pretrained Language Models and Custom Datasets
by: Barbu, Eduard, et al.
Published: (2025)
by: Barbu, Eduard, et al.
Published: (2025)
Code Pretraining Improves Entity Tracking Abilities of Language Models
by: Kim, Najoung, et al.
Published: (2024)
by: Kim, Najoung, et al.
Published: (2024)
The Hidden Cost of Thinking: Energy Use and Environmental Impact of LMs Beyond Pretraining
by: Morrison, Jacob, et al.
Published: (2026)
by: Morrison, Jacob, et al.
Published: (2026)
Carpe Diem: On the Evaluation of World Knowledge in Lifelong Language Models
by: Kim, Yujin, et al.
Published: (2023)
by: Kim, Yujin, et al.
Published: (2023)
A Causal Language Modeling Detour Improves Encoder Continued Pretraining
by: Touchent, Rian, et al.
Published: (2026)
by: Touchent, Rian, et al.
Published: (2026)
Language Models Improve When Pretraining Data Matches Target Tasks
by: Mizrahi, David, et al.
Published: (2025)
by: Mizrahi, David, et al.
Published: (2025)
Self-Regulation and Requesting Interventions
by: Min, So Yeon, et al.
Published: (2025)
by: Min, So Yeon, et al.
Published: (2025)
Exploring Pretraining via Active Forgetting for Improving Cross Lingual Transfer for Decoder Language Models
by: Aggarwal, Divyanshu, et al.
Published: (2024)
by: Aggarwal, Divyanshu, et al.
Published: (2024)
Pretraining Language Models Using Translationese
by: Doshi, Meet, et al.
Published: (2024)
by: Doshi, Meet, et al.
Published: (2024)
Geographic Adaptation of Pretrained Language Models
by: Hofmann, Valentin, et al.
Published: (2022)
by: Hofmann, Valentin, et al.
Published: (2022)
Back Attention: Understanding and Enhancing Multi-Hop Reasoning in Large Language Models
by: Yu, Zeping, et al.
Published: (2025)
by: Yu, Zeping, et al.
Published: (2025)
Planted in Pretraining, Swayed by Finetuning: A Case Study on the Origins of Cognitive Biases in LLMs
by: Itzhak, Itay, et al.
Published: (2025)
by: Itzhak, Itay, et al.
Published: (2025)
Just CHOP: Embarrassingly Simple LLM Compression
by: Jha, Ananya Harsh, et al.
Published: (2023)
by: Jha, Ananya Harsh, et al.
Published: (2023)
Kinetics: Rethinking Test-Time Scaling Laws
by: Sadhukhan, Ranajoy, et al.
Published: (2025)
by: Sadhukhan, Ranajoy, et al.
Published: (2025)
Knowledge in Superposition: Unveiling the Failures of Lifelong Knowledge Editing for Large Language Models
by: Hu, Chenhui, et al.
Published: (2024)
by: Hu, Chenhui, et al.
Published: (2024)
Improving Non-autoregressive Translation Quality with Pretrained Language Model, Embedding Distillation and Upsampling Strategy for CTC
by: Syu, Shen-sian, et al.
Published: (2023)
by: Syu, Shen-sian, et al.
Published: (2023)
Improving Language Plasticity via Pretraining with Active Forgetting
by: Chen, Yihong, et al.
Published: (2023)
by: Chen, Yihong, et al.
Published: (2023)
Towards Lifelong Learning of Large Language Models: A Survey
by: Zheng, Junhao, et al.
Published: (2024)
by: Zheng, Junhao, et al.
Published: (2024)
Lifelong Safety Alignment for Language Models
by: Wang, Haoyu, et al.
Published: (2025)
by: Wang, Haoyu, et al.
Published: (2025)
Progressive Residual Warmup for Language Model Pretraining
by: Chen, Tianhao, et al.
Published: (2026)
by: Chen, Tianhao, et al.
Published: (2026)
Similar Items
-
Energy Considerations of Large Language Model Inference and Efficiency Optimizations
by: Fernandez, Jared, et al.
Published: (2025) -
Language Models Need Inductive Biases to Count Inductively
by: Chang, Yingshan, et al.
Published: (2024) -
Hardware Scaling Trends and Diminishing Returns in Large-Scale Distributed Training
by: Fernandez, Jared, et al.
Published: (2024) -
AboutMe: Using Self-Descriptions in Webpages to Document the Effects of English Pretraining Data Filters
by: Lucy, Li, et al.
Published: (2024) -
Affordance RAG: Hierarchical Multimodal Retrieval with Affordance-Aware Embodied Memory for Mobile Manipulation
by: Korekata, Ryosuke, et al.
Published: (2025)