Full-Parameter Continual Pretraining of Gemma2: Insights into Fluency and Domain Knowledge
Fuente:
arXiv
Saved in:
| Main Authors: | Šliogeris, Vytenis, Daniušis, Povilas, Nakvosas, Artūras |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Open Llama2 Model for the Lithuanian Language
by: Nakvosas, Artūras, et al.
Published: (2024)
by: Nakvosas, Artūras, et al.
Published: (2024)
Gemma Scope: Open Sparse Autoencoders Everywhere All At Once on Gemma 2
by: Lieberum, Tom, et al.
Published: (2024)
by: Lieberum, Tom, et al.
Published: (2024)
ixi-GEN: Efficient Industrial sLLMs through Domain Adaptive Continual Pretraining
by: Kim, Seonwu, et al.
Published: (2025)
by: Kim, Seonwu, et al.
Published: (2025)
TxGemma: Efficient and Agentic LLMs for Therapeutics
by: Wang, Eric, et al.
Published: (2025)
by: Wang, Eric, et al.
Published: (2025)
Injecting Structured Biomedical Knowledge into Language Models: Continual Pretraining vs. GraphRAG
by: Klila, Jaafer, et al.
Published: (2026)
by: Klila, Jaafer, et al.
Published: (2026)
RecurrentGemma: Moving Past Transformers for Efficient Open Language Models
by: Botev, Aleksandar, et al.
Published: (2024)
by: Botev, Aleksandar, et al.
Published: (2024)
Discovering Knowledge-Critical Subnetworks in Pretrained Language Models
by: Bayazit, Deniz, et al.
Published: (2023)
by: Bayazit, Deniz, et al.
Published: (2023)
From Bytes to Borsch: Fine-Tuning Gemma and Mistral for the Ukrainian Language Representation
by: Kiulian, Artur, et al.
Published: (2024)
by: Kiulian, Artur, et al.
Published: (2024)
Cite Pretrain: Retrieval-Free Knowledge Attribution for Large Language Models
by: Huang, Yukun, et al.
Published: (2025)
by: Huang, Yukun, et al.
Published: (2025)
Stabilizing Reasoning in Medical LLMs with Continued Pretraining and Reasoning Preference Optimization
by: Kawakami, Wataru, et al.
Published: (2025)
by: Kawakami, Wataru, et al.
Published: (2025)
BIPEFT: Budget-Guided Iterative Search for Parameter Efficient Fine-Tuning of Large Pretrained Language Models
by: Chang, Aofei, et al.
Published: (2024)
by: Chang, Aofei, et al.
Published: (2024)
BioMistral: A Collection of Open-Source Pretrained Large Language Models for Medical Domains
by: Labrak, Yanis, et al.
Published: (2024)
by: Labrak, Yanis, et al.
Published: (2024)
Feeding Two Birds or Favoring One? Adequacy-Fluency Tradeoffs in Evaluation and Meta-Evaluation of Machine Translation
by: Shayegh, Behzad, et al.
Published: (2025)
by: Shayegh, Behzad, et al.
Published: (2025)
HMI: Hierarchical Knowledge Management for Efficient Multi-Tenant Inference in Pretrained Language Models
by: Zhang, Jun, et al.
Published: (2025)
by: Zhang, Jun, et al.
Published: (2025)
MiCA Learns More Knowledge Than LoRA and Full Fine-Tuning
by: Rüdiger, Sten, et al.
Published: (2026)
by: Rüdiger, Sten, et al.
Published: (2026)
HOP to the Next Tasks and Domains for Continual Learning in NLP
by: Michieli, Umberto, et al.
Published: (2024)
by: Michieli, Umberto, et al.
Published: (2024)
LLM-NEO: Parameter Efficient Knowledge Distillation for Large Language Models
by: Yang, Runming, et al.
Published: (2024)
by: Yang, Runming, et al.
Published: (2024)
Algorithm Selection with Zero Domain Knowledge via Text Embeddings
by: Szeider, Stefan
Published: (2026)
by: Szeider, Stefan
Published: (2026)
PPSEBM: An Energy-Based Model with Progressive Parameter Selection for Continual Learning
by: Li, Xiaodi, et al.
Published: (2025)
by: Li, Xiaodi, et al.
Published: (2025)
PaliGemma: A versatile 3B VLM for transfer
by: Beyer, Lucas, et al.
Published: (2024)
by: Beyer, Lucas, et al.
Published: (2024)
DINO Pre-training for Vision-based End-to-end Autonomous Driving
by: Juneja, Shubham, et al.
Published: (2024)
by: Juneja, Shubham, et al.
Published: (2024)
Training Domain Draft Models for Speculative Decoding: Best Practices and Insights
by: Hong, Fenglu, et al.
Published: (2025)
by: Hong, Fenglu, et al.
Published: (2025)
Pretrained Hybrids with MAD Skills
by: Roberts, Nicholas, et al.
Published: (2024)
by: Roberts, Nicholas, et al.
Published: (2024)
Parameter Efficient Diverse Paraphrase Generation Using Sequence-Level Knowledge Distillation
by: Jayawardena, Lasal, et al.
Published: (2024)
by: Jayawardena, Lasal, et al.
Published: (2024)
Knowledge is Not Enough: Injecting RL Skills for Continual Adaptation
by: Tang, Pingzhi, et al.
Published: (2026)
by: Tang, Pingzhi, et al.
Published: (2026)
Training Language Models on the Knowledge Graph: Insights on Hallucinations and Their Detectability
by: Hron, Jiri, et al.
Published: (2024)
by: Hron, Jiri, et al.
Published: (2024)
DKEC: Domain Knowledge Enhanced Multi-Label Classification for Diagnosis Prediction
by: Ge, Xueren, et al.
Published: (2023)
by: Ge, Xueren, et al.
Published: (2023)
Logic Programming on Knowledge Graph Networks And its Application in Medical Domain
by: Wang, Chuanqing, et al.
Published: (2026)
by: Wang, Chuanqing, et al.
Published: (2026)
From Wide to Deep: Dimension Lifting Network for Parameter-efficient Knowledge Graph Embedding
by: Cai, Borui, et al.
Published: (2023)
by: Cai, Borui, et al.
Published: (2023)
TPTT: Transforming Pretrained Transformers into Titans
by: Furfaro, Fabien
Published: (2025)
by: Furfaro, Fabien
Published: (2025)
RLP: Reinforcement as a Pretraining Objective
by: Hatamizadeh, Ali, et al.
Published: (2025)
by: Hatamizadeh, Ali, et al.
Published: (2025)
Memorization Dynamics of Fill-in-the-Middle Pretraining
by: von Arx, Tobias, et al.
Published: (2026)
by: von Arx, Tobias, et al.
Published: (2026)
Recurrent Knowledge Identification and Fusion for Language Model Continual Learning
by: Feng, Yujie, et al.
Published: (2025)
by: Feng, Yujie, et al.
Published: (2025)
On the Difficulty of Token-Level Modeling of Dysfluency and Fluency Shaping Artifacts
by: Gulzar, Kashaf, et al.
Published: (2025)
by: Gulzar, Kashaf, et al.
Published: (2025)
ProfBench: Multi-Domain Rubrics requiring Professional Knowledge to Answer and Judge
by: Wang, Zhilin, et al.
Published: (2025)
by: Wang, Zhilin, et al.
Published: (2025)
Data Mixing Agent: Learning to Re-weight Domains for Continual Pre-training
by: Yang, Kailai, et al.
Published: (2025)
by: Yang, Kailai, et al.
Published: (2025)
Scaling Fine-Grained MoE Beyond 50B Parameters: Empirical Evaluation and Practical Insights
by: Krajewski, Jakub, et al.
Published: (2025)
by: Krajewski, Jakub, et al.
Published: (2025)
MemEIC: A Step Toward Continual and Compositional Knowledge Editing
by: Seong, Jin, et al.
Published: (2025)
by: Seong, Jin, et al.
Published: (2025)
Pretraining Large Language Models with NVFP4
by: NVIDIA, et al.
Published: (2025)
by: NVIDIA, et al.
Published: (2025)
Patent Language Model Pretraining with ModernBERT
by: Yousefiramandi, Amirhossein, et al.
Published: (2025)
by: Yousefiramandi, Amirhossein, et al.
Published: (2025)
Similar Items
-
Open Llama2 Model for the Lithuanian Language
by: Nakvosas, Artūras, et al.
Published: (2024) -
Gemma Scope: Open Sparse Autoencoders Everywhere All At Once on Gemma 2
by: Lieberum, Tom, et al.
Published: (2024) -
ixi-GEN: Efficient Industrial sLLMs through Domain Adaptive Continual Pretraining
by: Kim, Seonwu, et al.
Published: (2025) -
TxGemma: Efficient and Agentic LLMs for Therapeutics
by: Wang, Eric, et al.
Published: (2025) -
Injecting Structured Biomedical Knowledge into Language Models: Continual Pretraining vs. GraphRAG
by: Klila, Jaafer, et al.
Published: (2026)