Language Models on a Diet: Cost-Efficient Development of Encoders for Closely-Related Languages via Additional Pretraining
Fuente:
arXiv
Saved in:
| Main Authors: | Ljubešić, Nikola, Suchomel, Vít, Rupnik, Peter, Kuzman, Taja, van Noord, Rik |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
The Growing Gains and Pains of Iterative Web Corpora Crawling: Insights from South Slavic CLASSLA-web 2.0 Corpora
by: Pungeršek, Taja Kuzman, et al.
Published: (2026)
by: Pungeršek, Taja Kuzman, et al.
Published: (2026)
CLASSLA-web: Comparable Web Corpora of South Slavic Languages Enriched with Linguistic and Genre Annotation
by: Ljubešić, Nikola, et al.
Published: (2024)
by: Ljubešić, Nikola, et al.
Published: (2024)
Do Language Models Care About Text Quality? Evaluating Web-Crawled Corpora Across 11 Languages
by: van Noord, Rik, et al.
Published: (2024)
by: van Noord, Rik, et al.
Published: (2024)
LLM Teacher-Student Framework for Text Classification With No Manually Annotated Data: A Case Study in IPTC News Topic Classification
by: Kuzman, Taja, et al.
Published: (2024)
by: Kuzman, Taja, et al.
Published: (2024)
Supercharging Agenda Setting Research: The ParlaCAP Dataset of 28 European Parliaments and a Scalable Multilingual LLM-Based Classification
by: Pungeršek, Taja Kuzman, et al.
Published: (2026)
by: Pungeršek, Taja Kuzman, et al.
Published: (2026)
ParlaSpeech 3.0: Richly Annotated Spoken Parliamentary Corpora of Croatian, Czech, Polish, and Serbian
by: Ljubešić, Nikola, et al.
Published: (2025)
by: Ljubešić, Nikola, et al.
Published: (2025)
State of the Art in Text Classification for South Slavic Languages: Fine-Tuning or Prompting?
by: Pungeršek, Taja Kuzman, et al.
Published: (2025)
by: Pungeršek, Taja Kuzman, et al.
Published: (2025)
Identifying Primary Stress Across Related Languages and Dialects with Transformer-based Speech Encoder Models
by: Ljubešić, Nikola, et al.
Published: (2025)
by: Ljubešić, Nikola, et al.
Published: (2025)
CLASSLA-Express: a Train of CLARIN.SI Workshops on Language Resources and Tools with Easily Expanding Route
by: Ljubešić, Nikola, et al.
Published: (2024)
by: Ljubešić, Nikola, et al.
Published: (2024)
Charting the European LLM Benchmarking Landscape: A New Taxonomy and a Set of Best Practices
by: Vintar, Špela, et al.
Published: (2025)
by: Vintar, Špela, et al.
Published: (2025)
The ParlaSent Multilingual Training Dataset for Sentiment Identification in Parliamentary Proceedings
by: Mochtak, Michal, et al.
Published: (2023)
by: Mochtak, Michal, et al.
Published: (2023)
Mići Princ -- A Little Boy Teaching Speech Technologies the Chakavian Dialect
by: Ljubešić, Nikola, et al.
Published: (2026)
by: Ljubešić, Nikola, et al.
Published: (2026)
The ParlaSpeech Collection of Automatically Generated Speech and Text Datasets from Parliamentary Proceedings
by: Ljubešić, Nikola, et al.
Published: (2024)
by: Ljubešić, Nikola, et al.
Published: (2024)
Geographic Adaptation of Pretrained Language Models
by: Hofmann, Valentin, et al.
Published: (2022)
by: Hofmann, Valentin, et al.
Published: (2022)
Should We Still Pretrain Encoders with Masked Language Modeling?
by: Gisserot-Boukhlef, Hippolyte, et al.
Published: (2025)
by: Gisserot-Boukhlef, Hippolyte, et al.
Published: (2025)
Creativity Bias: How Machine Evaluation Struggles with Creativity in Literary Translations
by: Gerrits, Kyo, et al.
Published: (2026)
by: Gerrits, Kyo, et al.
Published: (2026)
One Script Instead of Hundreds? On Pretraining Romanized Encoder Language Models
by: Ebing, Benedikt, et al.
Published: (2026)
by: Ebing, Benedikt, et al.
Published: (2026)
PMB5: Gaining More Insight into Neural Semantic Parsing with Challenging Benchmarks
by: Zhang, Xiao, et al.
Published: (2024)
by: Zhang, Xiao, et al.
Published: (2024)
Multi-perspective Alignment for Increasing Naturalness in Neural Machine Translation
by: Lai, Huiyuan, et al.
Published: (2024)
by: Lai, Huiyuan, et al.
Published: (2024)
Towards Tailored Recovery of Lexical Diversity in Literary Machine Translation
by: Ploeger, Esther, et al.
Published: (2024)
by: Ploeger, Esther, et al.
Published: (2024)
A Causal Language Modeling Detour Improves Encoder Continued Pretraining
by: Touchent, Rian, et al.
Published: (2026)
by: Touchent, Rian, et al.
Published: (2026)
Renaissance: Investigating the Pretraining of Vision-Language Encoders
by: Fields, Clayton, et al.
Published: (2024)
by: Fields, Clayton, et al.
Published: (2024)
On Multilingual Encoder Language Model Compression for Low-Resource Languages
by: Gurgurov, Daniil, et al.
Published: (2025)
by: Gurgurov, Daniil, et al.
Published: (2025)
Efficient Pretraining Data Selection for Language Models via Multi-Actor Collaboration
by: Bai, Tianyi, et al.
Published: (2024)
by: Bai, Tianyi, et al.
Published: (2024)
Accurate Retraining-free Pruning for Pretrained Encoder-based Language Models
by: Park, Seungcheol, et al.
Published: (2023)
by: Park, Seungcheol, et al.
Published: (2023)
Affective Polarization across European Parliaments
by: Evkoski, Bojan, et al.
Published: (2025)
by: Evkoski, Bojan, et al.
Published: (2025)
Multilingual Power and Ideology Identification in the Parliament: a Reference Dataset and Simple Baselines
by: Çöltekin, Çağrı, et al.
Published: (2024)
by: Çöltekin, Çağrı, et al.
Published: (2024)
Beyond Random Sampling: Efficient Language Model Pretraining via Curriculum Learning
by: Zhang, Yang, et al.
Published: (2025)
by: Zhang, Yang, et al.
Published: (2025)
A Comprehensive Evaluation of Semantic Relation Knowledge of Pretrained Language Models and Humans
by: Cao, Zhihan, et al.
Published: (2024)
by: Cao, Zhihan, et al.
Published: (2024)
New Textual Corpora for Serbian Language Modeling
by: Škorić, Mihailo, et al.
Published: (2024)
by: Škorić, Mihailo, et al.
Published: (2024)
Pretraining and Benchmarking Modern Encoders for Latvian
by: Znotins, Arturs
Published: (2026)
by: Znotins, Arturs
Published: (2026)
Pretraining Language Models Using Translationese
by: Doshi, Meet, et al.
Published: (2024)
by: Doshi, Meet, et al.
Published: (2024)
Memory-based Language Models: An Efficient, Explainable, and Eco-friendly Approach to Large Language Modeling
by: Bosch, Antal van den, et al.
Published: (2025)
by: Bosch, Antal van den, et al.
Published: (2025)
ESLM: Risk-Averse Selective Language Modeling for Efficient Pretraining
by: Bal, Melis Ilayda, et al.
Published: (2025)
by: Bal, Melis Ilayda, et al.
Published: (2025)
DA-Cramming: Enhancing Cost-Effective Language Model Pretraining with Dependency Agreement Integration
by: Kuo, Martin, et al.
Published: (2023)
by: Kuo, Martin, et al.
Published: (2023)
Encoder-Decoder or Decoder-Only? Revisiting Encoder-Decoder Large Language Model
by: Zhang, Biao, et al.
Published: (2025)
by: Zhang, Biao, et al.
Published: (2025)
Cross-Domain Bilingual Lexicon Induction via Pretrained Language Models
by: Ding, Qiuyu, et al.
Published: (2025)
by: Ding, Qiuyu, et al.
Published: (2025)
Vision-and-Language Pretraining
by: Nguyen, Thong, et al.
Published: (2022)
by: Nguyen, Thong, et al.
Published: (2022)
A Computational Model for the Assessment of Mutual Intelligibility Among Closely Related Languages
by: Nieder, Jessica, et al.
Published: (2024)
by: Nieder, Jessica, et al.
Published: (2024)
How Does Code Pretraining Affect Language Model Task Performance?
by: Petty, Jackson, et al.
Published: (2024)
by: Petty, Jackson, et al.
Published: (2024)
Similar Items
-
The Growing Gains and Pains of Iterative Web Corpora Crawling: Insights from South Slavic CLASSLA-web 2.0 Corpora
by: Pungeršek, Taja Kuzman, et al.
Published: (2026) -
CLASSLA-web: Comparable Web Corpora of South Slavic Languages Enriched with Linguistic and Genre Annotation
by: Ljubešić, Nikola, et al.
Published: (2024) -
Do Language Models Care About Text Quality? Evaluating Web-Crawled Corpora Across 11 Languages
by: van Noord, Rik, et al.
Published: (2024) -
LLM Teacher-Student Framework for Text Classification With No Manually Annotated Data: A Case Study in IPTC News Topic Classification
by: Kuzman, Taja, et al.
Published: (2024) -
Supercharging Agenda Setting Research: The ParlaCAP Dataset of 28 European Parliaments and a Scalable Multilingual LLM-Based Classification
by: Pungeršek, Taja Kuzman, et al.
Published: (2026)