Tending Towards Stability: Convergence Challenges in Small Language Models
Fuente:
arXiv
Saved in:
| Main Authors: | Martinez, Richard Diehl, Lesci, Pietro, Buttery, Paula |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Investigating ReLoRA: Effects on the Learning Dynamics of Small Language Models
by: Weiss, Yuval, et al.
Published: (2025)
by: Weiss, Yuval, et al.
Published: (2025)
Less is More: Pre-Training Cross-Lingual Small-Scale Language Models with Cognitively-Plausible Curriculum Learning Strategies
by: Salhan, Suchir, et al.
Published: (2024)
by: Salhan, Suchir, et al.
Published: (2024)
Learning Dynamics of Meta-Learning in Small Model Pretraining
by: Africa, David Demitri, et al.
Published: (2025)
by: Africa, David Demitri, et al.
Published: (2025)
Pico: A Modular Framework for Hypothesis-Driven Small Language Model Research
by: Martinez, Richard Diehl, et al.
Published: (2025)
by: Martinez, Richard Diehl, et al.
Published: (2025)
Mitigating Frequency Bias and Anisotropy in Language Model Pre-Training with Syntactic Smoothing
by: Martinez, Richard Diehl, et al.
Published: (2024)
by: Martinez, Richard Diehl, et al.
Published: (2024)
From Babble to Words: Pre-Training Language Models on Continuous Streams of Phonemes
by: Goriely, Zébulon, et al.
Published: (2024)
by: Goriely, Zébulon, et al.
Published: (2024)
ByteSpan: Information-Driven Subword Tokenisation
by: Goriely, Zébulon, et al.
Published: (2025)
by: Goriely, Zébulon, et al.
Published: (2025)
What is the Best Sequence Length for BABYLM?
by: Salhan, Suchir, et al.
Published: (2025)
by: Salhan, Suchir, et al.
Published: (2025)
Meta-Pretraining for Zero-Shot Cross-Lingual Named Entity Recognition in Low-Resource Philippine Languages
by: Africa, David Demitri, et al.
Published: (2025)
by: Africa, David Demitri, et al.
Published: (2025)
IPA-CHILDES & G2P+: Feature-Rich Resources for Cross-Lingual Phonology and Phonemic Language Modeling
by: Goriely, Zébulon, et al.
Published: (2025)
by: Goriely, Zébulon, et al.
Published: (2025)
BLiSS 1.0: Evaluating Bilingual Learner Competence in Second Language Small Language Models
by: Gao, Yuan, et al.
Published: (2025)
by: Gao, Yuan, et al.
Published: (2025)
BabyLM's First Words: Word Segmentation as a Phonological Probing Task
by: Goriely, Zébulon, et al.
Published: (2025)
by: Goriely, Zébulon, et al.
Published: (2025)
Bias Dynamics in BabyLMs: Towards a Compute-Efficient Sandbox for Democratising Pre-Training Debiasing
by: Trhlik, Filip, et al.
Published: (2026)
by: Trhlik, Filip, et al.
Published: (2026)
What Language is This? Ask Your Tokenizer
by: Meister, Clara, et al.
Published: (2026)
by: Meister, Clara, et al.
Published: (2026)
AnchorAL: Computationally Efficient Active Learning for Large and Imbalanced Datasets
by: Lesci, Pietro, et al.
Published: (2024)
by: Lesci, Pietro, et al.
Published: (2024)
Looking to Learn: Token-wise Dynamic Gating for Low-Resource Vision-Language Modelling
by: Ganescu, Bianca-Mihaela, et al.
Published: (2025)
by: Ganescu, Bianca-Mihaela, et al.
Published: (2025)
PolyPythias: Stability and Outliers across Fifty Language Model Pre-Training Runs
by: van der Wal, Oskar, et al.
Published: (2025)
by: van der Wal, Oskar, et al.
Published: (2025)
Self-Training Large Language Models for Tool-Use Without Demonstrations
by: Luo, Ne, et al.
Published: (2025)
by: Luo, Ne, et al.
Published: (2025)
SumTablets: A Transliteration Dataset of Sumerian Tablets
by: Simmons, Cole, et al.
Published: (2026)
by: Simmons, Cole, et al.
Published: (2026)
Why Safeguarded Ships Run Aground? Aligned Large Language Models' Safety Mechanisms Tend to Be Anchored in The Template Region
by: Leong, Chak Tou, et al.
Published: (2025)
by: Leong, Chak Tou, et al.
Published: (2025)
Towards Pareto Optimal Throughput in Small Language Model Serving
by: Recasens, Pol G., et al.
Published: (2024)
by: Recasens, Pol G., et al.
Published: (2024)
Causal Estimation of Tokenisation Bias
by: Lesci, Pietro, et al.
Published: (2025)
by: Lesci, Pietro, et al.
Published: (2025)
Towards Reasoning Ability of Small Language Models
by: Srivastava, Gaurav, et al.
Published: (2025)
by: Srivastava, Gaurav, et al.
Published: (2025)
Toward Cybersecurity-Expert Small Language Models
by: Levi, Matan, et al.
Published: (2025)
by: Levi, Matan, et al.
Published: (2025)
Diable: Efficient Dialogue State Tracking as Operations on Tables
by: Lesci, Pietro, et al.
Published: (2023)
by: Lesci, Pietro, et al.
Published: (2023)
Enhancing Vaccine Safety Surveillance: Extracting Vaccine Mentions from Emergency Department Triage Notes Using Fine-Tuned Large Language Models
by: Khademi, Sedigh, et al.
Published: (2025)
by: Khademi, Sedigh, et al.
Published: (2025)
Exploration of Plan-Guided Summarization for Narrative Texts: the Case of Small Language Models
by: Grenander, Matt, et al.
Published: (2025)
by: Grenander, Matt, et al.
Published: (2025)
Toward Informal Language Processing: Knowledge of Slang in Large Language Models
by: Sun, Zhewei, et al.
Published: (2024)
by: Sun, Zhewei, et al.
Published: (2024)
Teacher Demonstrations in a BabyLM's Zone of Proximal Development for Contingent Multi-Turn Interaction
by: Salhan, Suchir, et al.
Published: (2025)
by: Salhan, Suchir, et al.
Published: (2025)
Second Guess: Detecting Uncertainty Through Abstention and Answer Stability in Small Language Models
by: Aravindan, Ashwath Vaithinathan, et al.
Published: (2026)
by: Aravindan, Ashwath Vaithinathan, et al.
Published: (2026)
Towards Reliable Medical Question Answering: Techniques and Challenges in Mitigating Hallucinations in Language Models
by: Pham, Duy Khoa, et al.
Published: (2024)
by: Pham, Duy Khoa, et al.
Published: (2024)
Rethinking Data: Towards Better Performing Domain-Specific Small Language Models
by: Nazarov, Boris, et al.
Published: (2025)
by: Nazarov, Boris, et al.
Published: (2025)
Prompting open-source and commercial language models for grammatical error correction of English learner text
by: Davis, Christopher, et al.
Published: (2024)
by: Davis, Christopher, et al.
Published: (2024)
Small Language Models are Equation Reasoners
by: Kim, Bumjun, et al.
Published: (2024)
by: Kim, Bumjun, et al.
Published: (2024)
A Survey of Small Language Models
by: Van Nguyen, Chien, et al.
Published: (2024)
by: Van Nguyen, Chien, et al.
Published: (2024)
AS-ES Learning: Towards Efficient CoT Learning in Small Models
by: Xi, Nuwa, et al.
Published: (2024)
by: Xi, Nuwa, et al.
Published: (2024)
Measuring Stability Beyond Accuracy in Small Open-Source Medical Large Language Models for Pediatric Endocrinology
by: D'Amario, Vanessa, et al.
Published: (2025)
by: D'Amario, Vanessa, et al.
Published: (2025)
Tolerance Principle and Small Language Model Learning
by: Friedman, Adam E., et al.
Published: (2026)
by: Friedman, Adam E., et al.
Published: (2026)
Small Models, Big Impact: Efficient Corpus and Graph-Based Adaptation of Small Multilingual Language Models for Low-Resource Languages
by: Gurgurov, Daniil, et al.
Published: (2025)
by: Gurgurov, Daniil, et al.
Published: (2025)
Towards Modeling Learner Performance with Large Language Models
by: Neshaei, Seyed Parsa, et al.
Published: (2024)
by: Neshaei, Seyed Parsa, et al.
Published: (2024)
Similar Items
-
Investigating ReLoRA: Effects on the Learning Dynamics of Small Language Models
by: Weiss, Yuval, et al.
Published: (2025) -
Less is More: Pre-Training Cross-Lingual Small-Scale Language Models with Cognitively-Plausible Curriculum Learning Strategies
by: Salhan, Suchir, et al.
Published: (2024) -
Learning Dynamics of Meta-Learning in Small Model Pretraining
by: Africa, David Demitri, et al.
Published: (2025) -
Pico: A Modular Framework for Hypothesis-Driven Small Language Model Research
by: Martinez, Richard Diehl, et al.
Published: (2025) -
Mitigating Frequency Bias and Anisotropy in Language Model Pre-Training with Syntactic Smoothing
by: Martinez, Richard Diehl, et al.
Published: (2024)