Learning Dynamics of Meta-Learning in Small Model Pretraining
Fuente:
arXiv
Saved in:
| Main Authors: | Africa, David Demitri, Weiss, Yuval, Buttery, Paula, Martinez, Richard Diehl |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Investigating ReLoRA: Effects on the Learning Dynamics of Small Language Models
by: Weiss, Yuval, et al.
Published: (2025)
by: Weiss, Yuval, et al.
Published: (2025)
Meta-Pretraining for Zero-Shot Cross-Lingual Named Entity Recognition in Low-Resource Philippine Languages
by: Africa, David Demitri, et al.
Published: (2025)
by: Africa, David Demitri, et al.
Published: (2025)
Pico: A Modular Framework for Hypothesis-Driven Small Language Model Research
by: Martinez, Richard Diehl, et al.
Published: (2025)
by: Martinez, Richard Diehl, et al.
Published: (2025)
Less is More: Pre-Training Cross-Lingual Small-Scale Language Models with Cognitively-Plausible Curriculum Learning Strategies
by: Salhan, Suchir, et al.
Published: (2024)
by: Salhan, Suchir, et al.
Published: (2024)
LURE: Live-Usage Replay Evaluations for Reducing Evaluation Awareness
by: Ivanov, Igor, et al.
Published: (2026)
by: Ivanov, Igor, et al.
Published: (2026)
Steering Awareness: Detecting Activation Steering from Within
by: Rivera, Joshua Fonseca, et al.
Published: (2025)
by: Rivera, Joshua Fonseca, et al.
Published: (2025)
Consistency Training while Mitigating Obfuscation via Rate Matching
by: Imran, Sohaib, et al.
Published: (2026)
by: Imran, Sohaib, et al.
Published: (2026)
Tending Towards Stability: Convergence Challenges in Small Language Models
by: Martinez, Richard Diehl, et al.
Published: (2024)
by: Martinez, Richard Diehl, et al.
Published: (2024)
Looking to Learn: Token-wise Dynamic Gating for Low-Resource Vision-Language Modelling
by: Ganescu, Bianca-Mihaela, et al.
Published: (2025)
by: Ganescu, Bianca-Mihaela, et al.
Published: (2025)
Bias Dynamics in BabyLMs: Towards a Compute-Efficient Sandbox for Democratising Pre-Training Debiasing
by: Trhlik, Filip, et al.
Published: (2026)
by: Trhlik, Filip, et al.
Published: (2026)
No Answer Needed: Predicting LLM Answer Accuracy from Question-Only Linear Probes
by: Cencerrado, Iván Vicente Moreno, et al.
Published: (2025)
by: Cencerrado, Iván Vicente Moreno, et al.
Published: (2025)
Inoculation Prompting: Eliciting traits from LLMs during training can suppress them at test-time
by: Tan, Daniel, et al.
Published: (2025)
by: Tan, Daniel, et al.
Published: (2025)
Identifying a Circuit for Verb Conjugation in GPT-2
by: Africa, David Demitri
Published: (2025)
by: Africa, David Demitri
Published: (2025)
Batayan: A Filipino NLP benchmark for evaluating Large Language Models
by: Montalan, Jann Railey, et al.
Published: (2025)
by: Montalan, Jann Railey, et al.
Published: (2025)
Does Self-Evaluation Enable Wireheading in Language Models?
by: Africa, David Demitri, et al.
Published: (2025)
by: Africa, David Demitri, et al.
Published: (2025)
Alignment Pretraining: AI Discourse Causes Self-Fulfilling (Mis)alignment
by: Tice, Cameron, et al.
Published: (2026)
by: Tice, Cameron, et al.
Published: (2026)
What is the Best Sequence Length for BABYLM?
by: Salhan, Suchir, et al.
Published: (2025)
by: Salhan, Suchir, et al.
Published: (2025)
Discovering Knowledge-Critical Subnetworks in Pretrained Language Models
by: Bayazit, Deniz, et al.
Published: (2023)
by: Bayazit, Deniz, et al.
Published: (2023)
Pretrained LLMs Learn Multiple Types of Uncertainty
by: Cohen, Roi, et al.
Published: (2025)
by: Cohen, Roi, et al.
Published: (2025)
Beyond Random Sampling: Efficient Language Model Pretraining via Curriculum Learning
by: Zhang, Yang, et al.
Published: (2025)
by: Zhang, Yang, et al.
Published: (2025)
Mitigating Frequency Bias and Anisotropy in Language Model Pre-Training with Syntactic Smoothing
by: Martinez, Richard Diehl, et al.
Published: (2024)
by: Martinez, Richard Diehl, et al.
Published: (2024)
From Babble to Words: Pre-Training Language Models on Continuous Streams of Phonemes
by: Goriely, Zébulon, et al.
Published: (2024)
by: Goriely, Zébulon, et al.
Published: (2024)
Accelerating Large Language Model Pretraining via LFR Pedagogy: Learn, Focus, and Review
by: Prakriya, Neha, et al.
Published: (2024)
by: Prakriya, Neha, et al.
Published: (2024)
MachineLearningLM: Scaling Many-shot In-context Learning via Continued Pretraining
by: Dong, Haoyu, et al.
Published: (2025)
by: Dong, Haoyu, et al.
Published: (2025)
AI Managed Emergency Documentation with a Pretrained Model
by: Menzies, David, et al.
Published: (2024)
by: Menzies, David, et al.
Published: (2024)
The Tokenization Bottleneck: How Vocabulary Extension Improves Chemistry Representation Learning in Pretrained Language Models
by: Kalamkar, Prathamesh, et al.
Published: (2025)
by: Kalamkar, Prathamesh, et al.
Published: (2025)
Reanalyzing L2 Preposition Learning with Bayesian Mixed Effects and a Pretrained Language Model
by: Prange, Jakob, et al.
Published: (2023)
by: Prange, Jakob, et al.
Published: (2023)
Small Vocabularies, Big Gains: Pretraining and Tokenization in Time Series Models
by: Roger, Alexis, et al.
Published: (2025)
by: Roger, Alexis, et al.
Published: (2025)
Learning to Learn from Language Feedback with Social Meta-Learning
by: Cook, Jonathan, et al.
Published: (2026)
by: Cook, Jonathan, et al.
Published: (2026)
Dynamic Masking Rate Schedules for MLM Pretraining
by: Ankner, Zachary, et al.
Published: (2023)
by: Ankner, Zachary, et al.
Published: (2023)
Beyond Repetition: Text Simplification and Curriculum Learning for Data-Constrained Pretraining
by: Roque, Matthew Theodore, et al.
Published: (2025)
by: Roque, Matthew Theodore, et al.
Published: (2025)
Meta-aware Learning in text-to-SQL Large Language Model
by: Zhang, Wenda
Published: (2025)
by: Zhang, Wenda
Published: (2025)
Enhancing Vaccine Safety Surveillance: Extracting Vaccine Mentions from Emergency Department Triage Notes Using Fine-Tuned Large Language Models
by: Khademi, Sedigh, et al.
Published: (2025)
by: Khademi, Sedigh, et al.
Published: (2025)
Knowledge Graphs are Implicit Reward Models: Path-Derived Signals Enable Compositional Reasoning
by: Kansal, Yuval, et al.
Published: (2026)
by: Kansal, Yuval, et al.
Published: (2026)
Multilingual Pretraining for Pixel Language Models
by: Kesen, Ilker, et al.
Published: (2025)
by: Kesen, Ilker, et al.
Published: (2025)
Once Upon a Time: Interactive Learning for Storytelling with Small Language Models
by: Martins, Jonas Mayer, et al.
Published: (2025)
by: Martins, Jonas Mayer, et al.
Published: (2025)
Pretrained Generative Language Models as General Learning Frameworks for Sequence-Based Tasks
by: Fauber, Ben
Published: (2024)
by: Fauber, Ben
Published: (2024)
Meta-RTL: Reinforcement-Based Meta-Transfer Learning for Low-Resource Commonsense Reasoning
by: Fu, Yu, et al.
Published: (2024)
by: Fu, Yu, et al.
Published: (2024)
Regional Tiny Stories: Using Small Models to Compare Language Learning and Tokenizer Performance
by: Patil, Nirvan, et al.
Published: (2025)
by: Patil, Nirvan, et al.
Published: (2025)
Transfer Learning for Automated Feedback Generation on Small Datasets
by: Morris, Oscar
Published: (2025)
by: Morris, Oscar
Published: (2025)
Similar Items
-
Investigating ReLoRA: Effects on the Learning Dynamics of Small Language Models
by: Weiss, Yuval, et al.
Published: (2025) -
Meta-Pretraining for Zero-Shot Cross-Lingual Named Entity Recognition in Low-Resource Philippine Languages
by: Africa, David Demitri, et al.
Published: (2025) -
Pico: A Modular Framework for Hypothesis-Driven Small Language Model Research
by: Martinez, Richard Diehl, et al.
Published: (2025) -
Less is More: Pre-Training Cross-Lingual Small-Scale Language Models with Cognitively-Plausible Curriculum Learning Strategies
by: Salhan, Suchir, et al.
Published: (2024) -
LURE: Live-Usage Replay Evaluations for Reducing Evaluation Awareness
by: Ivanov, Igor, et al.
Published: (2026)