What is the Best Sequence Length for BABYLM?
Fuente:
arXiv
Salvato in:
| Autori principali: | Salhan, Suchir, Martinez, Richard Diehl, Goriely, Zébulon, Buttery, Paula |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Less is More: Pre-Training Cross-Lingual Small-Scale Language Models with Cognitively-Plausible Curriculum Learning Strategies
di: Salhan, Suchir, et al.
Pubblicazione: (2024)
di: Salhan, Suchir, et al.
Pubblicazione: (2024)
ByteSpan: Information-Driven Subword Tokenisation
di: Goriely, Zébulon, et al.
Pubblicazione: (2025)
di: Goriely, Zébulon, et al.
Pubblicazione: (2025)
IPA-CHILDES & G2P+: Feature-Rich Resources for Cross-Lingual Phonology and Phonemic Language Modeling
di: Goriely, Zébulon, et al.
Pubblicazione: (2025)
di: Goriely, Zébulon, et al.
Pubblicazione: (2025)
BabyLM's First Words: Word Segmentation as a Phonological Probing Task
di: Goriely, Zébulon, et al.
Pubblicazione: (2025)
di: Goriely, Zébulon, et al.
Pubblicazione: (2025)
Mitigating Frequency Bias and Anisotropy in Language Model Pre-Training with Syntactic Smoothing
di: Martinez, Richard Diehl, et al.
Pubblicazione: (2024)
di: Martinez, Richard Diehl, et al.
Pubblicazione: (2024)
From Babble to Words: Pre-Training Language Models on Continuous Streams of Phonemes
di: Goriely, Zébulon, et al.
Pubblicazione: (2024)
di: Goriely, Zébulon, et al.
Pubblicazione: (2024)
Meta-Pretraining for Zero-Shot Cross-Lingual Named Entity Recognition in Low-Resource Philippine Languages
di: Africa, David Demitri, et al.
Pubblicazione: (2025)
di: Africa, David Demitri, et al.
Pubblicazione: (2025)
Pico: A Modular Framework for Hypothesis-Driven Small Language Model Research
di: Martinez, Richard Diehl, et al.
Pubblicazione: (2025)
di: Martinez, Richard Diehl, et al.
Pubblicazione: (2025)
BLiSS 1.0: Evaluating Bilingual Learner Competence in Second Language Small Language Models
di: Gao, Yuan, et al.
Pubblicazione: (2025)
di: Gao, Yuan, et al.
Pubblicazione: (2025)
Looking to Learn: Token-wise Dynamic Gating for Low-Resource Vision-Language Modelling
di: Ganescu, Bianca-Mihaela, et al.
Pubblicazione: (2025)
di: Ganescu, Bianca-Mihaela, et al.
Pubblicazione: (2025)
A Computational Operationalisation of Competing Maturational Theories of Syntactic Development via Statistical Grammar Induction
di: Marcheva, Mila, et al.
Pubblicazione: (2026)
di: Marcheva, Mila, et al.
Pubblicazione: (2026)
Tending Towards Stability: Convergence Challenges in Small Language Models
di: Martinez, Richard Diehl, et al.
Pubblicazione: (2024)
di: Martinez, Richard Diehl, et al.
Pubblicazione: (2024)
Teacher Demonstrations in a BabyLM's Zone of Proximal Development for Contingent Multi-Turn Interaction
di: Salhan, Suchir, et al.
Pubblicazione: (2025)
di: Salhan, Suchir, et al.
Pubblicazione: (2025)
Modelling the Diachronic Emergence of Phoneme Frequency Distributions
di: Martín, Fermín Moscoso del Prado, et al.
Pubblicazione: (2026)
di: Martín, Fermín Moscoso del Prado, et al.
Pubblicazione: (2026)
The Distribution of Phoneme Frequencies across the World's Languages: Macroscopic and Microscopic Information-Theoretic Models
di: Martín, Fermín Moscoso del Prado, et al.
Pubblicazione: (2026)
di: Martín, Fermín Moscoso del Prado, et al.
Pubblicazione: (2026)
Learning Dynamics of Meta-Learning in Small Model Pretraining
di: Africa, David Demitri, et al.
Pubblicazione: (2025)
di: Africa, David Demitri, et al.
Pubblicazione: (2025)
Investigating ReLoRA: Effects on the Learning Dynamics of Small Language Models
di: Weiss, Yuval, et al.
Pubblicazione: (2025)
di: Weiss, Yuval, et al.
Pubblicazione: (2025)
Bias Dynamics in BabyLMs: Towards a Compute-Efficient Sandbox for Democratising Pre-Training Debiasing
di: Trhlik, Filip, et al.
Pubblicazione: (2026)
di: Trhlik, Filip, et al.
Pubblicazione: (2026)
BabyLM Turns 4 and Goes Multilingual: Call for Papers for the 2026 BabyLM Workshop
di: Choshen, Leshem, et al.
Pubblicazione: (2026)
di: Choshen, Leshem, et al.
Pubblicazione: (2026)
SumTablets: A Transliteration Dataset of Sumerian Tablets
di: Simmons, Cole, et al.
Pubblicazione: (2026)
di: Simmons, Cole, et al.
Pubblicazione: (2026)
Length-Unbiased Sequence Policy Optimization: Revealing and Controlling Response Length Variation in RLVR
di: Liu, Fanfan, et al.
Pubblicazione: (2026)
di: Liu, Fanfan, et al.
Pubblicazione: (2026)
What is the Best Way for ChatGPT to Translate Poetry?
di: Wang, Shanshan, et al.
Pubblicazione: (2024)
di: Wang, Shanshan, et al.
Pubblicazione: (2024)
Parallelizing Linear Transformers with the Delta Rule over Sequence Length
di: Yang, Songlin, et al.
Pubblicazione: (2024)
di: Yang, Songlin, et al.
Pubblicazione: (2024)
Clip Your Sequences Fairly: Enforcing Length Fairness for Sequence-Level RL
di: Mao, Hanyi, et al.
Pubblicazione: (2025)
di: Mao, Hanyi, et al.
Pubblicazione: (2025)
What Goes Into a LM Acceptability Judgment? Rethinking the Impact of Frequency and Length
di: Tjuatja, Lindia, et al.
Pubblicazione: (2024)
di: Tjuatja, Lindia, et al.
Pubblicazione: (2024)
Gecko: An Efficient Neural Architecture Inherently Processing Sequences with Arbitrary Lengths
di: Ma, Xuezhe, et al.
Pubblicazione: (2026)
di: Ma, Xuezhe, et al.
Pubblicazione: (2026)
Dataset Decomposition: Faster LLM Training with Variable Sequence Length Curriculum
di: Pouransari, Hadi, et al.
Pubblicazione: (2024)
di: Pouransari, Hadi, et al.
Pubblicazione: (2024)
Prompting open-source and commercial language models for grammatical error correction of English learner text
di: Davis, Christopher, et al.
Pubblicazione: (2024)
di: Davis, Christopher, et al.
Pubblicazione: (2024)
Provable Length Generalization in Sequence Prediction via Spectral Filtering
di: Marsden, Annie, et al.
Pubblicazione: (2024)
di: Marsden, Annie, et al.
Pubblicazione: (2024)
What Matters in Evaluating Book-Length Stories? A Systematic Study of Long Story Evaluation
di: Yang, Dingyi, et al.
Pubblicazione: (2025)
di: Yang, Dingyi, et al.
Pubblicazione: (2025)
What is the Best Process Model Representation? A Comparative Analysis for Process Modeling with Large Language Models
di: Brissard, Alexis, et al.
Pubblicazione: (2025)
di: Brissard, Alexis, et al.
Pubblicazione: (2025)
The What, Why, and How of Context Length Extension Techniques in Large Language Models -- A Detailed Survey
di: Pawar, Saurav, et al.
Pubblicazione: (2024)
di: Pawar, Saurav, et al.
Pubblicazione: (2024)
What's the Best Way to Retrieve Slides? A Comparative Study of Multimodal, Caption-Based, and Hybrid Retrieval Techniques
di: Giouroukis, Petros Stylianos, et al.
Pubblicazione: (2025)
di: Giouroukis, Petros Stylianos, et al.
Pubblicazione: (2025)
Efficient Interactive LLM Serving with Proxy Model-based Sequence Length Prediction
di: Qiu, Haoran, et al.
Pubblicazione: (2024)
di: Qiu, Haoran, et al.
Pubblicazione: (2024)
Rubrik's Cube: Testing a New Rubric for Evaluating Explanations on the CUBE dataset
di: Galvan-Sosa, Diana, et al.
Pubblicazione: (2025)
di: Galvan-Sosa, Diana, et al.
Pubblicazione: (2025)
Connecting the Dots: What Graph-Based Text Representations Work Best for Text Classification Using Graph Neural Networks?
di: Bugueño, Margarita, et al.
Pubblicazione: (2023)
di: Bugueño, Margarita, et al.
Pubblicazione: (2023)
What Makes an Evaluation Useful? Common Pitfalls and Best Practices
di: Gekker, Gil, et al.
Pubblicazione: (2025)
di: Gekker, Gil, et al.
Pubblicazione: (2025)
Lightning Attention-2: A Free Lunch for Handling Unlimited Sequence Lengths in Large Language Models
di: Qin, Zhen, et al.
Pubblicazione: (2024)
di: Qin, Zhen, et al.
Pubblicazione: (2024)
Making, not Taking, the Best of N
di: Khairi, Ammar, et al.
Pubblicazione: (2025)
di: Khairi, Ammar, et al.
Pubblicazione: (2025)
Length Value Model: Scalable Value Pretraining for Token-Level Length Modeling
di: Zhang, Zhen, et al.
Pubblicazione: (2026)
di: Zhang, Zhen, et al.
Pubblicazione: (2026)
Documenti analoghi
-
Less is More: Pre-Training Cross-Lingual Small-Scale Language Models with Cognitively-Plausible Curriculum Learning Strategies
di: Salhan, Suchir, et al.
Pubblicazione: (2024) -
ByteSpan: Information-Driven Subword Tokenisation
di: Goriely, Zébulon, et al.
Pubblicazione: (2025) -
IPA-CHILDES & G2P+: Feature-Rich Resources for Cross-Lingual Phonology and Phonemic Language Modeling
di: Goriely, Zébulon, et al.
Pubblicazione: (2025) -
BabyLM's First Words: Word Segmentation as a Phonological Probing Task
di: Goriely, Zébulon, et al.
Pubblicazione: (2025) -
Mitigating Frequency Bias and Anisotropy in Language Model Pre-Training with Syntactic Smoothing
di: Martinez, Richard Diehl, et al.
Pubblicazione: (2024)