CLASS-IT: Conversational and Lecture-Aligned Small-Scale Instruction Tuning for BabyLMs
Fuente:
arXiv
Saved in:
| Main Authors: | Capone, Luca, Bondielli, Alessandro, Lenci, Alessandro |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Are BabyLMs Second Language Learners?
by: Edman, Lukas, et al.
Published: (2024)
by: Edman, Lukas, et al.
Published: (2024)
BAMBI: Developing Baby Language Models for Italian
by: Suozzi, Alice, et al.
Published: (2025)
by: Suozzi, Alice, et al.
Published: (2025)
Child-directed speech facilitates production, not comprehension, in BabyLMs
by: Bunzeck, Bastian, et al.
Published: (2026)
by: Bunzeck, Bastian, et al.
Published: (2026)
Do Construction Distributions Shape Formal Language Learning In German BabyLMs?
by: Bunzeck, Bastian, et al.
Published: (2025)
by: Bunzeck, Bastian, et al.
Published: (2025)
Mask and You Shall Receive: Optimizing Masked Language Modeling For Pretraining BabyLMs
by: Edman, Lukas, et al.
Published: (2025)
by: Edman, Lukas, et al.
Published: (2025)
BabyLMs for isiXhosa: Data-Efficient Language Modelling in a Low-Resource Context
by: Matzopoulos, Alexis, et al.
Published: (2025)
by: Matzopoulos, Alexis, et al.
Published: (2025)
Are BabyLMs Deaf to Gricean Maxims? A Pragmatic Evaluation of Sample-efficient Language Models
by: Askari, Raha, et al.
Published: (2025)
by: Askari, Raha, et al.
Published: (2025)
BabyLM Turns 3: Call for papers for the 2025 BabyLM workshop
by: Charpentier, Lucas, et al.
Published: (2025)
by: Charpentier, Lucas, et al.
Published: (2025)
BabyLM Turns 4 and Goes Multilingual: Call for Papers for the 2026 BabyLM Workshop
by: Choshen, Leshem, et al.
Published: (2026)
by: Choshen, Leshem, et al.
Published: (2026)
Bias Dynamics in BabyLMs: Towards a Compute-Efficient Sandbox for Democratising Pre-Training Debiasing
by: Trhlik, Filip, et al.
Published: (2026)
by: Trhlik, Filip, et al.
Published: (2026)
Prompting Encoder Models for Zero-Shot Classification: A Cross-Domain Study in Italian
by: Auriemma, Serena, et al.
Published: (2024)
by: Auriemma, Serena, et al.
Published: (2024)
All-in-one: Understanding and Generation in Multimodal Reasoning with the MAIA Benchmark
by: Testa, Davide, et al.
Published: (2025)
by: Testa, Davide, et al.
Published: (2025)
Doing Things with Words: Rethinking Theory of Mind Simulation in Large Language Models
by: Lombardi, Agnese, et al.
Published: (2025)
by: Lombardi, Agnese, et al.
Published: (2025)
The quasi-semantic competence of LLMs: a case study on the part-whole relation
by: Proietti, Mattia, et al.
Published: (2025)
by: Proietti, Mattia, et al.
Published: (2025)
ExpliCa: Evaluating Explicit Causal Reasoning in Large Language Models
by: Miliani, Martina, et al.
Published: (2025)
by: Miliani, Martina, et al.
Published: (2025)
Bringing Up a Bilingual BabyLM: Investigating Multilingual Language Acquisition Using Small-Scale Models
by: Zeng, Linda, et al.
Published: (2026)
by: Zeng, Linda, et al.
Published: (2026)
BAMBINO-LM: (Bilingual-)Human-Inspired Continual Pretraining of BabyLM
by: Shen, Zhewen, et al.
Published: (2024)
by: Shen, Zhewen, et al.
Published: (2024)
Log Probabilities Are a Reliable Estimate of Semantic Plausibility in Base and Instruction-Tuned Language Models
by: Kauf, Carina, et al.
Published: (2024)
by: Kauf, Carina, et al.
Published: (2024)
BabyLM's First Constructions: Causal probing provides a signal of learning
by: Rozner, Joshua, et al.
Published: (2025)
by: Rozner, Joshua, et al.
Published: (2025)
Findings of the BabyLM Challenge: Sample-Efficient Pretraining on Developmentally Plausible Corpora
by: Warstadt, Alex, et al.
Published: (2025)
by: Warstadt, Alex, et al.
Published: (2025)
BabyLM's First Words: Word Segmentation as a Phonological Probing Task
by: Goriely, Zébulon, et al.
Published: (2025)
by: Goriely, Zébulon, et al.
Published: (2025)
BabyLM Challenge: Exploring the Effect of Variation Sets on Language Model Training Efficiency
by: Haga, Akari, et al.
Published: (2024)
by: Haga, Akari, et al.
Published: (2024)
Findings of the Second BabyLM Challenge: Sample-Efficient Pretraining on Developmentally Plausible Corpora
by: Hu, Michael Y., et al.
Published: (2024)
by: Hu, Michael Y., et al.
Published: (2024)
Dialogue Is Not Enough to Make a Communicative BabyLM (But Neither Is Developmentally Inspired Reinforcement Learning)
by: Padovani, Francesca, et al.
Published: (2025)
by: Padovani, Francesca, et al.
Published: (2025)
Teacher Demonstrations in a BabyLM's Zone of Proximal Development for Contingent Multi-Turn Interaction
by: Salhan, Suchir, et al.
Published: (2025)
by: Salhan, Suchir, et al.
Published: (2025)
[Call for Papers] The 2nd BabyLM Challenge: Sample-efficient pretraining on a developmentally plausible corpus
by: Choshen, Leshem, et al.
Published: (2024)
by: Choshen, Leshem, et al.
Published: (2024)
An Experimental Comparison of the Most Popular Approaches to Fake News Detection
by: Dell'Oglio, Pietro, et al.
Published: (2026)
by: Dell'Oglio, Pietro, et al.
Published: (2026)
Composing or Not Composing? Towards Distributional Construction Grammars
by: Blache, Philippe, et al.
Published: (2024)
by: Blache, Philippe, et al.
Published: (2024)
MathSpeech: Leveraging Small LMs for Accurate Conversion in Mathematical Speech-to-Formula
by: Hyeon, Sieun, et al.
Published: (2024)
by: Hyeon, Sieun, et al.
Published: (2024)
Enhancing Debunking Effectiveness through LLM-based Personality Adaptation
by: Dell'Oglio, Pietro, et al.
Published: (2026)
by: Dell'Oglio, Pietro, et al.
Published: (2026)
Probing for the Usage of Grammatical Number
by: Lasri, Karim, et al.
Published: (2022)
by: Lasri, Karim, et al.
Published: (2022)
Entropy-Aligned Decoding of LMs for Better Writing and Reasoning
by: Ahmed, Kareem, et al.
Published: (2026)
by: Ahmed, Kareem, et al.
Published: (2026)
HumanRankEval: Automatic Evaluation of LMs as Conversational Assistants
by: Gritta, Milan, et al.
Published: (2024)
by: Gritta, Milan, et al.
Published: (2024)
Safe to Serve: Aligning Instruction-Tuned Models for Safety and Helpfulness
by: Amballa, Avinash, et al.
Published: (2024)
by: Amballa, Avinash, et al.
Published: (2024)
Self-Refine Instruction-Tuning for Aligning Reasoning in Language Models
by: Ranaldi, Leonardo, et al.
Published: (2024)
by: Ranaldi, Leonardo, et al.
Published: (2024)
Charting a Decade of Computational Linguistics in Italy: The CLiC-it Corpus
by: Alzetta, Chiara, et al.
Published: (2025)
by: Alzetta, Chiara, et al.
Published: (2025)
Collocation in the Mind: Investigating Collocational Priming in Second Language Speakers of Italian
by: Irene Fioravanti, et al.
Published: (2024)
by: Irene Fioravanti, et al.
Published: (2024)
Reflect-RL: Two-Player Online RL Fine-Tuning for LMs
by: Zhou, Runlong, et al.
Published: (2024)
by: Zhou, Runlong, et al.
Published: (2024)
Dual-Track CoT: Budget-Aware Stepwise Guidance for Small LMs
by: Chatterjee, Sagnik, et al.
Published: (2026)
by: Chatterjee, Sagnik, et al.
Published: (2026)
Large-Scale Data Selection for Instruction Tuning
by: Ivison, Hamish, et al.
Published: (2025)
by: Ivison, Hamish, et al.
Published: (2025)
Similar Items
-
Are BabyLMs Second Language Learners?
by: Edman, Lukas, et al.
Published: (2024) -
BAMBI: Developing Baby Language Models for Italian
by: Suozzi, Alice, et al.
Published: (2025) -
Child-directed speech facilitates production, not comprehension, in BabyLMs
by: Bunzeck, Bastian, et al.
Published: (2026) -
Do Construction Distributions Shape Formal Language Learning In German BabyLMs?
by: Bunzeck, Bastian, et al.
Published: (2025) -
Mask and You Shall Receive: Optimizing Masked Language Modeling For Pretraining BabyLMs
by: Edman, Lukas, et al.
Published: (2025)