When Babies Teach Babies: Can student knowledge sharing outperform Teacher-Guided Distillation on small datasets?
Fuente:
arXiv
Gespeichert in:
| 1. Verfasser: | Iyer, Srikrishna |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Mini Minds: Exploring Bebeshka and Zlata Baby Models
von: Proskurina, Irina, et al.
Veröffentlicht: (2023)
von: Proskurina, Irina, et al.
Veröffentlicht: (2023)
What Should Baby Models Read? Exploring Sample-Efficient Data Composition on Model Performance
von: Yam, Hong Meng, et al.
Veröffentlicht: (2024)
von: Yam, Hong Meng, et al.
Veröffentlicht: (2024)
Baby Scale: Investigating Models Trained on Individual Children's Language Input
von: Feng, Steven Y., et al.
Veröffentlicht: (2026)
von: Feng, Steven Y., et al.
Veröffentlicht: (2026)
BabyReasoningBench: Generating Developmentally-Inspired Reasoning Tasks for Evaluating Baby Language Models
von: Dhole, Kaustubh D.
Veröffentlicht: (2026)
von: Dhole, Kaustubh D.
Veröffentlicht: (2026)
BabyVLM: Data-Efficient Pretraining of VLMs Inspired by Infant Learning
von: Wang, Shengao, et al.
Veröffentlicht: (2025)
von: Wang, Shengao, et al.
Veröffentlicht: (2025)
Bias Dynamics in BabyLMs: Towards a Compute-Efficient Sandbox for Democratising Pre-Training Debiasing
von: Trhlik, Filip, et al.
Veröffentlicht: (2026)
von: Trhlik, Filip, et al.
Veröffentlicht: (2026)
Bringing Up a Bilingual BabyLM: Investigating Multilingual Language Acquisition Using Small-Scale Models
von: Zeng, Linda, et al.
Veröffentlicht: (2026)
von: Zeng, Linda, et al.
Veröffentlicht: (2026)
EgoBabyVLM: Benchmarking Cross-Modal Learning from Naturalistic Egocentric Video Data
von: Lin, Dongyan, et al.
Veröffentlicht: (2026)
von: Lin, Dongyan, et al.
Veröffentlicht: (2026)
Auditing Google's AI Overviews and Featured Snippets: A Case Study on Baby Care and Pregnancy
von: Hu, Desheng, et al.
Veröffentlicht: (2025)
von: Hu, Desheng, et al.
Veröffentlicht: (2025)
Teaching-Assistant-in-the-Loop: Improving Knowledge Distillation from Imperfect Teacher Models in Low-Budget Scenarios
von: Zhou, Yuhang, et al.
Veröffentlicht: (2024)
von: Zhou, Yuhang, et al.
Veröffentlicht: (2024)
Your Teacher Can't Help You Here: Combating Supervision Fidelity Decay in On-Policy Distillation
von: Liu, Yanjiang, et al.
Veröffentlicht: (2026)
von: Liu, Yanjiang, et al.
Veröffentlicht: (2026)
Don't Kill the Baby: The Case for AI in Arbitration
von: Broyde, Michael, et al.
Veröffentlicht: (2024)
von: Broyde, Michael, et al.
Veröffentlicht: (2024)
Multi-agent AI systems outperform human teams in creativity
von: Hu, Tiancheng, et al.
Veröffentlicht: (2026)
von: Hu, Tiancheng, et al.
Veröffentlicht: (2026)
Decoding with Limited Teacher Supervision Requires Understanding When to Trust the Teacher
von: Ok, Hyunjong, et al.
Veröffentlicht: (2024)
von: Ok, Hyunjong, et al.
Veröffentlicht: (2024)
BabyLM Turns 3: Call for papers for the 2025 BabyLM workshop
von: Charpentier, Lucas, et al.
Veröffentlicht: (2025)
von: Charpentier, Lucas, et al.
Veröffentlicht: (2025)
Code-enabled language models can outperform reasoning models on diverse tasks
von: Zhang, Cedegao E., et al.
Veröffentlicht: (2025)
von: Zhang, Cedegao E., et al.
Veröffentlicht: (2025)
BabyLlama-2: Ensemble-Distilled Models Consistently Outperform Teachers With Limited Data
von: Tastet, Jean-Loup, et al.
Veröffentlicht: (2024)
von: Tastet, Jean-Loup, et al.
Veröffentlicht: (2024)
ASTRO: Teaching Language Models to Reason by Reflecting and Backtracking In-Context
von: Kim, Joongwon, et al.
Veröffentlicht: (2025)
von: Kim, Joongwon, et al.
Veröffentlicht: (2025)
MMG2Skill: Can Agents Distill In-the-Wild Guides into Self-Evolving Skills?
von: Che, Xinyu, et al.
Veröffentlicht: (2026)
von: Che, Xinyu, et al.
Veröffentlicht: (2026)
LLMs Can Teach Themselves to Better Predict the Future
von: Turtel, Benjamin, et al.
Veröffentlicht: (2025)
von: Turtel, Benjamin, et al.
Veröffentlicht: (2025)
Reliability Gated Multi-Teacher Distillation for Low Resource Abstractive Summarization
von: Sumit, Dipto, et al.
Veröffentlicht: (2026)
von: Sumit, Dipto, et al.
Veröffentlicht: (2026)
When Perplexity Lies: Generation-Focused Distillation of Hybrid Sequence Models
von: Kostelec, Juan Gabriel, et al.
Veröffentlicht: (2026)
von: Kostelec, Juan Gabriel, et al.
Veröffentlicht: (2026)
On Teacher Hacking in Language Model Distillation
von: Tiapkin, Daniil, et al.
Veröffentlicht: (2025)
von: Tiapkin, Daniil, et al.
Veröffentlicht: (2025)
Can Large Models Teach Student Models to Solve Mathematical Problems Like Human Beings? A Reasoning Distillation Method via Multi-LoRA Interaction
von: Li, Xinhe, et al.
Veröffentlicht: (2025)
von: Li, Xinhe, et al.
Veröffentlicht: (2025)
Can postgraduate translation students identify machine-generated text?
von: Farrell, Michael
Veröffentlicht: (2025)
von: Farrell, Michael
Veröffentlicht: (2025)
CLARity: Reasoning Consistency Alone Can Teach Reinforced Experts
von: Lin, Jiuheng, et al.
Veröffentlicht: (2025)
von: Lin, Jiuheng, et al.
Veröffentlicht: (2025)
BabyLM Turns 4 and Goes Multilingual: Call for Papers for the 2026 BabyLM Workshop
von: Choshen, Leshem, et al.
Veröffentlicht: (2026)
von: Choshen, Leshem, et al.
Veröffentlicht: (2026)
Speculative Knowledge Distillation: Bridging the Teacher-Student Gap Through Interleaved Sampling
von: Xu, Wenda, et al.
Veröffentlicht: (2024)
von: Xu, Wenda, et al.
Veröffentlicht: (2024)
Backtracking When It Strays: Mitigating Dual Exposure Biases in LLM Reasoning Distillation
von: Wang, Bing, et al.
Veröffentlicht: (2026)
von: Wang, Bing, et al.
Veröffentlicht: (2026)
Can LLMs Learn by Teaching for Better Reasoning? A Preliminary Study
von: Ning, Xuefei, et al.
Veröffentlicht: (2024)
von: Ning, Xuefei, et al.
Veröffentlicht: (2024)
Can RL Teach Long-Horizon Reasoning to LLMs? Expressiveness Is Key
von: Wang, Tianle, et al.
Veröffentlicht: (2026)
von: Wang, Tianle, et al.
Veröffentlicht: (2026)
Learning from Committee: Reasoning Distillation from a Mixture of Teachers with Peer-Review
von: Li, Zhuochun, et al.
Veröffentlicht: (2024)
von: Li, Zhuochun, et al.
Veröffentlicht: (2024)
Steering LLMs? Actually, Sparse Autoencoders can outperform simple baselines
von: Jørgensen, Mikkel Godsk, et al.
Veröffentlicht: (2026)
von: Jørgensen, Mikkel Godsk, et al.
Veröffentlicht: (2026)
ELAD: Explanation-Guided Large Language Models Active Distillation
von: Zhang, Yifei, et al.
Veröffentlicht: (2024)
von: Zhang, Yifei, et al.
Veröffentlicht: (2024)
TeacherLM: Teaching to Fish Rather Than Giving the Fish, Language Modeling Likewise
von: He, Nan, et al.
Veröffentlicht: (2023)
von: He, Nan, et al.
Veröffentlicht: (2023)
When Can Transformers Count to n?
von: Yehudai, Gilad, et al.
Veröffentlicht: (2024)
von: Yehudai, Gilad, et al.
Veröffentlicht: (2024)
When Silence Is Golden: Can LLMs Learn to Abstain in Temporal QA and Beyond?
von: Zhou, Xinyu, et al.
Veröffentlicht: (2026)
von: Zhou, Xinyu, et al.
Veröffentlicht: (2026)
From Fallback to Frontline: When Can LLMs be Superior Annotators of Human Perspectives?
von: Amin, Hasan, et al.
Veröffentlicht: (2026)
von: Amin, Hasan, et al.
Veröffentlicht: (2026)
ReAD: Reinforcement-Guided Capability Distillation for Large Language Models
von: Cheng, Xueqi, et al.
Veröffentlicht: (2026)
von: Cheng, Xueqi, et al.
Veröffentlicht: (2026)
Distilling to Hybrid Attention Models via KL-Guided Layer Selection
von: Li, Yanhong, et al.
Veröffentlicht: (2025)
von: Li, Yanhong, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Mini Minds: Exploring Bebeshka and Zlata Baby Models
von: Proskurina, Irina, et al.
Veröffentlicht: (2023) -
What Should Baby Models Read? Exploring Sample-Efficient Data Composition on Model Performance
von: Yam, Hong Meng, et al.
Veröffentlicht: (2024) -
Baby Scale: Investigating Models Trained on Individual Children's Language Input
von: Feng, Steven Y., et al.
Veröffentlicht: (2026) -
BabyReasoningBench: Generating Developmentally-Inspired Reasoning Tasks for Evaluating Baby Language Models
von: Dhole, Kaustubh D.
Veröffentlicht: (2026) -
BabyVLM: Data-Efficient Pretraining of VLMs Inspired by Infant Learning
von: Wang, Shengao, et al.
Veröffentlicht: (2025)