Pico: A Modular Framework for Hypothesis-Driven Small Language Model Research
Fuente:
arXiv
Guardado en:
| Autores principales: | Martinez, Richard Diehl, Africa, David Demitri, Weiss, Yuval, Salhan, Suchir, Daniels, Ryan, Buttery, Paula |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Meta-Pretraining for Zero-Shot Cross-Lingual Named Entity Recognition in Low-Resource Philippine Languages
por: Africa, David Demitri, et al.
Publicado: (2025)
por: Africa, David Demitri, et al.
Publicado: (2025)
Investigating ReLoRA: Effects on the Learning Dynamics of Small Language Models
por: Weiss, Yuval, et al.
Publicado: (2025)
por: Weiss, Yuval, et al.
Publicado: (2025)
Learning Dynamics of Meta-Learning in Small Model Pretraining
por: Africa, David Demitri, et al.
Publicado: (2025)
por: Africa, David Demitri, et al.
Publicado: (2025)
Less is More: Pre-Training Cross-Lingual Small-Scale Language Models with Cognitively-Plausible Curriculum Learning Strategies
por: Salhan, Suchir, et al.
Publicado: (2024)
por: Salhan, Suchir, et al.
Publicado: (2024)
What is the Best Sequence Length for BABYLM?
por: Salhan, Suchir, et al.
Publicado: (2025)
por: Salhan, Suchir, et al.
Publicado: (2025)
Looking to Learn: Token-wise Dynamic Gating for Low-Resource Vision-Language Modelling
por: Ganescu, Bianca-Mihaela, et al.
Publicado: (2025)
por: Ganescu, Bianca-Mihaela, et al.
Publicado: (2025)
BLiSS 1.0: Evaluating Bilingual Learner Competence in Second Language Small Language Models
por: Gao, Yuan, et al.
Publicado: (2025)
por: Gao, Yuan, et al.
Publicado: (2025)
ByteSpan: Information-Driven Subword Tokenisation
por: Goriely, Zébulon, et al.
Publicado: (2025)
por: Goriely, Zébulon, et al.
Publicado: (2025)
Tending Towards Stability: Convergence Challenges in Small Language Models
por: Martinez, Richard Diehl, et al.
Publicado: (2024)
por: Martinez, Richard Diehl, et al.
Publicado: (2024)
LURE: Live-Usage Replay Evaluations for Reducing Evaluation Awareness
por: Ivanov, Igor, et al.
Publicado: (2026)
por: Ivanov, Igor, et al.
Publicado: (2026)
Steering Awareness: Detecting Activation Steering from Within
por: Rivera, Joshua Fonseca, et al.
Publicado: (2025)
por: Rivera, Joshua Fonseca, et al.
Publicado: (2025)
The Distribution of Phoneme Frequencies across the World's Languages: Macroscopic and Microscopic Information-Theoretic Models
por: Martín, Fermín Moscoso del Prado, et al.
Publicado: (2026)
por: Martín, Fermín Moscoso del Prado, et al.
Publicado: (2026)
Identifying a Circuit for Verb Conjugation in GPT-2
por: Africa, David Demitri
Publicado: (2025)
por: Africa, David Demitri
Publicado: (2025)
Modelling the Diachronic Emergence of Phoneme Frequency Distributions
por: Martín, Fermín Moscoso del Prado, et al.
Publicado: (2026)
por: Martín, Fermín Moscoso del Prado, et al.
Publicado: (2026)
A Computational Operationalisation of Competing Maturational Theories of Syntactic Development via Statistical Grammar Induction
por: Marcheva, Mila, et al.
Publicado: (2026)
por: Marcheva, Mila, et al.
Publicado: (2026)
Consistency Training while Mitigating Obfuscation via Rate Matching
por: Imran, Sohaib, et al.
Publicado: (2026)
por: Imran, Sohaib, et al.
Publicado: (2026)
Teacher Demonstrations in a BabyLM's Zone of Proximal Development for Contingent Multi-Turn Interaction
por: Salhan, Suchir, et al.
Publicado: (2025)
por: Salhan, Suchir, et al.
Publicado: (2025)
Does Self-Evaluation Enable Wireheading in Language Models?
por: Africa, David Demitri, et al.
Publicado: (2025)
por: Africa, David Demitri, et al.
Publicado: (2025)
Mitigating Frequency Bias and Anisotropy in Language Model Pre-Training with Syntactic Smoothing
por: Martinez, Richard Diehl, et al.
Publicado: (2024)
por: Martinez, Richard Diehl, et al.
Publicado: (2024)
From Babble to Words: Pre-Training Language Models on Continuous Streams of Phonemes
por: Goriely, Zébulon, et al.
Publicado: (2024)
por: Goriely, Zébulon, et al.
Publicado: (2024)
Bias Dynamics in BabyLMs: Towards a Compute-Efficient Sandbox for Democratising Pre-Training Debiasing
por: Trhlik, Filip, et al.
Publicado: (2026)
por: Trhlik, Filip, et al.
Publicado: (2026)
Learning Modular Exponentiation with Transformers
por: Africa, David Demitri, et al.
Publicado: (2025)
por: Africa, David Demitri, et al.
Publicado: (2025)
Batayan: A Filipino NLP benchmark for evaluating Large Language Models
por: Montalan, Jann Railey, et al.
Publicado: (2025)
por: Montalan, Jann Railey, et al.
Publicado: (2025)
No Answer Needed: Predicting LLM Answer Accuracy from Question-Only Linear Probes
por: Cencerrado, Iván Vicente Moreno, et al.
Publicado: (2025)
por: Cencerrado, Iván Vicente Moreno, et al.
Publicado: (2025)
Inoculation Prompting: Eliciting traits from LLMs during training can suppress them at test-time
por: Tan, Daniel, et al.
Publicado: (2025)
por: Tan, Daniel, et al.
Publicado: (2025)
IPA-CHILDES & G2P+: Feature-Rich Resources for Cross-Lingual Phonology and Phonemic Language Modeling
por: Goriely, Zébulon, et al.
Publicado: (2025)
por: Goriely, Zébulon, et al.
Publicado: (2025)
Hypothesis-Driven Theory-of-Mind Reasoning for Large Language Models
por: Kim, Hyunwoo, et al.
Publicado: (2025)
por: Kim, Hyunwoo, et al.
Publicado: (2025)
BabyLM's First Words: Word Segmentation as a Phonological Probing Task
por: Goriely, Zébulon, et al.
Publicado: (2025)
por: Goriely, Zébulon, et al.
Publicado: (2025)
A Large-Language Model Framework for Relative Timeline Extraction from PubMed Case Reports
por: Wang, Jing, et al.
Publicado: (2025)
por: Wang, Jing, et al.
Publicado: (2025)
From Hypothesis to Publication: A Comprehensive Survey of AI-Driven Research Support Systems
por: Zhou, Zekun, et al.
Publicado: (2025)
por: Zhou, Zekun, et al.
Publicado: (2025)
Linear Representation Transferability Hypothesis: Leveraging Small Models to Steer Large Models
por: Bello, Femi, et al.
Publicado: (2025)
por: Bello, Femi, et al.
Publicado: (2025)
Enhancing Vaccine Safety Surveillance: Extracting Vaccine Mentions from Emergency Department Triage Notes Using Fine-Tuned Large Language Models
por: Khademi, Sedigh, et al.
Publicado: (2025)
por: Khademi, Sedigh, et al.
Publicado: (2025)
FreeEval: A Modular Framework for Trustworthy and Efficient Evaluation of Large Language Models
por: Yu, Zhuohao, et al.
Publicado: (2024)
por: Yu, Zhuohao, et al.
Publicado: (2024)
Federated Co-tuning Framework for Large and Small Language Models
por: Fan, Tao, et al.
Publicado: (2024)
por: Fan, Tao, et al.
Publicado: (2024)
Exploring the Knowledge Mismatch Hypothesis: Hallucination Propensity in Small Models Fine-tuned on Data from Larger Models
por: Wee, Phil, et al.
Publicado: (2024)
por: Wee, Phil, et al.
Publicado: (2024)
Hypothesis Search: Inductive Reasoning with Language Models
por: Wang, Ruocheng, et al.
Publicado: (2023)
por: Wang, Ruocheng, et al.
Publicado: (2023)
Improving Mathematical Reasoning Capabilities of Small Language Models via Feedback-Driven Distillation
por: Zhu, Xunyu, et al.
Publicado: (2024)
por: Zhu, Xunyu, et al.
Publicado: (2024)
Improving Scientific Hypothesis Generation with Knowledge Grounded Large Language Models
por: Xiong, Guangzhi, et al.
Publicado: (2024)
por: Xiong, Guangzhi, et al.
Publicado: (2024)
Hypothesis Testing Prompting Improves Deductive Reasoning in Large Language Models
por: Li, Yitian, et al.
Publicado: (2024)
por: Li, Yitian, et al.
Publicado: (2024)
Hypothesis Generation with Large Language Models
por: Zhou, Yangqiaoyu, et al.
Publicado: (2024)
por: Zhou, Yangqiaoyu, et al.
Publicado: (2024)
Ejemplares similares
-
Meta-Pretraining for Zero-Shot Cross-Lingual Named Entity Recognition in Low-Resource Philippine Languages
por: Africa, David Demitri, et al.
Publicado: (2025) -
Investigating ReLoRA: Effects on the Learning Dynamics of Small Language Models
por: Weiss, Yuval, et al.
Publicado: (2025) -
Learning Dynamics of Meta-Learning in Small Model Pretraining
por: Africa, David Demitri, et al.
Publicado: (2025) -
Less is More: Pre-Training Cross-Lingual Small-Scale Language Models with Cognitively-Plausible Curriculum Learning Strategies
por: Salhan, Suchir, et al.
Publicado: (2024) -
What is the Best Sequence Length for BABYLM?
por: Salhan, Suchir, et al.
Publicado: (2025)