Stacking Small Language Models for Generalizability
Fuente:
arXiv
Gespeichert in:
| 1. Verfasser: | Liang, Laurence |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Generalizable and Stable Finetuning of Pretrained Language Models on Low-Resource Texts
von: Somayajula, Sai Ashish, et al.
Veröffentlicht: (2024)
von: Somayajula, Sai Ashish, et al.
Veröffentlicht: (2024)
Large Language Models as Generalizable Policies for Embodied Tasks
von: Szot, Andrew, et al.
Veröffentlicht: (2023)
von: Szot, Andrew, et al.
Veröffentlicht: (2023)
Small LLMs Do Not Learn a Generalizable Theory of Mind via Reinforcement Learning
von: Sarangi, Sneheel, et al.
Veröffentlicht: (2025)
von: Sarangi, Sneheel, et al.
Veröffentlicht: (2025)
All Language Models Large and Small
von: Chen, Zhixun, et al.
Veröffentlicht: (2024)
von: Chen, Zhixun, et al.
Veröffentlicht: (2024)
Towards Understanding Multi-Round Large Language Model Reasoning: Approximability, Learnability and Generalizability
von: Xu, Chenhui, et al.
Veröffentlicht: (2025)
von: Xu, Chenhui, et al.
Veröffentlicht: (2025)
Kakugo: Distillation of Low-Resource Languages into Small Language Models
von: Devine, Peter, et al.
Veröffentlicht: (2026)
von: Devine, Peter, et al.
Veröffentlicht: (2026)
Small Language Models: Survey, Measurements, and Insights
von: Lu, Zhenyan, et al.
Veröffentlicht: (2024)
von: Lu, Zhenyan, et al.
Veröffentlicht: (2024)
Squat: Quant Small Language Models on the Edge
von: Shen, Xuan, et al.
Veröffentlicht: (2024)
von: Shen, Xuan, et al.
Veröffentlicht: (2024)
Towards Reasoning Ability of Small Language Models
von: Srivastava, Gaurav, et al.
Veröffentlicht: (2025)
von: Srivastava, Gaurav, et al.
Veröffentlicht: (2025)
Are Small Language Models Ready to Compete with Large Language Models for Practical Applications?
von: Sinha, Neelabh, et al.
Veröffentlicht: (2024)
von: Sinha, Neelabh, et al.
Veröffentlicht: (2024)
Light-IF: Endowing LLMs with Generalizable Reasoning via Preview and Self-Checking for Complex Instruction Following
von: Wang, Chenyang, et al.
Veröffentlicht: (2025)
von: Wang, Chenyang, et al.
Veröffentlicht: (2025)
RewardAnything: Generalizable Principle-Following Reward Models
von: Yu, Zhuohao, et al.
Veröffentlicht: (2025)
von: Yu, Zhuohao, et al.
Veröffentlicht: (2025)
Fox-1: Open Small Language Model for Cloud and Edge
von: Hu, Zijian, et al.
Veröffentlicht: (2024)
von: Hu, Zijian, et al.
Veröffentlicht: (2024)
Hymba: A Hybrid-head Architecture for Small Language Models
von: Dong, Xin, et al.
Veröffentlicht: (2024)
von: Dong, Xin, et al.
Veröffentlicht: (2024)
Small Language Models for Application Interactions: A Case Study
von: Li, Beibin, et al.
Veröffentlicht: (2024)
von: Li, Beibin, et al.
Veröffentlicht: (2024)
Technical Report: Small Language Model for Japanese Clinical and Medicine
von: Watanabe, Shogo
Veröffentlicht: (2024)
von: Watanabe, Shogo
Veröffentlicht: (2024)
Domain-Adapted Small Language Models for Reliable Clinical Triage
von: Aljohani, Manar, et al.
Veröffentlicht: (2026)
von: Aljohani, Manar, et al.
Veröffentlicht: (2026)
Weak-to-Strong Search: Align Large Language Models via Searching over Small Language Models
von: Zhou, Zhanhui, et al.
Veröffentlicht: (2024)
von: Zhou, Zhanhui, et al.
Veröffentlicht: (2024)
PARAMANU-GANITA: Can Small Math Language Models Rival with Large Language Models on Mathematical Reasoning?
von: Niyogi, Mitodru, et al.
Veröffentlicht: (2024)
von: Niyogi, Mitodru, et al.
Veröffentlicht: (2024)
Small Language Models for Privacy-Preserving Clinical Information Extraction in Low-Resource Languages
von: Ghaffarzadeh-Esfahani, Mohammadreza, et al.
Veröffentlicht: (2026)
von: Ghaffarzadeh-Esfahani, Mohammadreza, et al.
Veröffentlicht: (2026)
SmallToLarge (S2L): Scalable Data Selection for Fine-tuning Large Language Models by Summarizing Training Trajectories of Small Models
von: Yang, Yu, et al.
Veröffentlicht: (2024)
von: Yang, Yu, et al.
Veröffentlicht: (2024)
A Post-Training Enhanced Optimization Approach for Small Language Models
von: Zhai, Keke
Veröffentlicht: (2024)
von: Zhai, Keke
Veröffentlicht: (2024)
Mini-Giants: "Small" Language Models and Open Source Win-Win
von: Zhou, Zhengping, et al.
Veröffentlicht: (2023)
von: Zhou, Zhengping, et al.
Veröffentlicht: (2023)
PMPO: Probabilistic Metric Prompt Optimization for Small and Large Language Models
von: Zhao, Chenzhuo, et al.
Veröffentlicht: (2025)
von: Zhao, Chenzhuo, et al.
Veröffentlicht: (2025)
Luna-2: Scalable Single-Token Evaluation with Small Language Models
von: Goel, Vatsal, et al.
Veröffentlicht: (2026)
von: Goel, Vatsal, et al.
Veröffentlicht: (2026)
EffGen: Enabling Small Language Models as Capable Autonomous Agents
von: Srivastava, Gaurav, et al.
Veröffentlicht: (2026)
von: Srivastava, Gaurav, et al.
Veröffentlicht: (2026)
Scaling Smart: Accelerating Large Language Model Pre-training with Small Model Initialization
von: Samragh, Mohammad, et al.
Veröffentlicht: (2024)
von: Samragh, Mohammad, et al.
Veröffentlicht: (2024)
CaLM: Contrasting Large and Small Language Models to Verify Grounded Generation
von: Hsu, I-Hung, et al.
Veröffentlicht: (2024)
von: Hsu, I-Hung, et al.
Veröffentlicht: (2024)
Is Training Data Quality or Quantity More Impactful to Small Language Model Performance?
von: Sajith, Aryan, et al.
Veröffentlicht: (2024)
von: Sajith, Aryan, et al.
Veröffentlicht: (2024)
Readability $\ne$ Learnability: Rethinking the Role of Simplicity in Training Small Language Models
von: Lee, Ivan, et al.
Veröffentlicht: (2025)
von: Lee, Ivan, et al.
Veröffentlicht: (2025)
Self-Evolved Preference Optimization for Enhancing Mathematical Reasoning in Small Language Models
von: Singh, Joykirat, et al.
Veröffentlicht: (2025)
von: Singh, Joykirat, et al.
Veröffentlicht: (2025)
Adapt-Pruner: Adaptive Structural Pruning for Efficient Small Language Model Training
von: Pan, Rui, et al.
Veröffentlicht: (2025)
von: Pan, Rui, et al.
Veröffentlicht: (2025)
VocabTailor: Dynamic Vocabulary Selection for Downstream Tasks in Small Language Models
von: Zhang, Hanling, et al.
Veröffentlicht: (2025)
von: Zhang, Hanling, et al.
Veröffentlicht: (2025)
Meta-Tool: Efficient Few-Shot Tool Adaptation for Small Language Models
von: Kumar, Sachin
Veröffentlicht: (2026)
von: Kumar, Sachin
Veröffentlicht: (2026)
PlaSma: Making Small Language Models Better Procedural Knowledge Models for (Counterfactual) Planning
von: Brahman, Faeze, et al.
Veröffentlicht: (2023)
von: Brahman, Faeze, et al.
Veröffentlicht: (2023)
Limited Generalizability in Argument Mining: State-Of-The-Art Models Learn Datasets, Not Arguments
von: Feger, Marc, et al.
Veröffentlicht: (2025)
von: Feger, Marc, et al.
Veröffentlicht: (2025)
Accurate Prediction of Ligand-Protein Interaction Affinities with Fine-Tuned Small Language Models
von: Fauber, Ben
Veröffentlicht: (2024)
von: Fauber, Ben
Veröffentlicht: (2024)
When Correct Isn't Usable: Improving Structured Output Reliability in Small Language Models
von: Galeone, Cosimo, et al.
Veröffentlicht: (2026)
von: Galeone, Cosimo, et al.
Veröffentlicht: (2026)
On the Entropy Calibration of Language Models
von: Cao, Steven, et al.
Veröffentlicht: (2025)
von: Cao, Steven, et al.
Veröffentlicht: (2025)
PhoneLM:an Efficient and Capable Small Language Model Family through Principled Pre-training
von: Yi, Rongjie, et al.
Veröffentlicht: (2024)
von: Yi, Rongjie, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Generalizable and Stable Finetuning of Pretrained Language Models on Low-Resource Texts
von: Somayajula, Sai Ashish, et al.
Veröffentlicht: (2024) -
Large Language Models as Generalizable Policies for Embodied Tasks
von: Szot, Andrew, et al.
Veröffentlicht: (2023) -
Small LLMs Do Not Learn a Generalizable Theory of Mind via Reinforcement Learning
von: Sarangi, Sneheel, et al.
Veröffentlicht: (2025) -
All Language Models Large and Small
von: Chen, Zhixun, et al.
Veröffentlicht: (2024) -
Towards Understanding Multi-Round Large Language Model Reasoning: Approximability, Learnability and Generalizability
von: Xu, Chenhui, et al.
Veröffentlicht: (2025)