Gespeichert in:
| 1. Verfasser: | Rao, Praveen |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | https://arxiv.org/abs/2602.02632 |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Is Training Data Quality or Quantity More Impactful to Small Language Model Performance?
von: Sajith, Aryan, et al.
Veröffentlicht: (2024)
von: Sajith, Aryan, et al.
Veröffentlicht: (2024)
Adjusting Pretrained Backbones for Performativity
von: Demirel, Berker, et al.
Veröffentlicht: (2024)
von: Demirel, Berker, et al.
Veröffentlicht: (2024)
Reasoning Language Model Inference Serving Unveiled: An Empirical Study
von: Li, Qi, et al.
Veröffentlicht: (2025)
von: Li, Qi, et al.
Veröffentlicht: (2025)
Pretraining a Foundation Model for Small-Molecule Natural Products
von: Ding, Yuheng, et al.
Veröffentlicht: (2025)
von: Ding, Yuheng, et al.
Veröffentlicht: (2025)
LLM-Inspired Pretrain-Then-Finetune for Small-Data, Large-Scale Optimization
von: Zhang, Zishi, et al.
Veröffentlicht: (2026)
von: Zhang, Zishi, et al.
Veröffentlicht: (2026)
Finetune-Informed Pretraining Boosts Downstream Performance
von: Faysal, Atik, et al.
Veröffentlicht: (2026)
von: Faysal, Atik, et al.
Veröffentlicht: (2026)
Small Vocabularies, Big Gains: Pretraining and Tokenization in Time Series Models
von: Roger, Alexis, et al.
Veröffentlicht: (2025)
von: Roger, Alexis, et al.
Veröffentlicht: (2025)
Learning Transferable Sensor Models via Language-Informed Pretraining
von: Chen, Yuliang, et al.
Veröffentlicht: (2026)
von: Chen, Yuliang, et al.
Veröffentlicht: (2026)
How Does Controllability Emerge In Language Models During Pretraining?
von: She, Jianshu, et al.
Veröffentlicht: (2025)
von: She, Jianshu, et al.
Veröffentlicht: (2025)
Probing the Limits of Compressive Memory: A Study of Infini-Attention in Small-Scale Pretraining
von: Huang, Ruizhe, et al.
Veröffentlicht: (2025)
von: Huang, Ruizhe, et al.
Veröffentlicht: (2025)
On Training Large Language Models for Long-Horizon Tasks: An Empirical Study of Horizon Length
von: Kim, Sunghwan, et al.
Veröffentlicht: (2026)
von: Kim, Sunghwan, et al.
Veröffentlicht: (2026)
Why Representation Engineering Works: A Theoretical and Empirical Study in Vision-Language Models
von: Tian, Bowei, et al.
Veröffentlicht: (2025)
von: Tian, Bowei, et al.
Veröffentlicht: (2025)
A Step Toward Federated Pretraining of Multimodal Large Language Models
von: Xiong, Baochen, et al.
Veröffentlicht: (2026)
von: Xiong, Baochen, et al.
Veröffentlicht: (2026)
Modeling and Performance Analysis for Semantic Communications Based on Empirical Results
von: Ma, Shuai, et al.
Veröffentlicht: (2025)
von: Ma, Shuai, et al.
Veröffentlicht: (2025)
Pretraining Large Language Models with NVFP4
von: NVIDIA, et al.
Veröffentlicht: (2025)
von: NVIDIA, et al.
Veröffentlicht: (2025)
Patent Language Model Pretraining with ModernBERT
von: Yousefiramandi, Amirhossein, et al.
Veröffentlicht: (2025)
von: Yousefiramandi, Amirhossein, et al.
Veröffentlicht: (2025)
What Makes Quantization for Large Language Models Hard? An Empirical Study from the Lens of Perturbation
von: Gong, Zhuocheng, et al.
Veröffentlicht: (2024)
von: Gong, Zhuocheng, et al.
Veröffentlicht: (2024)
Enhancing Microgrid Performance Prediction with Attention-based Deep Learning Models
von: Maddineni, Vinod Kumar, et al.
Veröffentlicht: (2024)
von: Maddineni, Vinod Kumar, et al.
Veröffentlicht: (2024)
Dynamic Loss-Based Sample Reweighting for Improved Large Language Model Pretraining
von: Sow, Daouda, et al.
Veröffentlicht: (2025)
von: Sow, Daouda, et al.
Veröffentlicht: (2025)
Predicting LLM Reasoning Performance with Small Proxy Model
von: Koh, Woosung, et al.
Veröffentlicht: (2025)
von: Koh, Woosung, et al.
Veröffentlicht: (2025)
In-context Pretraining: Language Modeling Beyond Document Boundaries
von: Shi, Weijia, et al.
Veröffentlicht: (2023)
von: Shi, Weijia, et al.
Veröffentlicht: (2023)
Revisiting Multilingual Data Mixtures in Language Model Pretraining
von: Foroutan, Negar, et al.
Veröffentlicht: (2025)
von: Foroutan, Negar, et al.
Veröffentlicht: (2025)
Discovering Knowledge-Critical Subnetworks in Pretrained Language Models
von: Bayazit, Deniz, et al.
Veröffentlicht: (2023)
von: Bayazit, Deniz, et al.
Veröffentlicht: (2023)
Deep Ensembles Secretly Perform Empirical Bayes
von: Loaiza-Ganem, Gabriel, et al.
Veröffentlicht: (2025)
von: Loaiza-Ganem, Gabriel, et al.
Veröffentlicht: (2025)
Small Language Models for Application Interactions: A Case Study
von: Li, Beibin, et al.
Veröffentlicht: (2024)
von: Li, Beibin, et al.
Veröffentlicht: (2024)
Learnware of Language Models: Specialized Small Language Models Can Do Big
von: Tan, Zhi-Hao, et al.
Veröffentlicht: (2025)
von: Tan, Zhi-Hao, et al.
Veröffentlicht: (2025)
Caching Techniques for Reducing the Communication Cost of Federated Learning in IoT Environments
von: Alhonainy, Ahmad, et al.
Veröffentlicht: (2025)
von: Alhonainy, Ahmad, et al.
Veröffentlicht: (2025)
Robust Uncertainty Quantification for Self-Evolving Large Language Models via Continual Domain Pretraining
von: Zhou, Xiaofan, et al.
Veröffentlicht: (2025)
von: Zhou, Xiaofan, et al.
Veröffentlicht: (2025)
Retrieval Capabilities of Large Language Models Scale with Pretraining FLOPs
von: Portes, Jacob, et al.
Veröffentlicht: (2025)
von: Portes, Jacob, et al.
Veröffentlicht: (2025)
Teaching Pretrained Language Models to Think Deeper with Retrofitted Recurrence
von: McLeish, Sean, et al.
Veröffentlicht: (2025)
von: McLeish, Sean, et al.
Veröffentlicht: (2025)
Pretrained Vision-Language-Action Models are Surprisingly Resistant to Forgetting in Continual Learning
von: Liu, Huihan, et al.
Veröffentlicht: (2026)
von: Liu, Huihan, et al.
Veröffentlicht: (2026)
Multiple Physics Pretraining for Physical Surrogate Models
von: McCabe, Michael, et al.
Veröffentlicht: (2023)
von: McCabe, Michael, et al.
Veröffentlicht: (2023)
An Empirical Study of Realized GNN Expressiveness
von: Wang, Yanbo, et al.
Veröffentlicht: (2023)
von: Wang, Yanbo, et al.
Veröffentlicht: (2023)
Unlock the Potential of Large Language Models for Predictive Tabular Tasks in Data Science with Table-Specific Pretraining
von: Yang, Yazheng, et al.
Veröffentlicht: (2024)
von: Yang, Yazheng, et al.
Veröffentlicht: (2024)
Enhancing Two-Player Performance Through Single-Player Knowledge Transfer: An Empirical Study on Atari 2600 Games
von: Saadat, Kimiya, et al.
Veröffentlicht: (2024)
von: Saadat, Kimiya, et al.
Veröffentlicht: (2024)
Generalizable and Stable Finetuning of Pretrained Language Models on Low-Resource Texts
von: Somayajula, Sai Ashish, et al.
Veröffentlicht: (2024)
von: Somayajula, Sai Ashish, et al.
Veröffentlicht: (2024)
Cite Pretrain: Retrieval-Free Knowledge Attribution for Large Language Models
von: Huang, Yukun, et al.
Veröffentlicht: (2025)
von: Huang, Yukun, et al.
Veröffentlicht: (2025)
Improved Large Language Model Jailbreak Detection via Pretrained Embeddings
von: Galinkin, Erick, et al.
Veröffentlicht: (2024)
von: Galinkin, Erick, et al.
Veröffentlicht: (2024)
Tracing the Representation Geometry of Language Models from Pretraining to Post-training
von: Li, Melody Zixuan, et al.
Veröffentlicht: (2025)
von: Li, Melody Zixuan, et al.
Veröffentlicht: (2025)
FedNano: Toward Lightweight Federated Tuning for Pretrained Multimodal Large Language Models
von: Zhang, Yao, et al.
Veröffentlicht: (2025)
von: Zhang, Yao, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Is Training Data Quality or Quantity More Impactful to Small Language Model Performance?
von: Sajith, Aryan, et al.
Veröffentlicht: (2024) -
Adjusting Pretrained Backbones for Performativity
von: Demirel, Berker, et al.
Veröffentlicht: (2024) -
Reasoning Language Model Inference Serving Unveiled: An Empirical Study
von: Li, Qi, et al.
Veröffentlicht: (2025) -
Pretraining a Foundation Model for Small-Molecule Natural Products
von: Ding, Yuheng, et al.
Veröffentlicht: (2025) -
LLM-Inspired Pretrain-Then-Finetune for Small-Data, Large-Scale Optimization
von: Zhang, Zishi, et al.
Veröffentlicht: (2026)