Beyond URLs: Metadata Diversity and Position for Efficient LLM Pretraining
Fuente:
arXiv
Salvato in:
| Autori principali: | Fan, Dongyang, Hashemi, Diba, Karimireddy, Sai Praneeth, Jaggi, Martin |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
URLs Help, Topics Guide: Understanding Metadata Utility in LLM Training
di: Fan, Dongyang, et al.
Pubblicazione: (2025)
di: Fan, Dongyang, et al.
Pubblicazione: (2025)
Towards an empirical understanding of MoE design choices
di: Fan, Dongyang, et al.
Pubblicazione: (2024)
di: Fan, Dongyang, et al.
Pubblicazione: (2024)
Communication-Efficient Heterogeneous Federated Learning with Generalized Heavy-Ball Momentum
di: Zaccone, Riccardo, et al.
Pubblicazione: (2023)
di: Zaccone, Riccardo, et al.
Pubblicazione: (2023)
EPSVec: Efficient and Private Synthetic Data Generation via Dataset Vectors
di: Banayeeanzade, Amin, et al.
Pubblicazione: (2026)
di: Banayeeanzade, Amin, et al.
Pubblicazione: (2026)
DoGE: Domain Reweighting with Generalization Estimation
di: Fan, Simin, et al.
Pubblicazione: (2023)
di: Fan, Simin, et al.
Pubblicazione: (2023)
On the Limits of Momentum in Decentralized and Federated Optimization
di: Zaccone, Riccardo, et al.
Pubblicazione: (2025)
di: Zaccone, Riccardo, et al.
Pubblicazione: (2025)
Do Data Valuations Make Good Data Prices?
di: Fan, Dongyang, et al.
Pubblicazione: (2025)
di: Fan, Dongyang, et al.
Pubblicazione: (2025)
Tabular Embeddings for Tables with Bi-Dimensional Hierarchical Metadata and Nesting
di: Shrestha, Gyanendra, et al.
Pubblicazione: (2025)
di: Shrestha, Gyanendra, et al.
Pubblicazione: (2025)
Personalized Collaborative Fine-Tuning for On-Device Large Language Models
di: Wagner, Nicolas, et al.
Pubblicazione: (2024)
di: Wagner, Nicolas, et al.
Pubblicazione: (2024)
ABBA-Adapters: Efficient and Expressive Fine-Tuning of Foundation Models
di: Singhal, Raghav, et al.
Pubblicazione: (2025)
di: Singhal, Raghav, et al.
Pubblicazione: (2025)
Block-Attention for Efficient Prefilling
di: Ma, Dongyang, et al.
Pubblicazione: (2024)
di: Ma, Dongyang, et al.
Pubblicazione: (2024)
A Systematic Analysis of Base Model Choice for Reward Modeling
di: Ahrabian, Kian, et al.
Pubblicazione: (2025)
di: Ahrabian, Kian, et al.
Pubblicazione: (2025)
In-context Pretraining: Language Modeling Beyond Document Boundaries
di: Shi, Weijia, et al.
Pubblicazione: (2023)
di: Shi, Weijia, et al.
Pubblicazione: (2023)
Enhancing Multilingual LLM Pretraining with Model-Based Data Selection
di: Messmer, Bettina, et al.
Pubblicazione: (2025)
di: Messmer, Bettina, et al.
Pubblicazione: (2025)
LIA: Privacy-Preserving Data Quality Evaluation in Federated Learning Using a Lazy Influence Approximation
di: Rokvic, Ljubomir, et al.
Pubblicazione: (2022)
di: Rokvic, Ljubomir, et al.
Pubblicazione: (2022)
Pretrained Hybrids with MAD Skills
di: Roberts, Nicholas, et al.
Pubblicazione: (2024)
di: Roberts, Nicholas, et al.
Pubblicazione: (2024)
Output Embedding Centering for Stable LLM Pretraining
di: Stollenwerk, Felix, et al.
Pubblicazione: (2026)
di: Stollenwerk, Felix, et al.
Pubblicazione: (2026)
CoBo: Collaborative Learning via Bilevel Optimization
di: Hashemi, Diba, et al.
Pubblicazione: (2024)
di: Hashemi, Diba, et al.
Pubblicazione: (2024)
Ghosted Layers: Unconstrained Activation Alignment for Recovering Layer-Pruned LLMs
di: Yun, Vincent-Daniel, et al.
Pubblicazione: (2026)
di: Yun, Vincent-Daniel, et al.
Pubblicazione: (2026)
LML-DAP: Language Model Learning a Dataset for Data-Augmented Prediction
di: Vadlapati, Praneeth
Pubblicazione: (2024)
di: Vadlapati, Praneeth
Pubblicazione: (2024)
Generalizable and Stable Finetuning of Pretrained Language Models on Low-Resource Texts
di: Somayajula, Sai Ashish, et al.
Pubblicazione: (2024)
di: Somayajula, Sai Ashish, et al.
Pubblicazione: (2024)
Reconsidering LLM Uncertainty Estimation Methods in the Wild
di: Bakman, Yavuz, et al.
Pubblicazione: (2025)
di: Bakman, Yavuz, et al.
Pubblicazione: (2025)
Position-Aware Parameter Efficient Fine-Tuning Approach for Reducing Positional Bias in LLMs
di: Zhang, Zheng, et al.
Pubblicazione: (2024)
di: Zhang, Zheng, et al.
Pubblicazione: (2024)
On-Device Collaborative Language Modeling via a Mixture of Generalists and Specialists
di: Fan, Dongyang, et al.
Pubblicazione: (2024)
di: Fan, Dongyang, et al.
Pubblicazione: (2024)
Initialization using Update Approximation is a Silver Bullet for Extremely Efficient Low-Rank Fine-Tuning
di: Ponkshe, Kaustubh, et al.
Pubblicazione: (2024)
di: Ponkshe, Kaustubh, et al.
Pubblicazione: (2024)
Conformal Prediction Adaptive to Unknown Subpopulation Shifts
di: Wang, Nien-Shao, et al.
Pubblicazione: (2025)
di: Wang, Nien-Shao, et al.
Pubblicazione: (2025)
f-INE: A Hypothesis Testing Framework for Estimating Influence under Training Randomness
di: Panda, Subhodip, et al.
Pubblicazione: (2025)
di: Panda, Subhodip, et al.
Pubblicazione: (2025)
A Closer Look at Personalized Fine-Tuning in Heterogeneous Federated Learning
di: Chen, Minghui, et al.
Pubblicazione: (2025)
di: Chen, Minghui, et al.
Pubblicazione: (2025)
Mitigating Unintended Memorization with LoRA in Federated Learning for LLMs
di: Bossy, Thierry, et al.
Pubblicazione: (2025)
di: Bossy, Thierry, et al.
Pubblicazione: (2025)
Maximize Your Data's Potential: Enhancing LLM Accuracy with Two-Phase Pretraining
di: Feng, Steven, et al.
Pubblicazione: (2024)
di: Feng, Steven, et al.
Pubblicazione: (2024)
CDQuant: Greedy Coordinate Descent for Accurate LLM Quantization
di: Nair, Pranav Ajit, et al.
Pubblicazione: (2024)
di: Nair, Pranav Ajit, et al.
Pubblicazione: (2024)
LEAP: Layer-wise Exit-Aware Pretraining for Efficient Transformer Inference
di: Kapadia, Shashank, et al.
Pubblicazione: (2026)
di: Kapadia, Shashank, et al.
Pubblicazione: (2026)
DP-Fusion: Token-Level Differentially Private Inference for Large Language Models
di: Thareja, Rushil, et al.
Pubblicazione: (2025)
di: Thareja, Rushil, et al.
Pubblicazione: (2025)
Safety Subspaces are Not Linearly Distinct: A Fine-Tuning Case Study
di: Ponkshe, Kaustubh, et al.
Pubblicazione: (2025)
di: Ponkshe, Kaustubh, et al.
Pubblicazione: (2025)
Crosscoding Through Time: Tracking Emergence & Consolidation Of Linguistic Representations Throughout LLM Pretraining
di: Bayazit, Deniz, et al.
Pubblicazione: (2025)
di: Bayazit, Deniz, et al.
Pubblicazione: (2025)
Train Once, Answer All: Many Pretraining Experiments for the Cost of One
di: Bordt, Sebastian, et al.
Pubblicazione: (2025)
di: Bordt, Sebastian, et al.
Pubblicazione: (2025)
Semantic uncertainty in advanced decoding methods for LLM generation
di: Foodeei, Darius, et al.
Pubblicazione: (2025)
di: Foodeei, Darius, et al.
Pubblicazione: (2025)
Addressing LLM Diversity by Infusing Random Concepts
di: Agrawal, Pulin, et al.
Pubblicazione: (2026)
di: Agrawal, Pulin, et al.
Pubblicazione: (2026)
Understanding the Effects of RLHF on LLM Generalisation and Diversity
di: Kirk, Robert, et al.
Pubblicazione: (2023)
di: Kirk, Robert, et al.
Pubblicazione: (2023)
How Learning Rate Decay Wastes Your Best Data in Curriculum-Based LLM Pretraining
di: Luo, Kairong, et al.
Pubblicazione: (2025)
di: Luo, Kairong, et al.
Pubblicazione: (2025)
Documenti analoghi
-
URLs Help, Topics Guide: Understanding Metadata Utility in LLM Training
di: Fan, Dongyang, et al.
Pubblicazione: (2025) -
Towards an empirical understanding of MoE design choices
di: Fan, Dongyang, et al.
Pubblicazione: (2024) -
Communication-Efficient Heterogeneous Federated Learning with Generalized Heavy-Ball Momentum
di: Zaccone, Riccardo, et al.
Pubblicazione: (2023) -
EPSVec: Efficient and Private Synthetic Data Generation via Dataset Vectors
di: Banayeeanzade, Amin, et al.
Pubblicazione: (2026) -
DoGE: Domain Reweighting with Generalization Estimation
di: Fan, Simin, et al.
Pubblicazione: (2023)