Scaling Laws for Downstream Task Performance of Large Language Models
Fuente:
arXiv
Salvato in:
| Autori principali: | Isik, Berivan, Ponomareva, Natalia, Hazimeh, Hussein, Paparas, Dimitris, Vassilvitskii, Sergei, Koyejo, Sanmi |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
On Fairness of Low-Rank Adaptation of Large Models
di: Ding, Zhoujie, et al.
Pubblicazione: (2024)
di: Ding, Zhoujie, et al.
Pubblicazione: (2024)
Discovering Implicit Large Language Model Alignment Objectives
di: Chen, Edward, et al.
Pubblicazione: (2026)
di: Chen, Edward, et al.
Pubblicazione: (2026)
Differentially Private Adaptation of Diffusion Models via Noisy Aggregated Embeddings
di: Peetathawatchai, Pura, et al.
Pubblicazione: (2024)
di: Peetathawatchai, Pura, et al.
Pubblicazione: (2024)
Quantifying the Importance of Data Alignment in Downstream Model Performance
di: Chawla, Krrish, et al.
Pubblicazione: (2025)
di: Chawla, Krrish, et al.
Pubblicazione: (2025)
An Optimization Framework for Differentially Private Sparse Fine-Tuning
di: Makni, Mehdi, et al.
Pubblicazione: (2025)
di: Makni, Mehdi, et al.
Pubblicazione: (2025)
Why Has Predicting Downstream Capabilities of Frontier AI Models with Scale Remained Elusive?
di: Schaeffer, Rylan, et al.
Pubblicazione: (2024)
di: Schaeffer, Rylan, et al.
Pubblicazione: (2024)
Scaling Laws Are Unreliable for Downstream Tasks: A Reality Check
di: Lourie, Nicholas, et al.
Pubblicazione: (2025)
di: Lourie, Nicholas, et al.
Pubblicazione: (2025)
Adaptive Compression in Federated Learning via Side Information
di: Isik, Berivan, et al.
Pubblicazione: (2023)
di: Isik, Berivan, et al.
Pubblicazione: (2023)
OSSCAR: One-Shot Structured Pruning in Vision and Language Models with Combinatorial Optimization
di: Meng, Xiang, et al.
Pubblicazione: (2024)
di: Meng, Xiang, et al.
Pubblicazione: (2024)
Private prediction for large-scale synthetic text generation
di: Amin, Kareem, et al.
Pubblicazione: (2024)
di: Amin, Kareem, et al.
Pubblicazione: (2024)
Scaling Laws for Predicting Downstream Performance in LLMs
di: Chen, Yangyi, et al.
Pubblicazione: (2024)
di: Chen, Yangyi, et al.
Pubblicazione: (2024)
Reasoning Models Don't Just Think Longer, They Move Differently
di: Gjølbye, Anders, et al.
Pubblicazione: (2026)
di: Gjølbye, Anders, et al.
Pubblicazione: (2026)
The Magic Correlations: Understanding Knowledge Transfer from Pretraining to Supervised Fine-Tuning
di: Fan, Simin, et al.
Pubblicazione: (2026)
di: Fan, Simin, et al.
Pubblicazione: (2026)
Towards Modeling Learner Performance with Large Language Models
di: Neshaei, Seyed Parsa, et al.
Pubblicazione: (2024)
di: Neshaei, Seyed Parsa, et al.
Pubblicazione: (2024)
Escaping Collapse: The Strength of Weak Data for Large Language Model Training
di: Amin, Kareem, et al.
Pubblicazione: (2025)
di: Amin, Kareem, et al.
Pubblicazione: (2025)
Why Do Safety Guardrails Degrade Across Languages?
di: Zhang, Max, et al.
Pubblicazione: (2026)
di: Zhang, Max, et al.
Pubblicazione: (2026)
From Passive to Active Reasoning: Can Large Language Models Ask the Right Questions under Incomplete Information?
di: Zhou, Zhanke, et al.
Pubblicazione: (2025)
di: Zhou, Zhanke, et al.
Pubblicazione: (2025)
Pretraining Scaling Laws for Generative Evaluations of Language Models
di: Schaeffer, Rylan, et al.
Pubblicazione: (2025)
di: Schaeffer, Rylan, et al.
Pubblicazione: (2025)
Lottery Ticket Adaptation: Mitigating Destructive Interference in LLMs
di: Panda, Ashwinee, et al.
Pubblicazione: (2024)
di: Panda, Ashwinee, et al.
Pubblicazione: (2024)
Logits are All We Need to Adapt Closed Models
di: Hiranandani, Gaurush, et al.
Pubblicazione: (2025)
di: Hiranandani, Gaurush, et al.
Pubblicazione: (2025)
DART: A Principled Approach to Adversarially Robust Unsupervised Domain Adaptation
di: Wang, Yunjuan, et al.
Pubblicazione: (2024)
di: Wang, Yunjuan, et al.
Pubblicazione: (2024)
Can Character-based Language Models Improve Downstream Task Performance in Low-Resource and Noisy Language Scenarios?
di: Riabi, Arij, et al.
Pubblicazione: (2021)
di: Riabi, Arij, et al.
Pubblicazione: (2021)
Performance Law of Large Language Models
di: Wu, Chuhan, et al.
Pubblicazione: (2024)
di: Wu, Chuhan, et al.
Pubblicazione: (2024)
Revisiting the Scaling Properties of Downstream Metrics in Large Language Model Training
di: Krajewski, Jakub, et al.
Pubblicazione: (2025)
di: Krajewski, Jakub, et al.
Pubblicazione: (2025)
Adapting Decoder-Based Language Models for Diverse Encoder Downstream Tasks
di: Suganthan, Paul, et al.
Pubblicazione: (2025)
di: Suganthan, Paul, et al.
Pubblicazione: (2025)
Scaling Laws for Discriminative Classification in Large Language Models
di: Wyatte, Dean, et al.
Pubblicazione: (2024)
di: Wyatte, Dean, et al.
Pubblicazione: (2024)
Rethinking Machine Unlearning for Large Language Models
di: Liu, Sijia, et al.
Pubblicazione: (2024)
di: Liu, Sijia, et al.
Pubblicazione: (2024)
Beyond Scale: The Diversity Coefficient as a Data Quality Metric for Variability in Natural Language Data
di: Miranda, Brando, et al.
Pubblicazione: (2023)
di: Miranda, Brando, et al.
Pubblicazione: (2023)
Scalable Ensembling For Mitigating Reward Overoptimisation
di: Ahmed, Ahmed M., et al.
Pubblicazione: (2024)
di: Ahmed, Ahmed M., et al.
Pubblicazione: (2024)
FedEval-LLM: Federated Evaluation of Large Language Models on Downstream Tasks with Collective Wisdom
di: He, Yuanqin, et al.
Pubblicazione: (2024)
di: He, Yuanqin, et al.
Pubblicazione: (2024)
Investigating Data Contamination for Pre-training Language Models
di: Jiang, Minhao, et al.
Pubblicazione: (2024)
di: Jiang, Minhao, et al.
Pubblicazione: (2024)
Reliable and Efficient Amortized Model-based Evaluation
di: Truong, Sang, et al.
Pubblicazione: (2025)
di: Truong, Sang, et al.
Pubblicazione: (2025)
Exploring the Impact of a Transformer's Latent Space Geometry on Downstream Task Performance
di: Marbut, Anna C., et al.
Pubblicazione: (2024)
di: Marbut, Anna C., et al.
Pubblicazione: (2024)
Task-Stratified Knowledge Scaling Laws for Post-Training Quantized Large Language Models
di: Zhou, Chenxi, et al.
Pubblicazione: (2025)
di: Zhou, Chenxi, et al.
Pubblicazione: (2025)
Extracting books from production language models
di: Ahmed, Ahmed, et al.
Pubblicazione: (2026)
di: Ahmed, Ahmed, et al.
Pubblicazione: (2026)
Scaling Laws for Post Training Quantized Large Language Models
di: Xu, Zifei, et al.
Pubblicazione: (2024)
di: Xu, Zifei, et al.
Pubblicazione: (2024)
Learning to Poison Large Language Models for Downstream Manipulation
di: Zhou, Xiangyu, et al.
Pubblicazione: (2024)
di: Zhou, Xiangyu, et al.
Pubblicazione: (2024)
Exploring Scaling Laws for Local SGD in Large Language Model Training
di: He, Qiaozhi, et al.
Pubblicazione: (2024)
di: He, Qiaozhi, et al.
Pubblicazione: (2024)
Predicting Task Performance with Context-aware Scaling Laws
di: Montgomery, Kyle, et al.
Pubblicazione: (2025)
di: Montgomery, Kyle, et al.
Pubblicazione: (2025)
Observational Scaling Laws and the Predictability of Language Model Performance
di: Ruan, Yangjun, et al.
Pubblicazione: (2024)
di: Ruan, Yangjun, et al.
Pubblicazione: (2024)
Documenti analoghi
-
On Fairness of Low-Rank Adaptation of Large Models
di: Ding, Zhoujie, et al.
Pubblicazione: (2024) -
Discovering Implicit Large Language Model Alignment Objectives
di: Chen, Edward, et al.
Pubblicazione: (2026) -
Differentially Private Adaptation of Diffusion Models via Noisy Aggregated Embeddings
di: Peetathawatchai, Pura, et al.
Pubblicazione: (2024) -
Quantifying the Importance of Data Alignment in Downstream Model Performance
di: Chawla, Krrish, et al.
Pubblicazione: (2025) -
An Optimization Framework for Differentially Private Sparse Fine-Tuning
di: Makni, Mehdi, et al.
Pubblicazione: (2025)