Zero-Shot Performance Prediction for Probabilistic Scaling Laws
Fuente:
arXiv
Salvato in:
| Autori principali: | Schram, Viktoria, Hiller, Markus, Beck, Daniel, Cohn, Trevor |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Active Budget Allocation for Efficient Scaling Law Estimation via Surrogate-Guided Pruning
di: Schram, Viktoria, et al.
Pubblicazione: (2026)
di: Schram, Viktoria, et al.
Pubblicazione: (2026)
Don't Ignore the Tail: Decoupling top-K Probabilities for Efficient Language Model Distillation
di: Dasgupta, Sayantan, et al.
Pubblicazione: (2026)
di: Dasgupta, Sayantan, et al.
Pubblicazione: (2026)
Scaling Laws for Predicting Downstream Performance in LLMs
di: Chen, Yangyi, et al.
Pubblicazione: (2024)
di: Chen, Yangyi, et al.
Pubblicazione: (2024)
Predicting Task Performance with Context-aware Scaling Laws
di: Montgomery, Kyle, et al.
Pubblicazione: (2025)
di: Montgomery, Kyle, et al.
Pubblicazione: (2025)
Observational Scaling Laws and the Predictability of Language Model Performance
di: Ruan, Yangjun, et al.
Pubblicazione: (2024)
di: Ruan, Yangjun, et al.
Pubblicazione: (2024)
Knowledge Localization in Mixture-of-Experts LLMs Using Cross-Lingual Inconsistency
di: Bandarkar, Lucas, et al.
Pubblicazione: (2026)
di: Bandarkar, Lucas, et al.
Pubblicazione: (2026)
Perplexity-Aware Data Scaling Law: Perplexity Landscapes Predict Performance for Continual Pre-training
di: Liu, Lei, et al.
Pubblicazione: (2025)
di: Liu, Lei, et al.
Pubblicazione: (2025)
gzip Predicts Data-dependent Scaling Laws
di: Pandey, Rohan
Pubblicazione: (2024)
di: Pandey, Rohan
Pubblicazione: (2024)
Multi-EuP: The Multilingual European Parliament Dataset for Analysis of Bias in Information Retrieval
di: Yang, Jinrui, et al.
Pubblicazione: (2023)
di: Yang, Jinrui, et al.
Pubblicazione: (2023)
Scaling Laws for Downstream Task Performance of Large Language Models
di: Isik, Berivan, et al.
Pubblicazione: (2024)
di: Isik, Berivan, et al.
Pubblicazione: (2024)
Scaling Laws for Precision
di: Kumar, Tanishq, et al.
Pubblicazione: (2024)
di: Kumar, Tanishq, et al.
Pubblicazione: (2024)
Rethinking Cross-lingual Gaps from a Statistical Viewpoint
di: Piratla, Vihari, et al.
Pubblicazione: (2025)
di: Piratla, Vihari, et al.
Pubblicazione: (2025)
PTPP-Aware Adaptation Scaling Laws: Predicting Domain-Adaptation Performance at Unseen Pre-Training Budgets
di: Goffinet, Etienne, et al.
Pubblicazione: (2025)
di: Goffinet, Etienne, et al.
Pubblicazione: (2025)
Zero-Shot Conversational Stance Detection: Dataset and Approaches
di: Ding, Yuzhe, et al.
Pubblicazione: (2025)
di: Ding, Yuzhe, et al.
Pubblicazione: (2025)
Says Who? Effective Zero-Shot Annotation of Focalization
di: Hicke, Rebecca M. M., et al.
Pubblicazione: (2024)
di: Hicke, Rebecca M. M., et al.
Pubblicazione: (2024)
PREDILECT: Preferences Delineated with Zero-Shot Language-based Reasoning in Reinforcement Learning
di: Holk, Simon, et al.
Pubblicazione: (2024)
di: Holk, Simon, et al.
Pubblicazione: (2024)
Scaling Laws For Mixed Quantization
di: Cao, Zeyu, et al.
Pubblicazione: (2024)
di: Cao, Zeyu, et al.
Pubblicazione: (2024)
Neural Neural Scaling Laws
di: Hu, Michael Y., et al.
Pubblicazione: (2026)
di: Hu, Michael Y., et al.
Pubblicazione: (2026)
Loss-to-Loss Prediction: Scaling Laws for All Datasets
di: Brandfonbrener, David, et al.
Pubblicazione: (2024)
di: Brandfonbrener, David, et al.
Pubblicazione: (2024)
Optimizing GPT for Video Understanding: Zero-Shot Performance and Prompt Engineering
di: Beliaev, Mark, et al.
Pubblicazione: (2025)
di: Beliaev, Mark, et al.
Pubblicazione: (2025)
CMR Scaling Law: Predicting Critical Mixture Ratios for Continual Pre-training of Language Models
di: Gu, Jiawei, et al.
Pubblicazione: (2024)
di: Gu, Jiawei, et al.
Pubblicazione: (2024)
Zero-Shot Hierarchical Classification on the Common Procurement Vocabulary Taxonomy
di: Moiraghi, Federico, et al.
Pubblicazione: (2024)
di: Moiraghi, Federico, et al.
Pubblicazione: (2024)
Unified Scaling Laws for Compressed Representations
di: Panferov, Andrei, et al.
Pubblicazione: (2025)
di: Panferov, Andrei, et al.
Pubblicazione: (2025)
Parallel Scaling Law for Language Models
di: Chen, Mouxiang, et al.
Pubblicazione: (2025)
di: Chen, Mouxiang, et al.
Pubblicazione: (2025)
Scaling Law for Quantization-Aware Training
di: Chen, Mengzhao, et al.
Pubblicazione: (2025)
di: Chen, Mengzhao, et al.
Pubblicazione: (2025)
Reconciling Kaplan and Chinchilla Scaling Laws
di: Pearce, Tim, et al.
Pubblicazione: (2024)
di: Pearce, Tim, et al.
Pubblicazione: (2024)
Scaling Laws for Multilingual Language Models
di: He, Yifei, et al.
Pubblicazione: (2024)
di: He, Yifei, et al.
Pubblicazione: (2024)
Scaling Laws for Many-Shot In-Context Learning with Self-Generated Annotations
di: Gu, Zhengyao, et al.
Pubblicazione: (2025)
di: Gu, Zhengyao, et al.
Pubblicazione: (2025)
How Much Is One Recurrence Worth? Iso-Depth Scaling Laws for Looped Language Models
di: Schwethelm, Kristian, et al.
Pubblicazione: (2026)
di: Schwethelm, Kristian, et al.
Pubblicazione: (2026)
Zero-Shot Decision Tree Construction via Large Language Models
di: Carrasco, Lucas, et al.
Pubblicazione: (2025)
di: Carrasco, Lucas, et al.
Pubblicazione: (2025)
Learning to Generate Instruction Tuning Datasets for Zero-Shot Task Adaptation
di: Nayak, Nihal V., et al.
Pubblicazione: (2024)
di: Nayak, Nihal V., et al.
Pubblicazione: (2024)
Performance Law of Large Language Models
di: Wu, Chuhan, et al.
Pubblicazione: (2024)
di: Wu, Chuhan, et al.
Pubblicazione: (2024)
Kinetics: Rethinking Test-Time Scaling Laws
di: Sadhukhan, Ranajoy, et al.
Pubblicazione: (2025)
di: Sadhukhan, Ranajoy, et al.
Pubblicazione: (2025)
Compression Scaling Laws:Unifying Sparsity and Quantization
di: Frantar, Elias, et al.
Pubblicazione: (2025)
di: Frantar, Elias, et al.
Pubblicazione: (2025)
Prescriptive Scaling Laws for Data Constrained Training
di: Lovelace, Justin, et al.
Pubblicazione: (2026)
di: Lovelace, Justin, et al.
Pubblicazione: (2026)
Unraveling the Mystery of Scaling Laws: Part I
di: Su, Hui, et al.
Pubblicazione: (2024)
di: Su, Hui, et al.
Pubblicazione: (2024)
Zero-Shot Detection of LLM-Generated Code via Approximated Task Conditioning
di: Ashkenazi, Maor, et al.
Pubblicazione: (2025)
di: Ashkenazi, Maor, et al.
Pubblicazione: (2025)
ICXML: An In-Context Learning Framework for Zero-Shot Extreme Multi-Label Classification
di: Zhu, Yaxin, et al.
Pubblicazione: (2023)
di: Zhu, Yaxin, et al.
Pubblicazione: (2023)
Improving Zero-Shot Cross-Lingual Transfer via Progressive Code-Switching
di: Li, Zhuoran, et al.
Pubblicazione: (2024)
di: Li, Zhuoran, et al.
Pubblicazione: (2024)
Scaling Laws for Upcycling Mixture-of-Experts Language Models
di: Liew, Seng Pei, et al.
Pubblicazione: (2025)
di: Liew, Seng Pei, et al.
Pubblicazione: (2025)
Documenti analoghi
-
Active Budget Allocation for Efficient Scaling Law Estimation via Surrogate-Guided Pruning
di: Schram, Viktoria, et al.
Pubblicazione: (2026) -
Don't Ignore the Tail: Decoupling top-K Probabilities for Efficient Language Model Distillation
di: Dasgupta, Sayantan, et al.
Pubblicazione: (2026) -
Scaling Laws for Predicting Downstream Performance in LLMs
di: Chen, Yangyi, et al.
Pubblicazione: (2024) -
Predicting Task Performance with Context-aware Scaling Laws
di: Montgomery, Kyle, et al.
Pubblicazione: (2025) -
Observational Scaling Laws and the Predictability of Language Model Performance
di: Ruan, Yangjun, et al.
Pubblicazione: (2024)