Scaling Laws for Multilingual Language Models
Fuente:
arXiv
Salvato in:
| Autori principali: | He, Yifei, Benhaim, Alon, Patra, Barun, Vaddamanu, Praneetha, Ahuja, Sanchit, Chopra, Parul, Chaudhary, Vishrav, Zhao, Han, Song, Xia |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
EfficientXLang: Towards Improving Token Efficiency Through Cross-Lingual Reasoning
di: Ahuja, Sanchit, et al.
Pubblicazione: (2025)
di: Ahuja, Sanchit, et al.
Pubblicazione: (2025)
Scaling Optimal LR Across Token Horizons
di: Bjorck, Johan, et al.
Pubblicazione: (2024)
di: Bjorck, Johan, et al.
Pubblicazione: (2024)
A Practical Analysis of Human Alignment with *PO
di: Ahrabian, Kian, et al.
Pubblicazione: (2024)
di: Ahrabian, Kian, et al.
Pubblicazione: (2024)
On The Adaptation of Unlimiformer for Decoder-Only Transformers
di: Ahrabian, Kian, et al.
Pubblicazione: (2024)
di: Ahrabian, Kian, et al.
Pubblicazione: (2024)
POROver: Improving Safety and Reducing Overrefusal in Large Language Models with Overgeneration and Preference Optimization
di: Karaman, Batuhan K., et al.
Pubblicazione: (2024)
di: Karaman, Batuhan K., et al.
Pubblicazione: (2024)
A Glitch in the Matrix? Locating and Detecting Language Model Grounding with Fakepedia
di: Monea, Giovanni, et al.
Pubblicazione: (2023)
di: Monea, Giovanni, et al.
Pubblicazione: (2023)
Parameter Alignment Mitigates Catastrophic Forgetting in Multilingual Expert Language Models
di: Ahuja, Sanchit, et al.
Pubblicazione: (2026)
di: Ahuja, Sanchit, et al.
Pubblicazione: (2026)
sPhinX: Sample Efficient Multilingual Instruction Fine-Tuning Through N-shot Guided Prompting
di: Ahuja, Sanchit, et al.
Pubblicazione: (2024)
di: Ahuja, Sanchit, et al.
Pubblicazione: (2024)
S2-Attention: Hardware-Aware Context Sharding Among Attention Heads
di: Lin, Xihui, et al.
Pubblicazione: (2024)
di: Lin, Xihui, et al.
Pubblicazione: (2024)
Contamination Report for Multilingual Benchmarks
di: Ahuja, Sanchit, et al.
Pubblicazione: (2024)
di: Ahuja, Sanchit, et al.
Pubblicazione: (2024)
OWLS: Scaling Laws for Multilingual Speech Recognition and Translation Models
di: Chen, William, et al.
Pubblicazione: (2025)
di: Chen, William, et al.
Pubblicazione: (2025)
Language Models Entangle Language and Culture
di: Jain, Shourya, et al.
Pubblicazione: (2026)
di: Jain, Shourya, et al.
Pubblicazione: (2026)
ATLAS: Adaptive Transfer Scaling Laws for Multilingual Pretraining, Finetuning, and Decoding the Curse of Multilinguality
di: Longpre, Shayne, et al.
Pubblicazione: (2025)
di: Longpre, Shayne, et al.
Pubblicazione: (2025)
Can Small Language Models Use What They Retrieve? An Empirical Study of Retrieval Utilization Across Model Scale
di: Pandey, Sanchit
Pubblicazione: (2026)
di: Pandey, Sanchit
Pubblicazione: (2026)
Exploring Scaling Laws for Local SGD in Large Language Model Training
di: He, Qiaozhi, et al.
Pubblicazione: (2024)
di: He, Qiaozhi, et al.
Pubblicazione: (2024)
Parallel Scaling Law for Language Models
di: Chen, Mouxiang, et al.
Pubblicazione: (2025)
di: Chen, Mouxiang, et al.
Pubblicazione: (2025)
Biased or Flawed? Mitigating Stereotypes in Generative Language Models by Addressing Task-Specific Flaws
di: Jha, Akshita, et al.
Pubblicazione: (2024)
di: Jha, Akshita, et al.
Pubblicazione: (2024)
Theoretical Foundations of Scaling Law in Familial Models
di: Song, Huan, et al.
Pubblicazione: (2025)
di: Song, Huan, et al.
Pubblicazione: (2025)
Skywork-Math: Data Scaling Laws for Mathematical Reasoning in Large Language Models -- The Story Goes On
di: Zeng, Liang, et al.
Pubblicazione: (2024)
di: Zeng, Liang, et al.
Pubblicazione: (2024)
Scaling Laws for Upcycling Mixture-of-Experts Language Models
di: Liew, Seng Pei, et al.
Pubblicazione: (2025)
di: Liew, Seng Pei, et al.
Pubblicazione: (2025)
Scaling Laws for Discriminative Classification in Large Language Models
di: Wyatte, Dean, et al.
Pubblicazione: (2024)
di: Wyatte, Dean, et al.
Pubblicazione: (2024)
Scaling Laws for Conditional Emergence of Multilingual Image Captioning via Generalization from Translation
di: Spravil, Julian, et al.
Pubblicazione: (2025)
di: Spravil, Julian, et al.
Pubblicazione: (2025)
Integrating Large Language Models and Reinforcement Learning for Non-Linear Reasoning
di: Alon, Yoav, et al.
Pubblicazione: (2024)
di: Alon, Yoav, et al.
Pubblicazione: (2024)
Reconciling Kaplan and Chinchilla Scaling Laws
di: Pearce, Tim, et al.
Pubblicazione: (2024)
di: Pearce, Tim, et al.
Pubblicazione: (2024)
Reasoning Towards Fairness: Mitigating Bias in Language Models through Reasoning-Guided Fine-Tuning
di: Kabra, Sanchit, et al.
Pubblicazione: (2025)
di: Kabra, Sanchit, et al.
Pubblicazione: (2025)
Investigating Spatial Attention Bias in Vision-Language Models
di: Chaudhary, Aryan, et al.
Pubblicazione: (2025)
di: Chaudhary, Aryan, et al.
Pubblicazione: (2025)
Can Language Models Discover Scaling Laws?
di: Lin, Haowei, et al.
Pubblicazione: (2025)
di: Lin, Haowei, et al.
Pubblicazione: (2025)
Scaling Laws for Post Training Quantized Large Language Models
di: Xu, Zifei, et al.
Pubblicazione: (2024)
di: Xu, Zifei, et al.
Pubblicazione: (2024)
Scaling Law for Language Models Training Considering Batch Size
di: Shuai, Xian, et al.
Pubblicazione: (2024)
di: Shuai, Xian, et al.
Pubblicazione: (2024)
Scaling Laws for Downstream Task Performance of Large Language Models
di: Isik, Berivan, et al.
Pubblicazione: (2024)
di: Isik, Berivan, et al.
Pubblicazione: (2024)
CMR Scaling Law: Predicting Critical Mixture Ratios for Continual Pre-training of Language Models
di: Gu, Jiawei, et al.
Pubblicazione: (2024)
di: Gu, Jiawei, et al.
Pubblicazione: (2024)
Beyond Chinchilla-Optimal: Accounting for Inference in Language Model Scaling Laws
di: Sardana, Nikhil, et al.
Pubblicazione: (2023)
di: Sardana, Nikhil, et al.
Pubblicazione: (2023)
Towards Understanding the Fragility of Multilingual LLMs against Fine-Tuning Attacks
di: Poppi, Samuele, et al.
Pubblicazione: (2024)
di: Poppi, Samuele, et al.
Pubblicazione: (2024)
Relative-Based Scaling Law for Neural Language Models
di: Yue, Baoqing, et al.
Pubblicazione: (2025)
di: Yue, Baoqing, et al.
Pubblicazione: (2025)
Observational Scaling Laws and the Predictability of Language Model Performance
di: Ruan, Yangjun, et al.
Pubblicazione: (2024)
di: Ruan, Yangjun, et al.
Pubblicazione: (2024)
Uncovering Scaling Laws for Large Language Models via Inverse Problems
di: Verma, Arun, et al.
Pubblicazione: (2025)
di: Verma, Arun, et al.
Pubblicazione: (2025)
Task-Stratified Knowledge Scaling Laws for Post-Training Quantized Large Language Models
di: Zhou, Chenxi, et al.
Pubblicazione: (2025)
di: Zhou, Chenxi, et al.
Pubblicazione: (2025)
Scaling Laws For Mixed Quantization
di: Cao, Zeyu, et al.
Pubblicazione: (2024)
di: Cao, Zeyu, et al.
Pubblicazione: (2024)
In-Context Environments Induce Evaluation-Awareness in Language Models
di: Chaudhary, Maheep
Pubblicazione: (2026)
di: Chaudhary, Maheep
Pubblicazione: (2026)
Disentangling Language Roles in Multilingual LLM Task Execution
di: Zhan, Qishi, et al.
Pubblicazione: (2026)
di: Zhan, Qishi, et al.
Pubblicazione: (2026)
Documenti analoghi
-
EfficientXLang: Towards Improving Token Efficiency Through Cross-Lingual Reasoning
di: Ahuja, Sanchit, et al.
Pubblicazione: (2025) -
Scaling Optimal LR Across Token Horizons
di: Bjorck, Johan, et al.
Pubblicazione: (2024) -
A Practical Analysis of Human Alignment with *PO
di: Ahrabian, Kian, et al.
Pubblicazione: (2024) -
On The Adaptation of Unlimiformer for Decoder-Only Transformers
di: Ahrabian, Kian, et al.
Pubblicazione: (2024) -
POROver: Improving Safety and Reducing Overrefusal in Large Language Models with Overgeneration and Preference Optimization
di: Karaman, Batuhan K., et al.
Pubblicazione: (2024)