On the Scaling Laws of Geographical Representation in Language Models
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Godey, Nathan, de la Clergerie, Éric, Sagot, Benoît |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Why do small language models underperform? Studying Language Model Saturation via the Softmax Bottleneck
von: Godey, Nathan, et al.
Veröffentlicht: (2024)
von: Godey, Nathan, et al.
Veröffentlicht: (2024)
Gaperon: A Peppered English-French Generative Language Model Suite
von: Godey, Nathan, et al.
Veröffentlicht: (2025)
von: Godey, Nathan, et al.
Veröffentlicht: (2025)
Anisotropy Is Inherent to Self-Attention in Transformers
von: Godey, Nathan, et al.
Veröffentlicht: (2024)
von: Godey, Nathan, et al.
Veröffentlicht: (2024)
Q-Filters: Leveraging QK Geometry for Efficient KV Cache Compression
von: Godey, Nathan, et al.
Veröffentlicht: (2025)
von: Godey, Nathan, et al.
Veröffentlicht: (2025)
PatentEval: Understanding Errors in Patent Generation
von: Zuo, You, et al.
Veröffentlicht: (2024)
von: Zuo, You, et al.
Veröffentlicht: (2024)
A Causal Language Modeling Detour Improves Encoder Continued Pretraining
von: Touchent, Rian, et al.
Veröffentlicht: (2026)
von: Touchent, Rian, et al.
Veröffentlicht: (2026)
Biomed-Enriched: A Biomedical Dataset Enriched with LLMs for Pretraining and Extracting Rare and Hidden Content
von: Touchent, Rian, et al.
Veröffentlicht: (2025)
von: Touchent, Rian, et al.
Veröffentlicht: (2025)
CamemBERT-bio: Leveraging Continual Pre-training for Cost-Effective Models on French Biomedical Data
von: Touchent, Rian, et al.
Veröffentlicht: (2023)
von: Touchent, Rian, et al.
Veröffentlicht: (2023)
Patent Representation Learning via Self-supervision
von: Zuo, You, et al.
Veröffentlicht: (2025)
von: Zuo, You, et al.
Veröffentlicht: (2025)
CamemBERT 2.0: A Smarter French Language Model Aged to Perfection
von: Antoun, Wissam, et al.
Veröffentlicht: (2024)
von: Antoun, Wissam, et al.
Veröffentlicht: (2024)
Reading or Guessing? Visual Grounding Failures of Vision-Language Models for OCR in Ancient Greek Editions
von: Karamolegkou, Antonia, et al.
Veröffentlicht: (2026)
von: Karamolegkou, Antonia, et al.
Veröffentlicht: (2026)
BigO(Bench) -- Can LLMs Generate Code with Controlled Time and Space Complexity?
von: Chambon, Pierre, et al.
Veröffentlicht: (2025)
von: Chambon, Pierre, et al.
Veröffentlicht: (2025)
Scaling Laws of Synthetic Data for Language Models
von: Qin, Zeyu, et al.
Veröffentlicht: (2025)
von: Qin, Zeyu, et al.
Veröffentlicht: (2025)
Diagnosing Representation Dynamics in NER Model Extension
von: Zhang, Xirui, et al.
Veröffentlicht: (2025)
von: Zhang, Xirui, et al.
Veröffentlicht: (2025)
PLDR-LLM: Large Language Model from Power Law Decoder Representations
von: Gokden, Burc
Veröffentlicht: (2024)
von: Gokden, Burc
Veröffentlicht: (2024)
Can Language Models Discover Scaling Laws?
von: Lin, Haowei, et al.
Veröffentlicht: (2025)
von: Lin, Haowei, et al.
Veröffentlicht: (2025)
Observational Scaling Laws and the Predictability of Language Model Performance
von: Ruan, Yangjun, et al.
Veröffentlicht: (2024)
von: Ruan, Yangjun, et al.
Veröffentlicht: (2024)
Relative-Based Scaling Law for Neural Language Models
von: Yue, Baoqing, et al.
Veröffentlicht: (2025)
von: Yue, Baoqing, et al.
Veröffentlicht: (2025)
Weber's Law in Transformer Magnitude Representations: Efficient Coding, Representational Geometry, and Psychophysical Laws in Language Models
von: Cacioli, Jon-Paul
Veröffentlicht: (2026)
von: Cacioli, Jon-Paul
Veröffentlicht: (2026)
Large Language Models are Geographically Biased
von: Manvi, Rohin, et al.
Veröffentlicht: (2024)
von: Manvi, Rohin, et al.
Veröffentlicht: (2024)
Not-Just-Scaling Laws: Towards a Better Understanding of the Downstream Impact of Language Model Design Decisions
von: Liu, Emmy, et al.
Veröffentlicht: (2025)
von: Liu, Emmy, et al.
Veröffentlicht: (2025)
How to inject knowledge efficiently? Knowledge Infusion Scaling Law for Pre-training Large Language Models
von: Lv, Kangtao, et al.
Veröffentlicht: (2025)
von: Lv, Kangtao, et al.
Veröffentlicht: (2025)
Scaling Laws for Speculative Decoding
von: Yan, Siyuan, et al.
Veröffentlicht: (2025)
von: Yan, Siyuan, et al.
Veröffentlicht: (2025)
Scaling Laws with Vocabulary: Larger Models Deserve Larger Vocabularies
von: Tao, Chaofan, et al.
Veröffentlicht: (2024)
von: Tao, Chaofan, et al.
Veröffentlicht: (2024)
Provable Scaling Laws for the Test-Time Compute of Large Language Models
von: Chen, Yanxi, et al.
Veröffentlicht: (2024)
von: Chen, Yanxi, et al.
Veröffentlicht: (2024)
Uncovering Scaling Laws for Large Language Models via Inverse Problems
von: Verma, Arun, et al.
Veröffentlicht: (2025)
von: Verma, Arun, et al.
Veröffentlicht: (2025)
Finding the Minimal Parameter Budget for Implicit Reasoning: A Data Complexity Driven Scaling Law for Language Models
von: Wang, Xinyi, et al.
Veröffentlicht: (2025)
von: Wang, Xinyi, et al.
Veröffentlicht: (2025)
How do Scaling Laws Apply to Knowledge Graph Engineering Tasks? The Impact of Model Size on Large Language Model Performance
von: Heim, Desiree, et al.
Veröffentlicht: (2025)
von: Heim, Desiree, et al.
Veröffentlicht: (2025)
Codenames as a Benchmark for Large Language Models
von: Stephenson, Matthew, et al.
Veröffentlicht: (2024)
von: Stephenson, Matthew, et al.
Veröffentlicht: (2024)
Emergent Hierarchical Structure in Large Language Models: An Information-Theoretic Framework for Multi-Scale Representation
von: Zhang, Yukin, et al.
Veröffentlicht: (2025)
von: Zhang, Yukin, et al.
Veröffentlicht: (2025)
Scaling Laws of Decoder-Only Models on the Multilingual Machine Translation Task
von: Caillaut, Gaëtan, et al.
Veröffentlicht: (2024)
von: Caillaut, Gaëtan, et al.
Veröffentlicht: (2024)
Establishing Task Scaling Laws via Compute-Efficient Model Ladders
von: Bhagia, Akshita, et al.
Veröffentlicht: (2024)
von: Bhagia, Akshita, et al.
Veröffentlicht: (2024)
Physics of Language Models: Part 3.3, Knowledge Capacity Scaling Laws
von: Allen-Zhu, Zeyuan, et al.
Veröffentlicht: (2024)
von: Allen-Zhu, Zeyuan, et al.
Veröffentlicht: (2024)
Selecting Large Language Model to Fine-tune via Rectified Scaling Law
von: Lin, Haowei, et al.
Veröffentlicht: (2024)
von: Lin, Haowei, et al.
Veröffentlicht: (2024)
Factual Self-Awareness in Language Models: Representation, Robustness, and Scaling
von: Tamoyan, Hovhannes, et al.
Veröffentlicht: (2025)
von: Tamoyan, Hovhannes, et al.
Veröffentlicht: (2025)
Scaling Laws of RoPE-based Extrapolation
von: Liu, Xiaoran, et al.
Veröffentlicht: (2023)
von: Liu, Xiaoran, et al.
Veröffentlicht: (2023)
The Scaling Laws of Skills in LLM Agent Systems
von: Chen, Charles, et al.
Veröffentlicht: (2026)
von: Chen, Charles, et al.
Veröffentlicht: (2026)
Task-Stratified Knowledge Scaling Laws for Post-Training Quantized Large Language Models
von: Zhou, Chenxi, et al.
Veröffentlicht: (2025)
von: Zhou, Chenxi, et al.
Veröffentlicht: (2025)
Lost in Backpropagation: The LM Head is a Gradient Bottleneck
von: Godey, Nathan, et al.
Veröffentlicht: (2026)
von: Godey, Nathan, et al.
Veröffentlicht: (2026)
Automated Triaging and Transfer Learning of Incident Learning Safety Reports Using Large Language Representational Models
von: Beidler, Peter, et al.
Veröffentlicht: (2025)
von: Beidler, Peter, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Why do small language models underperform? Studying Language Model Saturation via the Softmax Bottleneck
von: Godey, Nathan, et al.
Veröffentlicht: (2024) -
Gaperon: A Peppered English-French Generative Language Model Suite
von: Godey, Nathan, et al.
Veröffentlicht: (2025) -
Anisotropy Is Inherent to Self-Attention in Transformers
von: Godey, Nathan, et al.
Veröffentlicht: (2024) -
Q-Filters: Leveraging QK Geometry for Efficient KV Cache Compression
von: Godey, Nathan, et al.
Veröffentlicht: (2025) -
PatentEval: Understanding Errors in Patent Generation
von: Zuo, You, et al.
Veröffentlicht: (2024)