On the importance of Data Scale in Pretraining Arabic Language Models
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Ghaddar, Abbas, Langlais, Philippe, Rezagholizadeh, Mehdi, Chen, Boxing |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
CHARP: Conversation History AwaReness Probing for Knowledge-grounded Dialogue Systems
von: Ghaddar, Abbas, et al.
Veröffentlicht: (2024)
von: Ghaddar, Abbas, et al.
Veröffentlicht: (2024)
ReGLA: Refining Gated Linear Attention
von: Lu, Peng, et al.
Veröffentlicht: (2025)
von: Lu, Peng, et al.
Veröffentlicht: (2025)
OTTAWA: Optimal TransporT Adaptive Word Aligner for Hallucination and Omission Translation Errors Detection
von: Huang, Chenyang, et al.
Veröffentlicht: (2024)
von: Huang, Chenyang, et al.
Veröffentlicht: (2024)
CHIQ: Contextual History Enhancement for Improving Query Rewriting in Conversational Search
von: Mo, Fengran, et al.
Veröffentlicht: (2024)
von: Mo, Fengran, et al.
Veröffentlicht: (2024)
Batch-Max: Higher LLM Throughput using Larger Batch Sizes and KV Cache Compression
von: Metel, Michael R., et al.
Veröffentlicht: (2024)
von: Metel, Michael R., et al.
Veröffentlicht: (2024)
Integral Transformer: Denoising Attention, Not Too Much Not Too Little
von: Kobyzev, Ivan, et al.
Veröffentlicht: (2025)
von: Kobyzev, Ivan, et al.
Veröffentlicht: (2025)
BOSCH: Black-Box Binary Optimization for Short-Context Attention-Head Selection in LLMs
von: Ghaddar, Abbas, et al.
Veröffentlicht: (2026)
von: Ghaddar, Abbas, et al.
Veröffentlicht: (2026)
LABO: Towards Learning Optimal Label Regularization via Bi-level Optimization
von: Lu, Peng, et al.
Veröffentlicht: (2023)
von: Lu, Peng, et al.
Veröffentlicht: (2023)
Sorted LLaMA: Unlocking the Potential of Intermediate Layers of Large Language Models for Dynamic Inference
von: Kavehzadeh, Parsa, et al.
Veröffentlicht: (2023)
von: Kavehzadeh, Parsa, et al.
Veröffentlicht: (2023)
Draft on the Fly: Adaptive Self-Speculative Decoding using Cosine Similarity
von: Metel, Michael R., et al.
Veröffentlicht: (2024)
von: Metel, Michael R., et al.
Veröffentlicht: (2024)
Do Robot Snakes Dream like Electric Sheep? Investigating the Effects of Architectural Inductive Biases on Hallucination
von: Huang, Jerry, et al.
Veröffentlicht: (2024)
von: Huang, Jerry, et al.
Veröffentlicht: (2024)
S2D: Sorted Speculative Decoding For More Efficient Deployment of Nested Large Language Models
von: Kavehzadeh, Parsa, et al.
Veröffentlicht: (2024)
von: Kavehzadeh, Parsa, et al.
Veröffentlicht: (2024)
EchoAtt: Attend, Copy, then Adjust for More Efficient Large Language Models
von: Rajabzadeh, Hossein, et al.
Veröffentlicht: (2024)
von: Rajabzadeh, Hossein, et al.
Veröffentlicht: (2024)
Balcony: A Lightweight Approach to Dynamic Inference of Generative Language Models
von: Jamialahmadi, Benyamin, et al.
Veröffentlicht: (2025)
von: Jamialahmadi, Benyamin, et al.
Veröffentlicht: (2025)
QDyLoRA: Quantized Dynamic Low-Rank Adaptation for Efficient Large Language Model Tuning
von: Rajabzadeh, Hossein, et al.
Veröffentlicht: (2024)
von: Rajabzadeh, Hossein, et al.
Veröffentlicht: (2024)
On Evaluation Protocols for Data Augmentation in a Limited Data Scenario
von: Piedboeuf, Frédéric, et al.
Veröffentlicht: (2024)
von: Piedboeuf, Frédéric, et al.
Veröffentlicht: (2024)
EWEK-QA: Enhanced Web and Efficient Knowledge Graph Retrieval for Citation-based Question Answering Systems
von: Dehghan, Mohammad, et al.
Veröffentlicht: (2024)
von: Dehghan, Mohammad, et al.
Veröffentlicht: (2024)
$\textit{BenchIE}^{FL}$ : A Manually Re-Annotated Fact-Based Open Information Extraction Benchmark
von: Lamarche, Fabrice, et al.
Veröffentlicht: (2024)
von: Lamarche, Fabrice, et al.
Veröffentlicht: (2024)
Part-Of-Speech Sensitivity of Routers in Mixture of Experts Models
von: Antoine, Elie, et al.
Veröffentlicht: (2024)
von: Antoine, Elie, et al.
Veröffentlicht: (2024)
Decomposing Retrieval Failures in RAG for Long-Document Financial Question Answering
von: Kobeissi, Amine, et al.
Veröffentlicht: (2026)
von: Kobeissi, Amine, et al.
Veröffentlicht: (2026)
Resonance RoPE: Improving Context Length Generalization of Large Language Models
von: Wang, Suyuchen, et al.
Veröffentlicht: (2024)
von: Wang, Suyuchen, et al.
Veröffentlicht: (2024)
DTRNet: Dynamic Token Routing Network to Reduce Quadratic Costs in Transformers
von: Sharma, Aman, et al.
Veröffentlicht: (2025)
von: Sharma, Aman, et al.
Veröffentlicht: (2025)
The Order Effect: Investigating Prompt Sensitivity to Input Order in LLMs
von: Guan, Bryan, et al.
Veröffentlicht: (2025)
von: Guan, Bryan, et al.
Veröffentlicht: (2025)
Beyond the Limits: A Survey of Techniques to Extend the Context Length in Large Language Models
von: Wang, Xindi, et al.
Veröffentlicht: (2024)
von: Wang, Xindi, et al.
Veröffentlicht: (2024)
AraPoemBERT: A Pretrained Language Model for Arabic Poetry Analysis
von: Qarah, Faisal
Veröffentlicht: (2024)
von: Qarah, Faisal
Veröffentlicht: (2024)
Zebra-Llama: Towards Extremely Efficient Hybrid Models
von: Yang, Mingyu, et al.
Veröffentlicht: (2025)
von: Yang, Mingyu, et al.
Veröffentlicht: (2025)
Enhancing Logical Reasoning in Large Language Models through Graph-based Synthetic Data
von: Zhou, Jiaming, et al.
Veröffentlicht: (2024)
von: Zhou, Jiaming, et al.
Veröffentlicht: (2024)
X-EcoMLA: Upcycling Pre-Trained Attention into MLA for Efficient and Extreme KV Compression
von: Li, Guihong, et al.
Veröffentlicht: (2025)
von: Li, Guihong, et al.
Veröffentlicht: (2025)
A linguistically-motivated evaluation methodology for unraveling model's abilities in reading comprehension tasks
von: Antoine, Elie, et al.
Veröffentlicht: (2025)
von: Antoine, Elie, et al.
Veröffentlicht: (2025)
Towards Practical Tool Usage for Continually Learning LLMs
von: Huang, Jerry, et al.
Veröffentlicht: (2024)
von: Huang, Jerry, et al.
Veröffentlicht: (2024)
Obscuring Data Contamination Through Translation: Evidence from Arabic Corpora
von: Abbas, Chaymaa, et al.
Veröffentlicht: (2026)
von: Abbas, Chaymaa, et al.
Veröffentlicht: (2026)
"Knowing When You Don't Know": A Multilingual Relevance Assessment Dataset for Robust Retrieval-Augmented Generation
von: Thakur, Nandan, et al.
Veröffentlicht: (2023)
von: Thakur, Nandan, et al.
Veröffentlicht: (2023)
Curriculum-Guided Layer Scaling for Language Model Pretraining
von: Singh, Karanpartap, et al.
Veröffentlicht: (2025)
von: Singh, Karanpartap, et al.
Veröffentlicht: (2025)
More Data, Fewer Diacritics: Scaling Arabic TTS
von: Musleh, Ahmed, et al.
Veröffentlicht: (2026)
von: Musleh, Ahmed, et al.
Veröffentlicht: (2026)
Thinking Long, but Short: Stable Sequential Test-Time Scaling for Large Reasoning Models
von: Metel, Michael R., et al.
Veröffentlicht: (2026)
von: Metel, Michael R., et al.
Veröffentlicht: (2026)
Multilingual Language Model Pretraining using Machine-translated Data
von: Wang, Jiayi, et al.
Veröffentlicht: (2025)
von: Wang, Jiayi, et al.
Veröffentlicht: (2025)
Judging Quality Across Languages: A Multilingual Approach to Pretraining Data Filtering with Language Models
von: Ali, Mehdi, et al.
Veröffentlicht: (2025)
von: Ali, Mehdi, et al.
Veröffentlicht: (2025)
The Landscape of Arabic Large Language Models (ALLMs): A New Era for Arabic Language Technology
von: Al-Khalifa, Shahad, et al.
Veröffentlicht: (2025)
von: Al-Khalifa, Shahad, et al.
Veröffentlicht: (2025)
AceGPT, Localizing Large Language Models in Arabic
von: Huang, Huang, et al.
Veröffentlicht: (2023)
von: Huang, Huang, et al.
Veröffentlicht: (2023)
BhashaKritika: Building Synthetic Pretraining Data at Scale for Indic Languages
von: Manoj, Guduru, et al.
Veröffentlicht: (2025)
von: Manoj, Guduru, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
CHARP: Conversation History AwaReness Probing for Knowledge-grounded Dialogue Systems
von: Ghaddar, Abbas, et al.
Veröffentlicht: (2024) -
ReGLA: Refining Gated Linear Attention
von: Lu, Peng, et al.
Veröffentlicht: (2025) -
OTTAWA: Optimal TransporT Adaptive Word Aligner for Hallucination and Omission Translation Errors Detection
von: Huang, Chenyang, et al.
Veröffentlicht: (2024) -
CHIQ: Contextual History Enhancement for Improving Query Rewriting in Conversational Search
von: Mo, Fengran, et al.
Veröffentlicht: (2024) -
Batch-Max: Higher LLM Throughput using Larger Batch Sizes and KV Cache Compression
von: Metel, Michael R., et al.
Veröffentlicht: (2024)