SaudiBERT: A Large Language Model Pretrained on Saudi Dialect Corpora
Fuente:
arXiv
Guardado en:
| Autor principal: | Qarah, Faisal |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
EgyBERT: A Large Language Model Pretrained on Egyptian Dialect Corpora
por: Qarah, Faisal
Publicado: (2024)
por: Qarah, Faisal
Publicado: (2024)
AraPoemBERT: A Pretrained Language Model for Arabic Poetry Analysis
por: Qarah, Faisal
Publicado: (2024)
por: Qarah, Faisal
Publicado: (2024)
SaudiCulture: A Benchmark for Evaluating Large Language Models Cultural Competence within Saudi Arabia
por: Ayash, Lama, et al.
Publicado: (2025)
por: Ayash, Lama, et al.
Publicado: (2025)
From Words to Proverbs: Evaluating LLMs Linguistic and Cultural Competence in Saudi Dialects with Absher
por: Al-Monef, Renad, et al.
Publicado: (2025)
por: Al-Monef, Renad, et al.
Publicado: (2025)
Continuous Saudi Sign Language Recognition: A Vision Transformer Approach
por: Elhassen, Soukeina, et al.
Publicado: (2025)
por: Elhassen, Soukeina, et al.
Publicado: (2025)
Patent Language Model Pretraining with ModernBERT
por: Yousefiramandi, Amirhossein, et al.
Publicado: (2025)
por: Yousefiramandi, Amirhossein, et al.
Publicado: (2025)
Attributing Culture-Conditioned Generations to Pretraining Corpora
por: Li, Huihan, et al.
Publicado: (2024)
por: Li, Huihan, et al.
Publicado: (2024)
Data-Augmentation-Based Dialectal Adaptation for LLMs
por: Faisal, Fahim, et al.
Publicado: (2024)
por: Faisal, Fahim, et al.
Publicado: (2024)
Beyond Line-Level Filtering for the Pretraining Corpora of LLMs
por: Park, Chanwoo, et al.
Publicado: (2025)
por: Park, Chanwoo, et al.
Publicado: (2025)
Low-Resource Dialect Adaptation of Large Language Models: A French Dialect Case-Study
por: Khan, Eeham, et al.
Publicado: (2025)
por: Khan, Eeham, et al.
Publicado: (2025)
A Survey on Multilingual Large Language Models: Corpora, Alignment, and Bias
por: Xu, Yuemei, et al.
Publicado: (2024)
por: Xu, Yuemei, et al.
Publicado: (2024)
PhayaThaiBERT: Enhancing a Pretrained Thai Language Model with Unassimilated Loanwords
por: Sriwirote, Panyut, et al.
Publicado: (2023)
por: Sriwirote, Panyut, et al.
Publicado: (2023)
A Survey of Large Language Models for Arabic Language and its Dialects
por: Mashaabi, Malak, et al.
Publicado: (2024)
por: Mashaabi, Malak, et al.
Publicado: (2024)
DIALECTBENCH: A NLP Benchmark for Dialects, Varieties, and Closely-Related Languages
por: Faisal, Fahim, et al.
Publicado: (2024)
por: Faisal, Fahim, et al.
Publicado: (2024)
Dallah: A Dialect-Aware Multimodal Large Language Model for Arabic
por: Alwajih, Fakhraddin, et al.
Publicado: (2024)
por: Alwajih, Fakhraddin, et al.
Publicado: (2024)
Generative AI in Saudi Arabia: A National Survey of Adoption, Risks, and Public Perceptions
por: AlDakheel, Abdulaziz, et al.
Publicado: (2026)
por: AlDakheel, Abdulaziz, et al.
Publicado: (2026)
DialectalArabicMMLU: Benchmarking Dialectal Capabilities in Arabic and Multilingual Language Models
por: Altakrori, Malik H., et al.
Publicado: (2025)
por: Altakrori, Malik H., et al.
Publicado: (2025)
Exploring Narrative Clustering in Large Language Models: A Layerwise Analysis of BERT
por: Banerjee, Awritrojit, et al.
Publicado: (2025)
por: Banerjee, Awritrojit, et al.
Publicado: (2025)
Gender Stereotypes in Professional Roles Among Saudis: An Analytical Study of AI-Generated Images Using Language Models
por: AlKhalifah, Khaloud S., et al.
Publicado: (2025)
por: AlKhalifah, Khaloud S., et al.
Publicado: (2025)
Arabic Dialect Classification using RNNs, Transformers, and Large Language Models: A Comparative Analysis
por: Essameldin, Omar A., et al.
Publicado: (2025)
por: Essameldin, Omar A., et al.
Publicado: (2025)
Grounding Synthetic Data Evaluations of Language Models in Unsupervised Document Corpora
por: Majurski, Michael, et al.
Publicado: (2025)
por: Majurski, Michael, et al.
Publicado: (2025)
Emergent Abilities of Large Language Models under Continued Pretraining for Language Adaptation
por: Elhady, Ahmed, et al.
Publicado: (2025)
por: Elhady, Ahmed, et al.
Publicado: (2025)
Personal Intelligence System UniLM: Hybrid On-Device Small Language Model and Server-Based Large Language Model for Malay Nusantara
por: Nazri, Azree, et al.
Publicado: (2024)
por: Nazri, Azree, et al.
Publicado: (2024)
Memory Decoder: A Pretrained, Plug-and-Play Memory for Large Language Models
por: Cao, Jiaqi, et al.
Publicado: (2025)
por: Cao, Jiaqi, et al.
Publicado: (2025)
RooseBERT: A New Deal For Political Language Modelling
por: Dore, Deborah, et al.
Publicado: (2025)
por: Dore, Deborah, et al.
Publicado: (2025)
Side-by-side Comparison Amplifies Dialect Bias in Language Models
por: Kondapally, Kritee, et al.
Publicado: (2026)
por: Kondapally, Kritee, et al.
Publicado: (2026)
CodePMP: Scalable Preference Model Pretraining for Large Language Model Reasoning
por: Yu, Huimu, et al.
Publicado: (2024)
por: Yu, Huimu, et al.
Publicado: (2024)
Analyzing Narrative Processing in Large Language Models (LLMs): Using GPT4 to test BERT
por: Krauss, Patrick, et al.
Publicado: (2024)
por: Krauss, Patrick, et al.
Publicado: (2024)
Exploring the Benefits of Domain-Pretraining of Generative Large Language Models for Chemistry
por: Acharya, Anurag, et al.
Publicado: (2024)
por: Acharya, Anurag, et al.
Publicado: (2024)
Pretraining Large Language Models with NVFP4
por: NVIDIA, et al.
Publicado: (2025)
por: NVIDIA, et al.
Publicado: (2025)
Utilizing Large Language Models to Generate Synthetic Data to Increase the Performance of BERT-Based Neural Networks
por: Woolsey, Chancellor R., et al.
Publicado: (2024)
por: Woolsey, Chancellor R., et al.
Publicado: (2024)
Analysis and Visualization of Linguistic Structures in Large Language Models: Neural Representations of Verb-Particle Constructions in BERT
por: Kissane, Hassane, et al.
Publicado: (2024)
por: Kissane, Hassane, et al.
Publicado: (2024)
Chinese Tiny LLM: Pretraining a Chinese-Centric Large Language Model
por: Du, Xinrun, et al.
Publicado: (2024)
por: Du, Xinrun, et al.
Publicado: (2024)
Deciphering the Impact of Pretraining Data on Large Language Models through Machine Unlearning
por: Zhao, Yang, et al.
Publicado: (2024)
por: Zhao, Yang, et al.
Publicado: (2024)
Exploring Polyglot Harmony: On Multilingual Data Allocation for Large Language Models Pretraining
por: Guo, Ping, et al.
Publicado: (2025)
por: Guo, Ping, et al.
Publicado: (2025)
Fake News Detection: Comparative Evaluation of BERT-like Models and Large Language Models with Generative AI-Annotated Data
por: Raza, Shaina, et al.
Publicado: (2024)
por: Raza, Shaina, et al.
Publicado: (2024)
NeoBERT: A Next-Generation BERT
por: Breton, Lola Le, et al.
Publicado: (2025)
por: Breton, Lola Le, et al.
Publicado: (2025)
CCI4.0: A Bilingual Pretraining Dataset for Enhancing Reasoning in Large Language Models
por: Liu, Guang, et al.
Publicado: (2025)
por: Liu, Guang, et al.
Publicado: (2025)
Multilingual Pretraining for Pixel Language Models
por: Kesen, Ilker, et al.
Publicado: (2025)
por: Kesen, Ilker, et al.
Publicado: (2025)
KliniskVestBERT: BERT Model Specialised to Norwegian Clinical Texts
por: Autenried, Christian, et al.
Publicado: (2026)
por: Autenried, Christian, et al.
Publicado: (2026)
Ejemplares similares
-
EgyBERT: A Large Language Model Pretrained on Egyptian Dialect Corpora
por: Qarah, Faisal
Publicado: (2024) -
AraPoemBERT: A Pretrained Language Model for Arabic Poetry Analysis
por: Qarah, Faisal
Publicado: (2024) -
SaudiCulture: A Benchmark for Evaluating Large Language Models Cultural Competence within Saudi Arabia
por: Ayash, Lama, et al.
Publicado: (2025) -
From Words to Proverbs: Evaluating LLMs Linguistic and Cultural Competence in Saudi Dialects with Absher
por: Al-Monef, Renad, et al.
Publicado: (2025) -
Continuous Saudi Sign Language Recognition: A Vision Transformer Approach
por: Elhassen, Soukeina, et al.
Publicado: (2025)