EgyBERT: A Large Language Model Pretrained on Egyptian Dialect Corpora
Fuente:
arXiv
Saved in:
| Main Author: | Qarah, Faisal |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
SaudiBERT: A Large Language Model Pretrained on Saudi Dialect Corpora
by: Qarah, Faisal
Published: (2024)
by: Qarah, Faisal
Published: (2024)
AraPoemBERT: A Pretrained Language Model for Arabic Poetry Analysis
by: Qarah, Faisal
Published: (2024)
by: Qarah, Faisal
Published: (2024)
Data Caricatures: On the Representation of African American Language in Pretraining Corpora
by: Deas, Nicholas, et al.
Published: (2025)
by: Deas, Nicholas, et al.
Published: (2025)
Patent Language Model Pretraining with ModernBERT
by: Yousefiramandi, Amirhossein, et al.
Published: (2025)
by: Yousefiramandi, Amirhossein, et al.
Published: (2025)
Cross-Dialect Text-To-Speech in Pitch-Accent Language Incorporating Multi-Dialect Phoneme-Level BERT
by: Yamauchi, Kazuki, et al.
Published: (2024)
by: Yamauchi, Kazuki, et al.
Published: (2024)
Attributing Culture-Conditioned Generations to Pretraining Corpora
by: Li, Huihan, et al.
Published: (2024)
by: Li, Huihan, et al.
Published: (2024)
A Recipe of Parallel Corpora Exploitation for Multilingual Large Language Models
by: Lin, Peiqin, et al.
Published: (2024)
by: Lin, Peiqin, et al.
Published: (2024)
Make Every Letter Count: Building Dialect Variation Dictionaries from Monolingual Corpora
by: Litschko, Robert, et al.
Published: (2025)
by: Litschko, Robert, et al.
Published: (2025)
Dialectal Toxicity Detection: Evaluating LLM-as-a-Judge Consistency Across Language Varieties
by: Faisal, Fahim, et al.
Published: (2024)
by: Faisal, Fahim, et al.
Published: (2024)
Data-Augmentation-Based Dialectal Adaptation for LLMs
by: Faisal, Fahim, et al.
Published: (2024)
by: Faisal, Fahim, et al.
Published: (2024)
ManufactuBERT: Efficient Continual Pretraining for Manufacturing
by: Armingaud, Robin, et al.
Published: (2025)
by: Armingaud, Robin, et al.
Published: (2025)
Beyond Line-Level Filtering for the Pretraining Corpora of LLMs
by: Park, Chanwoo, et al.
Published: (2025)
by: Park, Chanwoo, et al.
Published: (2025)
Low-Resource Dialect Adaptation of Large Language Models: A French Dialect Case-Study
by: Khan, Eeham, et al.
Published: (2025)
by: Khan, Eeham, et al.
Published: (2025)
Large Language Models Discriminate Against Speakers of German Dialects
by: Bui, Minh Duc, et al.
Published: (2025)
by: Bui, Minh Duc, et al.
Published: (2025)
DialUp! Modeling the Language Continuum by Adapting Models to Dialects and Dialects to Models
by: Bafna, Niyati, et al.
Published: (2025)
by: Bafna, Niyati, et al.
Published: (2025)
A Survey on Multilingual Large Language Models: Corpora, Alignment, and Bias
by: Xu, Yuemei, et al.
Published: (2024)
by: Xu, Yuemei, et al.
Published: (2024)
PhayaThaiBERT: Enhancing a Pretrained Thai Language Model with Unassimilated Loanwords
by: Sriwirote, Panyut, et al.
Published: (2023)
by: Sriwirote, Panyut, et al.
Published: (2023)
Leveraging Large Language Models to Measure Gender Representation Bias in Gendered Language Corpora
by: Derner, Erik, et al.
Published: (2024)
by: Derner, Erik, et al.
Published: (2024)
A Survey of Large Language Models for Arabic Language and its Dialects
by: Mashaabi, Malak, et al.
Published: (2024)
by: Mashaabi, Malak, et al.
Published: (2024)
New Textual Corpora for Serbian Language Modeling
by: Škorić, Mihailo, et al.
Published: (2024)
by: Škorić, Mihailo, et al.
Published: (2024)
Findings of the BabyLM Challenge: Sample-Efficient Pretraining on Developmentally Plausible Corpora
by: Warstadt, Alex, et al.
Published: (2025)
by: Warstadt, Alex, et al.
Published: (2025)
DIALECTBENCH: A NLP Benchmark for Dialects, Varieties, and Closely-Related Languages
by: Faisal, Fahim, et al.
Published: (2024)
by: Faisal, Fahim, et al.
Published: (2024)
A Hierarchical and Attentional Analysis of Argument Structure Constructions in BERT Using Naturalistic Corpora
by: Kaipeng, Liu, et al.
Published: (2026)
by: Kaipeng, Liu, et al.
Published: (2026)
Analysis of Argument Structure Constructions in the Large Language Model BERT
by: Ramezani, Pegah, et al.
Published: (2024)
by: Ramezani, Pegah, et al.
Published: (2024)
Dallah: A Dialect-Aware Multimodal Large Language Model for Arabic
by: Alwajih, Fakhraddin, et al.
Published: (2024)
by: Alwajih, Fakhraddin, et al.
Published: (2024)
Findings of the Second BabyLM Challenge: Sample-Efficient Pretraining on Developmentally Plausible Corpora
by: Hu, Michael Y., et al.
Published: (2024)
by: Hu, Michael Y., et al.
Published: (2024)
Meenz bleibt Meenz, but Large Language Models Do Not Speak Its Dialect
by: Bui, Minh Duc, et al.
Published: (2026)
by: Bui, Minh Duc, et al.
Published: (2026)
A Review of the Challenges with Massive Web-mined Corpora Used in Large Language Models Pre-Training
by: Perełkiewicz, Michał, et al.
Published: (2024)
by: Perełkiewicz, Michał, et al.
Published: (2024)
Self-reflecting Large Language Models: A Hegelian Dialectical Approach
by: Abdali, Sara, et al.
Published: (2025)
by: Abdali, Sara, et al.
Published: (2025)
Validating and Exploring Large Geographic Corpora
by: Dunn, Jonathan
Published: (2024)
by: Dunn, Jonathan
Published: (2024)
Detecting Bias in Large Language Models: Fine-tuned KcBERT
by: Lee, J. K., et al.
Published: (2024)
by: Lee, J. K., et al.
Published: (2024)
Exploring Large Language Models in Healthcare: Insights into Corpora Sources, Customization Strategies, and Evaluation Metrics
by: Yang, Shuqi, et al.
Published: (2025)
by: Yang, Shuqi, et al.
Published: (2025)
MosaicBERT: A Bidirectional Encoder Optimized for Fast Pretraining
by: Portes, Jacob, et al.
Published: (2023)
by: Portes, Jacob, et al.
Published: (2023)
Identifying Emerging Concepts in Large Corpora
by: Ma, Sibo, et al.
Published: (2025)
by: Ma, Sibo, et al.
Published: (2025)
Dialect Normalization using Large Language Models and Morphological Rules
by: Dimakis, Antonios, et al.
Published: (2025)
by: Dimakis, Antonios, et al.
Published: (2025)
Investigating the 'Autoencoder Behavior' in Speech Self-Supervised Models: a focus on HuBERT's Pretraining
by: Vielzeuf, Valentin
Published: (2024)
by: Vielzeuf, Valentin
Published: (2024)
Atlas-Chat: Adapting Large Language Models for Low-Resource Moroccan Arabic Dialect
by: Shang, Guokan, et al.
Published: (2024)
by: Shang, Guokan, et al.
Published: (2024)
Assessing Dialect Fairness and Robustness of Large Language Models in Reasoning Tasks
by: Lin, Fangru, et al.
Published: (2024)
by: Lin, Fangru, et al.
Published: (2024)
TabiBERT: A Large-Scale ModernBERT Foundation Model and A Unified Benchmark for Turkish
by: Türker, Melikşah, et al.
Published: (2025)
by: Türker, Melikşah, et al.
Published: (2025)
DialectalArabicMMLU: Benchmarking Dialectal Capabilities in Arabic and Multilingual Language Models
by: Altakrori, Malik H., et al.
Published: (2025)
by: Altakrori, Malik H., et al.
Published: (2025)
Similar Items
-
SaudiBERT: A Large Language Model Pretrained on Saudi Dialect Corpora
by: Qarah, Faisal
Published: (2024) -
AraPoemBERT: A Pretrained Language Model for Arabic Poetry Analysis
by: Qarah, Faisal
Published: (2024) -
Data Caricatures: On the Representation of African American Language in Pretraining Corpora
by: Deas, Nicholas, et al.
Published: (2025) -
Patent Language Model Pretraining with ModernBERT
by: Yousefiramandi, Amirhossein, et al.
Published: (2025) -
Cross-Dialect Text-To-Speech in Pitch-Accent Language Incorporating Multi-Dialect Phoneme-Level BERT
by: Yamauchi, Kazuki, et al.
Published: (2024)