Ensemble of pre-trained language models and data augmentation for hate speech detection from Arabic tweets
Fuente:
arXiv
Saved in:
| Main Authors: | Daouadi, Kheir Eddine, Boualleg, Yaakoub, Haouaouchi, Kheir Eddine |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
STF: Sentence Transformer Fine-Tuning For Topic Categorization With Limited Data
by: Daouadi, Kheir Eddine, et al.
Published: (2024)
by: Daouadi, Kheir Eddine, et al.
Published: (2024)
SciMantify -- A Hybrid Approach for the Evolving Semantification of Scientific Knowledge
by: John, Lena, et al.
Published: (2025)
by: John, Lena, et al.
Published: (2025)
Classification is a RAG problem: A case study on hate speech detection
by: Willats, Richard, et al.
Published: (2025)
by: Willats, Richard, et al.
Published: (2025)
MorphBPE: A Morpho-Aware Tokenizer Bridging Linguistic Complexity for Efficient LLM Training Across Morphologies
by: Asgari, Ehsaneddin, et al.
Published: (2025)
by: Asgari, Ehsaneddin, et al.
Published: (2025)
Can pre-trained language models generate titles for research papers?
by: Rehman, Tohida, et al.
Published: (2024)
by: Rehman, Tohida, et al.
Published: (2024)
The potential -- and the pitfalls -- of using pre-trained language models as cognitive science theories
by: Shah, Raj Sanjay, et al.
Published: (2025)
by: Shah, Raj Sanjay, et al.
Published: (2025)
What do self-supervised speech models know about Dutch? Analyzing advantages of language-specific pre-training
by: Kloots, Marianne de Heer, et al.
Published: (2025)
by: Kloots, Marianne de Heer, et al.
Published: (2025)
Assessment and manipulation of latent constructs in pre-trained language models using psychometric scales
by: Reuben, Maor, et al.
Published: (2024)
by: Reuben, Maor, et al.
Published: (2024)
WizardLM: Empowering large pre-trained language models to follow complex instructions
by: Xu, Can, et al.
Published: (2023)
by: Xu, Can, et al.
Published: (2023)
Retrieval-augmented reasoning with lean language models
by: Chan, Ryan Sze-Yin, et al.
Published: (2025)
by: Chan, Ryan Sze-Yin, et al.
Published: (2025)
Transformers and Ensemble methods: A solution for Hate Speech Detection in Arabic languages
by: de Paula, Angel Felipe Magnossão, et al.
Published: (2023)
by: de Paula, Angel Felipe Magnossão, et al.
Published: (2023)
Using LLMs to discover emerging coded antisemitic hate-speech in extremist social media
by: Kikkisetti, Dhanush, et al.
Published: (2024)
by: Kikkisetti, Dhanush, et al.
Published: (2024)
QUARTZ : QA-based Unsupervised Abstractive Refinement for Task-oriented Dialogue Summarization
by: Ghebriout, Mohamed Imed Eddine, et al.
Published: (2025)
by: Ghebriout, Mohamed Imed Eddine, et al.
Published: (2025)
Improving accuracy of GPT-3/4 results on biomedical data using a retrieval-augmented language model
by: Soong, David, et al.
Published: (2023)
by: Soong, David, et al.
Published: (2023)
Aleph-Alpha-GermanWeb: Improving German-language LLM pre-training with model-based data curation and synthetic data generation
by: Burns, Thomas F, et al.
Published: (2025)
by: Burns, Thomas F, et al.
Published: (2025)
Bridging the gap in online hate speech detection: a comparative analysis of BERT and traditional models for homophobic content identification on X/Twitter
by: McGiff, Josh, et al.
Published: (2024)
by: McGiff, Josh, et al.
Published: (2024)
Bringing legal knowledge to the public by constructing a legal question bank using large-scale pre-trained language model
by: Yuan, Mingruo, et al.
Published: (2025)
by: Yuan, Mingruo, et al.
Published: (2025)
AugSumm: towards generalizable speech summarization using synthetic labels from large language model
by: Jung, Jee-weon, et al.
Published: (2024)
by: Jung, Jee-weon, et al.
Published: (2024)
Enhancing textual textbook question answering with large language models and retrieval augmented generation
by: Alawwad, Hessa Abdulrahman, et al.
Published: (2024)
by: Alawwad, Hessa Abdulrahman, et al.
Published: (2024)
Arabic Tweet Act: A Weighted Ensemble Pre-Trained Transformer Model for Classifying Arabic Speech Acts on Twitter
by: Alshehri, Khadejaa, et al.
Published: (2024)
by: Alshehri, Khadejaa, et al.
Published: (2024)
Named Entity Recognition in COVID-19 tweets with Entity Knowledge Augmentation
by: Zhang, Xuankang, et al.
Published: (2025)
by: Zhang, Xuankang, et al.
Published: (2025)
What augmentations are sensitive to hyper-parameters and why?
by: Awais, Ch Muhammad, et al.
Published: (2021)
by: Awais, Ch Muhammad, et al.
Published: (2021)
Can social media provide early warning of retraction? Evidence from critical tweets identified by human annotation and large language models
by: Zheng, Er-Te, et al.
Published: (2024)
by: Zheng, Er-Te, et al.
Published: (2024)
A survey of textual cyber abuse detection using cutting-edge language models and large language models
by: Diaz-Garcia, Jose A., et al.
Published: (2025)
by: Diaz-Garcia, Jose A., et al.
Published: (2025)
"HOT" ChatGPT: The promise of ChatGPT in detecting and discriminating hateful, offensive, and toxic comments on social media
by: Li, Lingyao, et al.
Published: (2023)
by: Li, Lingyao, et al.
Published: (2023)
Augmenting emotion features in irony detection with Large language modeling
by: Lin, Yucheng, et al.
Published: (2024)
by: Lin, Yucheng, et al.
Published: (2024)
Vocabulary embeddings organize linguistic structure early in language model training
by: Papadimitriou, Isabel, et al.
Published: (2025)
by: Papadimitriou, Isabel, et al.
Published: (2025)
Language translation, and change of accent for speech-to-speech task using diffusion model
by: Mishra, Abhishek, et al.
Published: (2025)
by: Mishra, Abhishek, et al.
Published: (2025)
Falcon Mamba: The First Competitive Attention-free 7B Language Model
by: Zuo, Jingwei, et al.
Published: (2024)
by: Zuo, Jingwei, et al.
Published: (2024)
Retrieval augmented generation based dynamic prompting for few-shot biomedical named entity recognition using large language models
by: Ge, Yao, et al.
Published: (2025)
by: Ge, Yao, et al.
Published: (2025)
SmolKalam: Ensemble Quality-Filtered Translation at Scale for High Quality Arabic Post-Training Data
by: Alrashed, Sultan, et al.
Published: (2025)
by: Alrashed, Sultan, et al.
Published: (2025)
In-domain SSL pre-training and streaming ASR
by: Duret, Jarod, et al.
Published: (2025)
by: Duret, Jarod, et al.
Published: (2025)
Deep Memory Search: A Metaheuristic Approach for Optimizing Heuristic Search
by: Hedar, Abdel-Rahman, et al.
Published: (2024)
by: Hedar, Abdel-Rahman, et al.
Published: (2024)
Explainable cognitive decline detection in free dialogues with a Machine Learning approach based on pre-trained Large Language Models
by: de Arriba-Pérez, Francisco, et al.
Published: (2024)
by: de Arriba-Pérez, Francisco, et al.
Published: (2024)
Hope Speech Detection in code-mixed Roman Urdu tweets: A Positive Turn in Natural Language Processing
by: Ahmad, Muhammad, et al.
Published: (2025)
by: Ahmad, Muhammad, et al.
Published: (2025)
DepressLLM: Interpretable domain-adapted language model for depression detection from real-world narratives
by: Moon, Sehwan, et al.
Published: (2025)
by: Moon, Sehwan, et al.
Published: (2025)
StatLLaMA: Multi-Stage training for domain-optimized statistical large language models
by: Zeng, Jing-Yi, et al.
Published: (2025)
by: Zeng, Jing-Yi, et al.
Published: (2025)
Ukrainian-to-English folktale corpus: Parallel corpus creation and augmentation for machine translation in low-resource languages
by: Burda-Lassen, Olena
Published: (2024)
by: Burda-Lassen, Olena
Published: (2024)
LLMs and Finetuning: Benchmarking cross-domain performance for hate speech detection
by: Nasir, Ahmad, et al.
Published: (2023)
by: Nasir, Ahmad, et al.
Published: (2023)
The Arabic Generality Score: Another Dimension of Modeling Arabic Dialectness
by: Shaban, Sanad, et al.
Published: (2025)
by: Shaban, Sanad, et al.
Published: (2025)
Similar Items
-
STF: Sentence Transformer Fine-Tuning For Topic Categorization With Limited Data
by: Daouadi, Kheir Eddine, et al.
Published: (2024) -
SciMantify -- A Hybrid Approach for the Evolving Semantification of Scientific Knowledge
by: John, Lena, et al.
Published: (2025) -
Classification is a RAG problem: A case study on hate speech detection
by: Willats, Richard, et al.
Published: (2025) -
MorphBPE: A Morpho-Aware Tokenizer Bridging Linguistic Complexity for Efficient LLM Training Across Morphologies
by: Asgari, Ehsaneddin, et al.
Published: (2025) -
Can pre-trained language models generate titles for research papers?
by: Rehman, Tohida, et al.
Published: (2024)