Ensemble of pre-trained language models and data augmentation for hate speech detection from Arabic tweets
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Daouadi, Kheir Eddine, Boualleg, Yaakoub, Haouaouchi, Kheir Eddine |
|---|---|
| Format: | Preprint |
| Publié: |
2024
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
STF: Sentence Transformer Fine-Tuning For Topic Categorization With Limited Data
par: Daouadi, Kheir Eddine, et autres
Publié: (2024)
par: Daouadi, Kheir Eddine, et autres
Publié: (2024)
SciMantify -- A Hybrid Approach for the Evolving Semantification of Scientific Knowledge
par: John, Lena, et autres
Publié: (2025)
par: John, Lena, et autres
Publié: (2025)
Classification is a RAG problem: A case study on hate speech detection
par: Willats, Richard, et autres
Publié: (2025)
par: Willats, Richard, et autres
Publié: (2025)
MorphBPE: A Morpho-Aware Tokenizer Bridging Linguistic Complexity for Efficient LLM Training Across Morphologies
par: Asgari, Ehsaneddin, et autres
Publié: (2025)
par: Asgari, Ehsaneddin, et autres
Publié: (2025)
Can pre-trained language models generate titles for research papers?
par: Rehman, Tohida, et autres
Publié: (2024)
par: Rehman, Tohida, et autres
Publié: (2024)
The potential -- and the pitfalls -- of using pre-trained language models as cognitive science theories
par: Shah, Raj Sanjay, et autres
Publié: (2025)
par: Shah, Raj Sanjay, et autres
Publié: (2025)
What do self-supervised speech models know about Dutch? Analyzing advantages of language-specific pre-training
par: Kloots, Marianne de Heer, et autres
Publié: (2025)
par: Kloots, Marianne de Heer, et autres
Publié: (2025)
Assessment and manipulation of latent constructs in pre-trained language models using psychometric scales
par: Reuben, Maor, et autres
Publié: (2024)
par: Reuben, Maor, et autres
Publié: (2024)
WizardLM: Empowering large pre-trained language models to follow complex instructions
par: Xu, Can, et autres
Publié: (2023)
par: Xu, Can, et autres
Publié: (2023)
Retrieval-augmented reasoning with lean language models
par: Chan, Ryan Sze-Yin, et autres
Publié: (2025)
par: Chan, Ryan Sze-Yin, et autres
Publié: (2025)
Transformers and Ensemble methods: A solution for Hate Speech Detection in Arabic languages
par: de Paula, Angel Felipe Magnossão, et autres
Publié: (2023)
par: de Paula, Angel Felipe Magnossão, et autres
Publié: (2023)
Using LLMs to discover emerging coded antisemitic hate-speech in extremist social media
par: Kikkisetti, Dhanush, et autres
Publié: (2024)
par: Kikkisetti, Dhanush, et autres
Publié: (2024)
QUARTZ : QA-based Unsupervised Abstractive Refinement for Task-oriented Dialogue Summarization
par: Ghebriout, Mohamed Imed Eddine, et autres
Publié: (2025)
par: Ghebriout, Mohamed Imed Eddine, et autres
Publié: (2025)
Improving accuracy of GPT-3/4 results on biomedical data using a retrieval-augmented language model
par: Soong, David, et autres
Publié: (2023)
par: Soong, David, et autres
Publié: (2023)
Aleph-Alpha-GermanWeb: Improving German-language LLM pre-training with model-based data curation and synthetic data generation
par: Burns, Thomas F, et autres
Publié: (2025)
par: Burns, Thomas F, et autres
Publié: (2025)
Bridging the gap in online hate speech detection: a comparative analysis of BERT and traditional models for homophobic content identification on X/Twitter
par: McGiff, Josh, et autres
Publié: (2024)
par: McGiff, Josh, et autres
Publié: (2024)
Bringing legal knowledge to the public by constructing a legal question bank using large-scale pre-trained language model
par: Yuan, Mingruo, et autres
Publié: (2025)
par: Yuan, Mingruo, et autres
Publié: (2025)
AugSumm: towards generalizable speech summarization using synthetic labels from large language model
par: Jung, Jee-weon, et autres
Publié: (2024)
par: Jung, Jee-weon, et autres
Publié: (2024)
Enhancing textual textbook question answering with large language models and retrieval augmented generation
par: Alawwad, Hessa Abdulrahman, et autres
Publié: (2024)
par: Alawwad, Hessa Abdulrahman, et autres
Publié: (2024)
Arabic Tweet Act: A Weighted Ensemble Pre-Trained Transformer Model for Classifying Arabic Speech Acts on Twitter
par: Alshehri, Khadejaa, et autres
Publié: (2024)
par: Alshehri, Khadejaa, et autres
Publié: (2024)
Named Entity Recognition in COVID-19 tweets with Entity Knowledge Augmentation
par: Zhang, Xuankang, et autres
Publié: (2025)
par: Zhang, Xuankang, et autres
Publié: (2025)
What augmentations are sensitive to hyper-parameters and why?
par: Awais, Ch Muhammad, et autres
Publié: (2021)
par: Awais, Ch Muhammad, et autres
Publié: (2021)
Can social media provide early warning of retraction? Evidence from critical tweets identified by human annotation and large language models
par: Zheng, Er-Te, et autres
Publié: (2024)
par: Zheng, Er-Te, et autres
Publié: (2024)
A survey of textual cyber abuse detection using cutting-edge language models and large language models
par: Diaz-Garcia, Jose A., et autres
Publié: (2025)
par: Diaz-Garcia, Jose A., et autres
Publié: (2025)
"HOT" ChatGPT: The promise of ChatGPT in detecting and discriminating hateful, offensive, and toxic comments on social media
par: Li, Lingyao, et autres
Publié: (2023)
par: Li, Lingyao, et autres
Publié: (2023)
Augmenting emotion features in irony detection with Large language modeling
par: Lin, Yucheng, et autres
Publié: (2024)
par: Lin, Yucheng, et autres
Publié: (2024)
Vocabulary embeddings organize linguistic structure early in language model training
par: Papadimitriou, Isabel, et autres
Publié: (2025)
par: Papadimitriou, Isabel, et autres
Publié: (2025)
Language translation, and change of accent for speech-to-speech task using diffusion model
par: Mishra, Abhishek, et autres
Publié: (2025)
par: Mishra, Abhishek, et autres
Publié: (2025)
Falcon Mamba: The First Competitive Attention-free 7B Language Model
par: Zuo, Jingwei, et autres
Publié: (2024)
par: Zuo, Jingwei, et autres
Publié: (2024)
Retrieval augmented generation based dynamic prompting for few-shot biomedical named entity recognition using large language models
par: Ge, Yao, et autres
Publié: (2025)
par: Ge, Yao, et autres
Publié: (2025)
SmolKalam: Ensemble Quality-Filtered Translation at Scale for High Quality Arabic Post-Training Data
par: Alrashed, Sultan, et autres
Publié: (2025)
par: Alrashed, Sultan, et autres
Publié: (2025)
In-domain SSL pre-training and streaming ASR
par: Duret, Jarod, et autres
Publié: (2025)
par: Duret, Jarod, et autres
Publié: (2025)
Deep Memory Search: A Metaheuristic Approach for Optimizing Heuristic Search
par: Hedar, Abdel-Rahman, et autres
Publié: (2024)
par: Hedar, Abdel-Rahman, et autres
Publié: (2024)
Explainable cognitive decline detection in free dialogues with a Machine Learning approach based on pre-trained Large Language Models
par: de Arriba-Pérez, Francisco, et autres
Publié: (2024)
par: de Arriba-Pérez, Francisco, et autres
Publié: (2024)
Hope Speech Detection in code-mixed Roman Urdu tweets: A Positive Turn in Natural Language Processing
par: Ahmad, Muhammad, et autres
Publié: (2025)
par: Ahmad, Muhammad, et autres
Publié: (2025)
DepressLLM: Interpretable domain-adapted language model for depression detection from real-world narratives
par: Moon, Sehwan, et autres
Publié: (2025)
par: Moon, Sehwan, et autres
Publié: (2025)
StatLLaMA: Multi-Stage training for domain-optimized statistical large language models
par: Zeng, Jing-Yi, et autres
Publié: (2025)
par: Zeng, Jing-Yi, et autres
Publié: (2025)
Ukrainian-to-English folktale corpus: Parallel corpus creation and augmentation for machine translation in low-resource languages
par: Burda-Lassen, Olena
Publié: (2024)
par: Burda-Lassen, Olena
Publié: (2024)
LLMs and Finetuning: Benchmarking cross-domain performance for hate speech detection
par: Nasir, Ahmad, et autres
Publié: (2023)
par: Nasir, Ahmad, et autres
Publié: (2023)
The Arabic Generality Score: Another Dimension of Modeling Arabic Dialectness
par: Shaban, Sanad, et autres
Publié: (2025)
par: Shaban, Sanad, et autres
Publié: (2025)
Documents similaires
-
STF: Sentence Transformer Fine-Tuning For Topic Categorization With Limited Data
par: Daouadi, Kheir Eddine, et autres
Publié: (2024) -
SciMantify -- A Hybrid Approach for the Evolving Semantification of Scientific Knowledge
par: John, Lena, et autres
Publié: (2025) -
Classification is a RAG problem: A case study on hate speech detection
par: Willats, Richard, et autres
Publié: (2025) -
MorphBPE: A Morpho-Aware Tokenizer Bridging Linguistic Complexity for Efficient LLM Training Across Morphologies
par: Asgari, Ehsaneddin, et autres
Publié: (2025) -
Can pre-trained language models generate titles for research papers?
par: Rehman, Tohida, et autres
Publié: (2024)