TuPy-E: detecting hate speech in Brazilian Portuguese social media with a novel dataset and comprehensive analysis of models
Fuente:
arXiv
Saved in:
| Main Authors: | Oliveira, Felipe, Reis, Victoria, Ebecken, Nelson |
|---|---|
| Format: | Preprint |
| Published: |
2023
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
What is the social benefit of hate speech detection research? A Systematic Review
by: Wong, Sidney Gig-Jan
Published: (2024)
by: Wong, Sidney Gig-Jan
Published: (2024)
LLMs and Finetuning: Benchmarking cross-domain performance for hate speech detection
by: Nasir, Ahmad, et al.
Published: (2023)
by: Nasir, Ahmad, et al.
Published: (2023)
A multilingual dataset for offensive language and hate speech detection for hausa, yoruba and igbo languages
by: Aliyu, Saminu Mohammad, et al.
Published: (2024)
by: Aliyu, Saminu Mohammad, et al.
Published: (2024)
Using LLMs to discover emerging coded antisemitic hate-speech in extremist social media
by: Kikkisetti, Dhanush, et al.
Published: (2024)
by: Kikkisetti, Dhanush, et al.
Published: (2024)
Ensemble of pre-trained language models and data augmentation for hate speech detection from Arabic tweets
by: Daouadi, Kheir Eddine, et al.
Published: (2024)
by: Daouadi, Kheir Eddine, et al.
Published: (2024)
Classification is a RAG problem: A case study on hate speech detection
by: Willats, Richard, et al.
Published: (2025)
by: Willats, Richard, et al.
Published: (2025)
Tagarela - A Portuguese speech dataset from podcasts
by: de Oliveira, Frederico Santos, et al.
Published: (2026)
by: de Oliveira, Frederico Santos, et al.
Published: (2026)
Exploring the topics, sentiments and hate speech in the Spanish information environment
by: LOPEZ, ALEJANDRO BUITRAGO, et al.
Published: (2024)
by: LOPEZ, ALEJANDRO BUITRAGO, et al.
Published: (2024)
AtteSTNet -- An attention and subword tokenization based approach for code-switched text hate speech detection
by: Shingi, Geet, et al.
Published: (2021)
by: Shingi, Geet, et al.
Published: (2021)
"HOT" ChatGPT: The promise of ChatGPT in detecting and discriminating hateful, offensive, and toxic comments on social media
by: Li, Lingyao, et al.
Published: (2023)
by: Li, Lingyao, et al.
Published: (2023)
Bridging the gap in online hate speech detection: a comparative analysis of BERT and traditional models for homophobic content identification on X/Twitter
by: McGiff, Josh, et al.
Published: (2024)
by: McGiff, Josh, et al.
Published: (2024)
Digital Guardians: Can GPT-4, Perspective API, and Moderation API reliably detect hate speech in reader comments of German online newspapers?
by: Weber, Manuel, et al.
Published: (2025)
by: Weber, Manuel, et al.
Published: (2025)
Child-directed speech facilitates production, not comprehension, in BabyLMs
by: Bunzeck, Bastian, et al.
Published: (2026)
by: Bunzeck, Bastian, et al.
Published: (2026)
TeenyTinyLlama: open-source tiny language models trained in Brazilian Portuguese
by: Corrêa, Nicholas Kluge, et al.
Published: (2024)
by: Corrêa, Nicholas Kluge, et al.
Published: (2024)
Zero-shot Performance of Generative AI in Brazilian Portuguese Medical Exam
by: Truyts, Cesar Augusto Madid, et al.
Published: (2025)
by: Truyts, Cesar Augusto Madid, et al.
Published: (2025)
From Brazilian Portuguese to European Portuguese
by: Sanches, João, et al.
Published: (2024)
by: Sanches, João, et al.
Published: (2024)
Image captioning for Brazilian Portuguese using GRIT model
by: de Alencar, Rafael Silva, et al.
Published: (2024)
by: de Alencar, Rafael Silva, et al.
Published: (2024)
Improving code-mixed hate detection by native sample mixing: A case study for Hindi-English code-mixed scenario
by: Mazumder, Debajyoti, et al.
Published: (2024)
by: Mazumder, Debajyoti, et al.
Published: (2024)
PeLLE: Encoder-based language models for Brazilian Portuguese based on open data
by: de Mello, Guilherme Lamartine, et al.
Published: (2024)
by: de Mello, Guilherme Lamartine, et al.
Published: (2024)
Depression detection in social media posts using transformer-based models and auxiliary features
by: Kerasiotis, Marios, et al.
Published: (2024)
by: Kerasiotis, Marios, et al.
Published: (2024)
Disentangling segmental and prosodic factors to non-native speech comprehensibility
by: Quamer, Waris, et al.
Published: (2024)
by: Quamer, Waris, et al.
Published: (2024)
Leveraging language models for summarizing mental state examinations: A comprehensive evaluation and dataset release
by: Sahu, Nilesh Kumar, et al.
Published: (2024)
by: Sahu, Nilesh Kumar, et al.
Published: (2024)
TuCo: Measuring the Contribution of Fine-Tuning to Individual Responses of LLMs
by: Nuti, Felipe, et al.
Published: (2025)
by: Nuti, Felipe, et al.
Published: (2025)
Beyond surface form: A pipeline for semantic analysis in Alzheimer's Disease detection from spontaneous speech
by: Phelps, Dylan, et al.
Published: (2025)
by: Phelps, Dylan, et al.
Published: (2025)
Optimal strategies to perform multilingual analysis of social content for a novel dataset in the tourism domain
by: Masson, Maxime, et al.
Published: (2023)
by: Masson, Maxime, et al.
Published: (2023)
Hate speech detection in algerian dialect using deep learning
by: Lanasri, Dihia, et al.
Published: (2023)
by: Lanasri, Dihia, et al.
Published: (2023)
A comprehensive cross-language framework for harmful content detection with the aid of sentiment analysis
by: Dehghani, Mohammad
Published: (2024)
by: Dehghani, Mohammad
Published: (2024)
TheBlueScrubs-v1, a comprehensive curated medical dataset derived from the internet
by: Felipe, Luis, et al.
Published: (2025)
by: Felipe, Luis, et al.
Published: (2025)
A layer-wise analysis of Mandarin and English suprasegmentals in SSL speech models
by: de la Fuente, Antón, et al.
Published: (2024)
by: de la Fuente, Antón, et al.
Published: (2024)
Performance in a dialectal profiling task of LLMs for varieties of Brazilian Portuguese
by: Freitag, Raquel Meister Ko, et al.
Published: (2024)
by: Freitag, Raquel Meister Ko, et al.
Published: (2024)
The emojification of sentiment on social media: Collection and analysis of a longitudinal Twitter sentiment dataset
by: Yin, Wenjie, et al.
Published: (2021)
by: Yin, Wenjie, et al.
Published: (2021)
CAPITU: A Benchmark for Evaluating Instruction-Following in Brazilian Portuguese with Literary Context
by: Bonás, Giovana Kerche, et al.
Published: (2026)
by: Bonás, Giovana Kerche, et al.
Published: (2026)
MariNER: A Dataset for Historical Brazilian Portuguese Named Entity Recognition
by: Sarcinelli, João Lucas Luz Lima, et al.
Published: (2025)
by: Sarcinelli, João Lucas Luz Lima, et al.
Published: (2025)
Evaluating Named Entity Recognition: A comparative analysis of mono- and multilingual transformer models on a novel Brazilian corporate earnings call transcripts dataset
by: Abilio, Ramon, et al.
Published: (2024)
by: Abilio, Ramon, et al.
Published: (2024)
JurisTCU: A Brazilian Portuguese Information Retrieval Dataset with Query Relevance Judgments
by: Fernandes, Leandro Carísio, et al.
Published: (2025)
by: Fernandes, Leandro Carísio, et al.
Published: (2025)
LLM generated responses to mitigate the impact of hate speech
by: Podolak, Jakub, et al.
Published: (2023)
by: Podolak, Jakub, et al.
Published: (2023)
Brazilian Portuguese Image Captioning with Transformers: A Study on Cross-Native-Translated Dataset
by: Bromonschenkel, Gabriel, et al.
Published: (2026)
by: Bromonschenkel, Gabriel, et al.
Published: (2026)
MedPT: A Massive Medical Question Answering Dataset for Brazilian-Portuguese Speakers
by: Färber, Fernanda Bufon, et al.
Published: (2025)
by: Färber, Fernanda Bufon, et al.
Published: (2025)
Empowering machine learning models with contextual knowledge for enhancing the detection of eating disorders in social media posts
by: Benítez-Andrades, José Alberto, et al.
Published: (2024)
by: Benítez-Andrades, José Alberto, et al.
Published: (2024)
Language translation, and change of accent for speech-to-speech task using diffusion model
by: Mishra, Abhishek, et al.
Published: (2025)
by: Mishra, Abhishek, et al.
Published: (2025)
Similar Items
-
What is the social benefit of hate speech detection research? A Systematic Review
by: Wong, Sidney Gig-Jan
Published: (2024) -
LLMs and Finetuning: Benchmarking cross-domain performance for hate speech detection
by: Nasir, Ahmad, et al.
Published: (2023) -
A multilingual dataset for offensive language and hate speech detection for hausa, yoruba and igbo languages
by: Aliyu, Saminu Mohammad, et al.
Published: (2024) -
Using LLMs to discover emerging coded antisemitic hate-speech in extremist social media
by: Kikkisetti, Dhanush, et al.
Published: (2024) -
Ensemble of pre-trained language models and data augmentation for hate speech detection from Arabic tweets
by: Daouadi, Kheir Eddine, et al.
Published: (2024)