Towards Better Inclusivity: A Diverse Tweet Corpus of English Varieties
Fuente:
arXiv
Saved in:
| Main Authors: | Pham, Nhi, Pham, Lachlan, Meyers, Adam L. |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
EmoScan: Automatic Screening of Depression Symptoms in Romanized Sinhala Tweets
by: Hewapathirana, Jayathi, et al.
Published: (2024)
by: Hewapathirana, Jayathi, et al.
Published: (2024)
Towards Better Health Conversations: The Benefits of Context-seeking
by: Sayres, Rory, et al.
Published: (2025)
by: Sayres, Rory, et al.
Published: (2025)
MAFIA: Multi-Adapter Fused Inclusive LanguAge Models
by: Jain, Prachi, et al.
Published: (2024)
by: Jain, Prachi, et al.
Published: (2024)
Mapping the Podcast Ecosystem with the Structured Podcast Research Corpus
by: Litterer, Benjamin, et al.
Published: (2024)
by: Litterer, Benjamin, et al.
Published: (2024)
Reading Between the Tweets: Deciphering Ideological Stances of Interconnected Mixed-Ideology Communities
by: He, Zihao, et al.
Published: (2024)
by: He, Zihao, et al.
Published: (2024)
The Cambridge Law Corpus: A Dataset for Legal AI Research
by: Östling, Andreas, et al.
Published: (2023)
by: Östling, Andreas, et al.
Published: (2023)
An Annotated Corpus of Arabic Tweets for Hate Speech Analysis
by: Zaghouani, Wajdi, et al.
Published: (2025)
by: Zaghouani, Wajdi, et al.
Published: (2025)
SocialNLP Fake-EmoReact 2021 Challenge Overview: Predicting Fake Tweets from Their Replies and GIFs
by: Huang, Chien-Kun, et al.
Published: (2024)
by: Huang, Chien-Kun, et al.
Published: (2024)
The Moral Foundations Reddit Corpus
by: Trager, Jackson, et al.
Published: (2022)
by: Trager, Jackson, et al.
Published: (2022)
SPOT: An Annotated French Corpus and Benchmark for Detecting Critical Interventions in Online Conversations
by: Berriche, Manon, et al.
Published: (2025)
by: Berriche, Manon, et al.
Published: (2025)
Mechanism Plausibility in Generative Agent-Based Modeling
by: Zhao, Patrick, et al.
Published: (2026)
by: Zhao, Patrick, et al.
Published: (2026)
Better Together: Quantifying the Benefits of AI-Assisted Recruitment
by: Aka, Ada, et al.
Published: (2025)
by: Aka, Ada, et al.
Published: (2025)
ARCADE: A City-Scale Corpus for Fine-Grained Arabic Dialect Tagging
by: Nacar, Omer, et al.
Published: (2026)
by: Nacar, Omer, et al.
Published: (2026)
Divergent Emotional Patterns in Disinformation on Social Media? An Analysis of Tweets and TikToks about the DANA in Valencia
by: Arcos, Iván, et al.
Published: (2025)
by: Arcos, Iván, et al.
Published: (2025)
Improving Stance Detection by Leveraging Measurement Knowledge from Social Sciences: A Case Study of Dutch Political Tweets and Traditional Gender Role Division
by: Fang, Qixiang, et al.
Published: (2022)
by: Fang, Qixiang, et al.
Published: (2022)
SLM-Bench: A Comprehensive Benchmark of Small Language Models on Environmental Impacts--Extended Version
by: Pham, Nghiem Thanh, et al.
Published: (2025)
by: Pham, Nghiem Thanh, et al.
Published: (2025)
SMILE: Single-turn to Multi-turn Inclusive Language Expansion via ChatGPT for Mental Health Support
by: Qiu, Huachuan, et al.
Published: (2023)
by: Qiu, Huachuan, et al.
Published: (2023)
Corpus-Based Approaches to Igbo Diacritic Restoration
by: Ezeani, Ignatius
Published: (2026)
by: Ezeani, Ignatius
Published: (2026)
The Better Angels of Machine Personality: How Personality Relates to LLM Safety
by: Zhang, Jie, et al.
Published: (2024)
by: Zhang, Jie, et al.
Published: (2024)
Better Call GPT, Comparing Large Language Models Against Lawyers
by: Martin, Lauren, et al.
Published: (2024)
by: Martin, Lauren, et al.
Published: (2024)
Extending a Parliamentary Corpus with MPs' Tweets: Automatic Annotation and Evaluation Using MultiParTweet
by: Bagci, Mevlüt, et al.
Published: (2025)
by: Bagci, Mevlüt, et al.
Published: (2025)
The Homogenization Problem in LLMs: Towards Meaningful Diversity in AI Safety
by: Rios-Sialer, Ian
Published: (2026)
by: Rios-Sialer, Ian
Published: (2026)
Extracting O*NET Features from the NLx Corpus to Build Public Use Aggregate Labor Market Data
by: Meisenbacher, Stephen, et al.
Published: (2025)
by: Meisenbacher, Stephen, et al.
Published: (2025)
EquiSumm : A Gender Bias-Aware Framework for Inclusive Tweet Summarization
by: Wanjari, Chaitanya, et al.
Published: (2026)
by: Wanjari, Chaitanya, et al.
Published: (2026)
Evaluating Digital Inclusiveness of Digital Agri-Food Tools Using Large Language Models: A Comparative Analysis Between Human and AI-Based Evaluations
by: Pewinya, Githma, et al.
Published: (2026)
by: Pewinya, Githma, et al.
Published: (2026)
Building Better: Avoiding Pitfalls in Developing Language Resources when Data is Scarce
by: Ousidhoum, Nedjma, et al.
Published: (2024)
by: Ousidhoum, Nedjma, et al.
Published: (2024)
Ethical Concern Identification in NLP: A Corpus of ACL Anthology Ethics Statements
by: Karamolegkou, Antonia, et al.
Published: (2024)
by: Karamolegkou, Antonia, et al.
Published: (2024)
GPT-HateCheck: Can LLMs Write Better Functional Tests for Hate Speech Detection?
by: Jin, Yiping, et al.
Published: (2024)
by: Jin, Yiping, et al.
Published: (2024)
Do Llamas Work in English? On the Latent Language of Multilingual Transformers
by: Wendler, Chris, et al.
Published: (2024)
by: Wendler, Chris, et al.
Published: (2024)
Beyond English: Unveiling Multilingual Bias in LLM Copyright Compliance
by: Chen, Yupeng, et al.
Published: (2025)
by: Chen, Yupeng, et al.
Published: (2025)
Benchmarking Sociolinguistic Diversity in Swahili NLP: A Taxonomy-Guided Approach
by: Oketch, Kezia, et al.
Published: (2025)
by: Oketch, Kezia, et al.
Published: (2025)
Robust Pronoun Fidelity with English LLMs: Are they Reasoning, Repeating, or Just Biased?
by: Gautam, Vagrant, et al.
Published: (2024)
by: Gautam, Vagrant, et al.
Published: (2024)
Multilingual Prompting for Improving LLM Generation Diversity
by: Wang, Qihan, et al.
Published: (2025)
by: Wang, Qihan, et al.
Published: (2025)
Can Large Language Models Understand You Better? An MBTI Personality Detection Dataset Aligned with Population Traits
by: Li, Bohan, et al.
Published: (2024)
by: Li, Bohan, et al.
Published: (2024)
People Make Better Edits: Measuring the Efficacy of LLM-Generated Counterfactually Augmented Data for Harmful Language Detection
by: Sen, Indira, et al.
Published: (2023)
by: Sen, Indira, et al.
Published: (2023)
Words of Warmth: Trust and Sociability Norms for over 26k English Words
by: Mohammad, Saif M.
Published: (2025)
by: Mohammad, Saif M.
Published: (2025)
Inverse Scaling: When Bigger Isn't Better
by: McKenzie, Ian R., et al.
Published: (2023)
by: McKenzie, Ian R., et al.
Published: (2023)
LangLingual: A Personalised, Exercise-oriented English Language Learning Tool Leveraging Large Language Models
by: Gupta, Sammriddh, et al.
Published: (2025)
by: Gupta, Sammriddh, et al.
Published: (2025)
What is Stigma Attributed to? A Theory-Grounded, Expert-Annotated Interview Corpus for Demystifying Mental-Health Stigma
by: Meng, Han, et al.
Published: (2025)
by: Meng, Han, et al.
Published: (2025)
Language of Thought Shapes Output Diversity in Large Language Models
by: Xu, Shaoyang, et al.
Published: (2026)
by: Xu, Shaoyang, et al.
Published: (2026)
Similar Items
-
EmoScan: Automatic Screening of Depression Symptoms in Romanized Sinhala Tweets
by: Hewapathirana, Jayathi, et al.
Published: (2024) -
Towards Better Health Conversations: The Benefits of Context-seeking
by: Sayres, Rory, et al.
Published: (2025) -
MAFIA: Multi-Adapter Fused Inclusive LanguAge Models
by: Jain, Prachi, et al.
Published: (2024) -
Mapping the Podcast Ecosystem with the Structured Podcast Research Corpus
by: Litterer, Benjamin, et al.
Published: (2024) -
Reading Between the Tweets: Deciphering Ideological Stances of Interconnected Mixed-Ideology Communities
by: He, Zihao, et al.
Published: (2024)