DADIT: A Dataset for Demographic Classification of Italian Twitter Users and a Comparison of Prediction Methods
Fuente:
arXiv
Saved in:
| Main Authors: | Lupo, Lorenzo, Bose, Paul, Habibi, Mahyar, Hovy, Dirk, Schwarz, Carlo |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
The Content Moderator's Dilemma: Removal of Toxic Content and Distortions to Online Discourse
by: Habibi, Mahyar, et al.
Published: (2024)
by: Habibi, Mahyar, et al.
Published: (2024)
Compromesso! Italian Many-Shot Jailbreaks Undermine the Safety of Large Language Models
by: Pernisi, Fabio, et al.
Published: (2024)
by: Pernisi, Fabio, et al.
Published: (2024)
Towards Human-Level Text Coding with LLMs: The Case of Fatherhood Roles in Public Policy Documents
by: Lupo, Lorenzo, et al.
Published: (2023)
by: Lupo, Lorenzo, et al.
Published: (2023)
SimBench: Benchmarking the Ability of Large Language Models to Simulate Human Behaviors
by: Hu, Tiancheng, et al.
Published: (2025)
by: Hu, Tiancheng, et al.
Published: (2025)
Beyond Demographics: Fine-tuning Large Language Models to Predict Individuals' Subjective Text Perceptions
by: Orlikowski, Matthias, et al.
Published: (2025)
by: Orlikowski, Matthias, et al.
Published: (2025)
Beyond Flesch-Kincaid: Prompt-based Metrics Improve Difficulty Classification of Educational Texts
by: Rooein, Donya, et al.
Published: (2024)
by: Rooein, Donya, et al.
Published: (2024)
SafetyPrompts: a Systematic Review of Open Datasets for Evaluating and Improving Large Language Model Safety
by: Röttger, Paul, et al.
Published: (2024)
by: Röttger, Paul, et al.
Published: (2024)
Conversations as a Source for Teaching Scientific Concepts at Different Education Levels
by: Rooein, Donya, et al.
Published: (2024)
by: Rooein, Donya, et al.
Published: (2024)
Narratives at Conflict: Computational Analysis of News Framing in Multilingual Disinformation Campaigns
by: Sinelnik, Antonina, et al.
Published: (2024)
by: Sinelnik, Antonina, et al.
Published: (2024)
ITALIC: An Italian Intent Classification Dataset
by: Koudounas, Alkis, et al.
Published: (2023)
by: Koudounas, Alkis, et al.
Published: (2023)
The Ecological Fallacy in Annotation: Modelling Human Label Variation goes beyond Sociodemographics
by: Orlikowski, Matthias, et al.
Published: (2023)
by: Orlikowski, Matthias, et al.
Published: (2023)
Do Prompts Reshape Representations? An Empirical Study of Prompting Effects on Embeddings
by: Gonzalez-Gutierrez, Cesar, et al.
Published: (2025)
by: Gonzalez-Gutierrez, Cesar, et al.
Published: (2025)
The Pluralistic Moral Gap: Understanding Judgment and Value Differences between Humans and Large Language Models
by: Russo, Giuseppe, et al.
Published: (2025)
by: Russo, Giuseppe, et al.
Published: (2025)
The AI Gap: How Socioeconomic Status Affects Language Technology Interactions
by: Bassignana, Elisa, et al.
Published: (2025)
by: Bassignana, Elisa, et al.
Published: (2025)
Principled Personas: Defining and Measuring the Intended Effects of Persona Prompting on Task Performance
by: de Araujo, Pedro Henrique Luz, et al.
Published: (2025)
by: de Araujo, Pedro Henrique Luz, et al.
Published: (2025)
The Call for Socially Aware Language Technologies
by: Yang, Diyi, et al.
Published: (2024)
by: Yang, Diyi, et al.
Published: (2024)
Twists, Humps, and Pebbles: Multilingual Speech Recognition Models Exhibit Gender Performance Gaps
by: Attanasio, Giuseppe, et al.
Published: (2024)
by: Attanasio, Giuseppe, et al.
Published: (2024)
Biased Tales: Cultural and Topic Bias in Generating Children's Stories
by: Rooein, Donya, et al.
Published: (2025)
by: Rooein, Donya, et al.
Published: (2025)
Impoverished Language Technology: The Lack of (Social) Class in NLP
by: Curry, Amanda Cercas, et al.
Published: (2024)
by: Curry, Amanda Cercas, et al.
Published: (2024)
Classist Tools: Social Class Correlates with Performance in NLP
by: Curry, Amanda Cercas, et al.
Published: (2024)
by: Curry, Amanda Cercas, et al.
Published: (2024)
Wisdom of Instruction-Tuned Language Model Crowds. Exploring Model Label Variation
by: Plaza-del-Arco, Flor Miriam, et al.
Published: (2023)
by: Plaza-del-Arco, Flor Miriam, et al.
Published: (2023)
Do Large Language Models Adapt to Language Variation across Socioeconomic Status?
by: Bassignana, Elisa, et al.
Published: (2026)
by: Bassignana, Elisa, et al.
Published: (2026)
LLMs for Argument Mining: Detection, Extraction, and Relationship Classification of pre-defined Arguments in Online Comments
by: Guida, Matteo, et al.
Published: (2025)
by: Guida, Matteo, et al.
Published: (2025)
Exploring Subjective Tasks in Farsi: A Survey Analysis and Evaluation of Language Models
by: Rooein, Donya, et al.
Published: (2025)
by: Rooein, Donya, et al.
Published: (2025)
EcoVerse: An Annotated Twitter Dataset for Eco-Relevance Classification, Environmental Impact Analysis, and Stance Detection
by: Grasso, Francesca, et al.
Published: (2024)
by: Grasso, Francesca, et al.
Published: (2024)
XSTest: A Test Suite for Identifying Exaggerated Safety Behaviours in Large Language Models
by: Röttger, Paul, et al.
Published: (2023)
by: Röttger, Paul, et al.
Published: (2023)
Diffusion Language Models Are Natively Length-Aware
by: Rossi, Vittorio, et al.
Published: (2026)
by: Rossi, Vittorio, et al.
Published: (2026)
IssueBench: Millions of Realistic Prompts for Measuring Issue Bias in LLM Writing Assistance
by: Röttger, Paul, et al.
Published: (2025)
by: Röttger, Paul, et al.
Published: (2025)
Consistency is Key: Disentangling Label Variation in Natural Language Processing with Intra-Annotator Agreement
by: Abercrombie, Gavin, et al.
Published: (2023)
by: Abercrombie, Gavin, et al.
Published: (2023)
No for Some, Yes for Others: Persona Prompts and Other Sources of False Refusal in Language Models
by: Plaza-del-Arco, Flor Miriam, et al.
Published: (2025)
by: Plaza-del-Arco, Flor Miriam, et al.
Published: (2025)
HateDay: Insights from a Global Hate Speech Dataset Representative of a Day on Twitter
by: Tonneau, Manuel, et al.
Published: (2024)
by: Tonneau, Manuel, et al.
Published: (2024)
RTI-Bench: A Structured Dataset for Indian Right-to-Information Decision Analysis
by: Bose, Joy
Published: (2026)
by: Bose, Joy
Published: (2026)
Emotion Analysis in NLP: Trends, Gaps and Roadmap for Future Directions
by: Plaza-del-Arco, Flor Miriam, et al.
Published: (2024)
by: Plaza-del-Arco, Flor Miriam, et al.
Published: (2024)
Comparing Pre-trained Human Language Models: Is it Better with Human Context as Groups, Individual Traits, or Both?
by: Soni, Nikita, et al.
Published: (2024)
by: Soni, Nikita, et al.
Published: (2024)
ProvocationProbe: Instigating Hate Speech Dataset from Twitter
by: Kumar, Abhay, et al.
Published: (2024)
by: Kumar, Abhay, et al.
Published: (2024)
"My Answer is C": First-Token Probabilities Do Not Match Text Answers in Instruction-Tuned Language Models
by: Wang, Xinpeng, et al.
Published: (2024)
by: Wang, Xinpeng, et al.
Published: (2024)
Enriching Datasets with Demographics through Large Language Models: What's in a Name?
by: AlNuaimi, Khaled, et al.
Published: (2024)
by: AlNuaimi, Khaled, et al.
Published: (2024)
Triggered: A Statistical Analysis of Environmental Influences on Extremist Groups
by: de Kock, Christine, et al.
Published: (2026)
by: de Kock, Christine, et al.
Published: (2026)
Political Compass or Spinning Arrow? Towards More Meaningful Evaluations for Values and Opinions in Large Language Models
by: Röttger, Paul, et al.
Published: (2024)
by: Röttger, Paul, et al.
Published: (2024)
RideKE: Leveraging Low-Resource, User-Generated Twitter Content for Sentiment and Emotion Detection in Kenyan Code-Switched Dataset
by: Etori, Naome A., et al.
Published: (2025)
by: Etori, Naome A., et al.
Published: (2025)
Similar Items
-
The Content Moderator's Dilemma: Removal of Toxic Content and Distortions to Online Discourse
by: Habibi, Mahyar, et al.
Published: (2024) -
Compromesso! Italian Many-Shot Jailbreaks Undermine the Safety of Large Language Models
by: Pernisi, Fabio, et al.
Published: (2024) -
Towards Human-Level Text Coding with LLMs: The Case of Fatherhood Roles in Public Policy Documents
by: Lupo, Lorenzo, et al.
Published: (2023) -
SimBench: Benchmarking the Ability of Large Language Models to Simulate Human Behaviors
by: Hu, Tiancheng, et al.
Published: (2025) -
Beyond Demographics: Fine-tuning Large Language Models to Predict Individuals' Subjective Text Perceptions
by: Orlikowski, Matthias, et al.
Published: (2025)