Curating corpora with classifiers: A case study of clean energy sentiment online
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Arnold, Michael V., Dodds, Peter Sheridan, Danforth, Christopher M. |
|---|---|
| Format: | Preprint |
| Publié: |
2023
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
LLMs for Low-Resource Dialect Translation Using Context-Aware Prompting: A Case Study on Sylheti
par: Prama, Tabia Tanzin, et autres
Publié: (2025)
par: Prama, Tabia Tanzin, et autres
Publié: (2025)
Misalignment of LLM-Generated Personas with Human Perceptions in Low-Resource Settings
par: Prama, Tabia Tanzin, et autres
Publié: (2025)
par: Prama, Tabia Tanzin, et autres
Publié: (2025)
Hollywood's misrepresentation of death: A comparison of overall and by-gender mortality causes in film and the real world
par: Beauregard, Calla, et autres
Publié: (2024)
par: Beauregard, Calla, et autres
Publié: (2024)
BanglaMATH : A Bangla benchmark dataset for testing LLM mathematical reasoning at grades 6, 7, and 8
par: Prama, Tabia Tanzin, et autres
Publié: (2025)
par: Prama, Tabia Tanzin, et autres
Publié: (2025)
Story and essential meaning dynamics in Bangladesh's July 2024 Student-People's Uprising
par: Prama, Tabia Tanzin, et autres
Publié: (2025)
par: Prama, Tabia Tanzin, et autres
Publié: (2025)
Statistical laws and linguistics inform meaning in naturalistic and fictional conversation
par: Fehr, Ashley M. A., et autres
Publié: (2025)
par: Fehr, Ashley M. A., et autres
Publié: (2025)
Us-vs-Them bias in Large Language Models
par: Prama, Tabia Tanzin, et autres
Publié: (2025)
par: Prama, Tabia Tanzin, et autres
Publié: (2025)
Complete asymptotic type-token relationship for growing complex systems with inverse power-law count rankings
par: Rosillo-Rodes, Pablo, et autres
Publié: (2025)
par: Rosillo-Rodes, Pablo, et autres
Publié: (2025)
Archetypes and gender in fiction: A data-driven mapping of gender stereotypes in stories
par: Beauregard, Calla Glavin, et autres
Publié: (2026)
par: Beauregard, Calla Glavin, et autres
Publié: (2026)
Global brain drain and gain in high-potential student mobility
par: Pramaa, Tabia Tanzin, et autres
Publié: (2026)
par: Pramaa, Tabia Tanzin, et autres
Publié: (2026)
A suite of allotaxonometric tools for the comparison of complex systems using rank-turbulence divergence
par: St-Onge, Jonathan, et autres
Publié: (2025)
par: St-Onge, Jonathan, et autres
Publié: (2025)
Detecting sub-populations in online health communities: A mixed-methods exploration of breastfeeding messages in BabyCenter Birth Clubs
par: Beauregard, Calla, et autres
Publié: (2025)
par: Beauregard, Calla, et autres
Publié: (2025)
Aim High, Stay Private: Differentially Private Synthetic Data Enables Public Release of Behavioral Health Information with High Utility
par: Ghasemizade, Mohsen, et autres
Publié: (2025)
par: Ghasemizade, Mohsen, et autres
Publié: (2025)
Entropy and type-token ratio in gigaword corpora
par: Rosillo-Rodes, Pablo, et autres
Publié: (2024)
par: Rosillo-Rodes, Pablo, et autres
Publié: (2024)
Leveraging Machine Learning to Detect Data Curation Activities
par: Lafia, Sara, et autres
Publié: (2021)
par: Lafia, Sara, et autres
Publié: (2021)
Crisis-induced differences in attention towards Ukraine in Twitter 2008-2023
par: Mets, Mark, et autres
Publié: (2026)
par: Mets, Mark, et autres
Publié: (2026)
A Formative Study of Brief Affective Text as a Complement to Wearable Sensing for Longitudinal Student Health Monitoring
par: Harry, Tamunotonye, et autres
Publié: (2026)
par: Harry, Tamunotonye, et autres
Publié: (2026)
A blind spot for large language models: Supradiegetic linguistic information
par: Zimmerman, Julia Witte, et autres
Publié: (2023)
par: Zimmerman, Julia Witte, et autres
Publié: (2023)
Involvement drives complexity of language in online debates
par: Amadori, Eleonora, et autres
Publié: (2025)
par: Amadori, Eleonora, et autres
Publié: (2025)
Identifying Body Composition Measures That Correlate with Self-Compassion and Social Support
par: Poon, Enerson, et autres
Publié: (2026)
par: Poon, Enerson, et autres
Publié: (2026)
Bloom-epistemic and sentiment analysis hierarchical classification in course discussion forums
par: Toba, H., et autres
Publié: (2024)
par: Toba, H., et autres
Publié: (2024)
Using LLMs to create analytical datasets: A case study of reconstructing the historical memory of Colombia
par: Anderson, David, et autres
Publié: (2025)
par: Anderson, David, et autres
Publié: (2025)
Taste for Privacy: How Context, Identity, and Lived-Experience Shape Information Sharing Preferences
par: Lovato, Juniper, et autres
Publié: (2026)
par: Lovato, Juniper, et autres
Publié: (2026)
The Three Books of Science
par: Dodds, Peter Sheridan
Publié: (2025)
par: Dodds, Peter Sheridan
Publié: (2025)
Divided by discipline? A systematic literature review on the quantification of online sexism and misogyny using a semi-automated approach
par: Dutta, Aditi, et autres
Publié: (2024)
par: Dutta, Aditi, et autres
Publié: (2024)
Polarization by Default: Auditing Recommendation Bias in LLM-Based Content Curation
par: Pagan, Nicolò, et autres
Publié: (2026)
par: Pagan, Nicolò, et autres
Publié: (2026)
Large Language Models Require Curated Context for Reliable Political Fact-Checking -- Even with Reasoning and Web Search
par: DeVerna, Matthew R., et autres
Publié: (2025)
par: DeVerna, Matthew R., et autres
Publié: (2025)
Building low-resource African language corpora: A case study of Kidawida, Kalenjin and Dholuo
par: Mbogho, Audrey, et autres
Publié: (2025)
par: Mbogho, Audrey, et autres
Publié: (2025)
Comparing energy consumption and accuracy in text classification inference
par: Zschache, Johannes, et autres
Publié: (2025)
par: Zschache, Johannes, et autres
Publié: (2025)
Peeking Behind Closed Doors: Risks of LLM Evaluation by Private Data Curators
par: Bansal, Hritik, et autres
Publié: (2025)
par: Bansal, Hritik, et autres
Publié: (2025)
Collective sleep and activity patterns of college students from wearable devices
par: Fudolig, Mikaela Irene, et autres
Publié: (2024)
par: Fudolig, Mikaela Irene, et autres
Publié: (2024)
How Large Language Models play humans in online conversations: a simulated study of the 2016 US politics on Reddit
par: Cirulli, Daniele, et autres
Publié: (2025)
par: Cirulli, Daniele, et autres
Publié: (2025)
Mapping Toxic Comments Across Demographics: A Dataset from German Public Broadcasting
par: Fillies, Jan, et autres
Publié: (2025)
par: Fillies, Jan, et autres
Publié: (2025)
Using generative AI to support standardization work -- the case of 3GPP
par: Staron, Miroslaw, et autres
Publié: (2024)
par: Staron, Miroslaw, et autres
Publié: (2024)
Semi-automated analysis of audio-recorded lessons: The case of teachers' engaging messages
par: Falcon, Samuel, et autres
Publié: (2024)
par: Falcon, Samuel, et autres
Publié: (2024)
The Life Cycle of Large Language Models: A Review of Biases in Education
par: Lee, Jinsook, et autres
Publié: (2024)
par: Lee, Jinsook, et autres
Publié: (2024)
Autoscoring Anticlimax: A Meta-analytic Understanding of AI's Short-answer Shortcomings and Wording Weaknesses
par: Hardy, Michael
Publié: (2026)
par: Hardy, Michael
Publié: (2026)
Politicians vs ChatGPT. A study of presuppositions in French and Italian political communication
par: Garassino, Davide, et autres
Publié: (2024)
par: Garassino, Davide, et autres
Publié: (2024)
Sentiment analysis and random forest to classify LLM versus human source applied to Scientific Texts
par: Sanchez-Medina, Javier J.
Publié: (2024)
par: Sanchez-Medina, Javier J.
Publié: (2024)
Multilingual corpora for the study of new concepts in the social sciences and humanities:
par: Kyriakoglou, Revekka, et autres
Publié: (2025)
par: Kyriakoglou, Revekka, et autres
Publié: (2025)
Documents similaires
-
LLMs for Low-Resource Dialect Translation Using Context-Aware Prompting: A Case Study on Sylheti
par: Prama, Tabia Tanzin, et autres
Publié: (2025) -
Misalignment of LLM-Generated Personas with Human Perceptions in Low-Resource Settings
par: Prama, Tabia Tanzin, et autres
Publié: (2025) -
Hollywood's misrepresentation of death: A comparison of overall and by-gender mortality causes in film and the real world
par: Beauregard, Calla, et autres
Publié: (2024) -
BanglaMATH : A Bangla benchmark dataset for testing LLM mathematical reasoning at grades 6, 7, and 8
par: Prama, Tabia Tanzin, et autres
Publié: (2025) -
Story and essential meaning dynamics in Bangladesh's July 2024 Student-People's Uprising
par: Prama, Tabia Tanzin, et autres
Publié: (2025)