Leveraging Machine Learning to Detect Data Curation Activities
Fuente:
arXiv
Saved in:
| Main Authors: | Lafia, Sara, Thomer, Andrea, Bleckley, David, Akmon, Dharma, Hemphill, Libby |
|---|---|
| Format: | Preprint |
| Published: |
2021
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Evaluating how LLM annotations represent diverse views on contentious topics
by: Brown, Megan A., et al.
Published: (2025)
by: Brown, Megan A., et al.
Published: (2025)
The Effects of Moral Framing on Online Fundraising Outcomes: Evidence from GoFundMe Campaigns
by: Kim, Ji Eun, et al.
Published: (2025)
by: Kim, Ji Eun, et al.
Published: (2025)
War and Peace (WarAgent): Large Language Model-based Multi-Agent Simulation of World Wars
by: Hua, Wenyue, et al.
Published: (2023)
by: Hua, Wenyue, et al.
Published: (2023)
Detecting AI-Generated Text in Educational Content: Leveraging Machine Learning and Explainable AI for Academic Integrity
by: Najjar, Ayat A., et al.
Published: (2025)
by: Najjar, Ayat A., et al.
Published: (2025)
Crowdsourced reviews reveal substantial disparities in public perceptions of parking
by: Li, Lingyao, et al.
Published: (2024)
by: Li, Lingyao, et al.
Published: (2024)
AppealMod: Inducing Friction to Reduce Moderator Workload of Handling User Appeals
by: Atreja, Shubham, et al.
Published: (2023)
by: Atreja, Shubham, et al.
Published: (2023)
Curating corpora with classifiers: A case study of clean energy sentiment online
by: Arnold, Michael V., et al.
Published: (2023)
by: Arnold, Michael V., et al.
Published: (2023)
The Spread of Virtual Gifting in Live Streaming: The Case of Twitch
by: Kim, Ji Eun, et al.
Published: (2025)
by: Kim, Ji Eun, et al.
Published: (2025)
From Measurement Instruments to Data: Leveraging Theory-Driven Synthetic Training Data for Classifying Social Constructs
by: Birkenmaier, Lukas, et al.
Published: (2024)
by: Birkenmaier, Lukas, et al.
Published: (2024)
Leveraging Large Language Models to Measure Gender Representation Bias in Gendered Language Corpora
by: Derner, Erik, et al.
Published: (2024)
by: Derner, Erik, et al.
Published: (2024)
“Unnecessarily cumbersome”: Researchers' Opinions on Restricted Data Access Systems
by: Megan A Brown, et al.
Published: (2025)
by: Megan A Brown, et al.
Published: (2025)
SoK: Machine Learning for Misinformation Detection
by: Xiao, Madelyne, et al.
Published: (2023)
by: Xiao, Madelyne, et al.
Published: (2023)
Machine Learning for Detection and Analysis of Novel LLM Jailbreaks
by: Hawkins, John, et al.
Published: (2025)
by: Hawkins, John, et al.
Published: (2025)
OpenTuringBench: An Open-Model-based Benchmark and Framework for Machine-Generated Text Detection and Attribution
by: La Cava, Lucio, et al.
Published: (2025)
by: La Cava, Lucio, et al.
Published: (2025)
Peeking Behind Closed Doors: Risks of LLM Evaluation by Private Data Curators
by: Bansal, Hritik, et al.
Published: (2025)
by: Bansal, Hritik, et al.
Published: (2025)
Characterizing Online Toxicity During the 2022 Mpox Outbreak: A Computational Analysis of Topical and Network Dynamics
by: Fan, Lizhou, et al.
Published: (2024)
by: Fan, Lizhou, et al.
Published: (2024)
LM$^2$otifs : An Explainable Framework for Machine-Generated Texts Detection
by: Zheng, Xu, et al.
Published: (2025)
by: Zheng, Xu, et al.
Published: (2025)
Luminol-AIDetect: Fast Zero-shot Machine-Generated Text Detection based on Perplexity under Text Shuffling
by: La Cava, Lucio, et al.
Published: (2026)
by: La Cava, Lucio, et al.
Published: (2026)
ChatEd: A Chatbot Leveraging ChatGPT for an Enhanced Learning Experience in Higher Education
by: Wang, Kevin, et al.
Published: (2023)
by: Wang, Kevin, et al.
Published: (2023)
Landscape of Generative AI in Global News: Topics, Sentiments, and Spatiotemporal Analysis
by: Xian, Lu, et al.
Published: (2024)
by: Xian, Lu, et al.
Published: (2024)
Leveraging Machine Learning to Identify Gendered Stereotypes and Body Image Concerns on Diet and Fitness Online Forums
by: Chu, Minh Duc, et al.
Published: (2024)
by: Chu, Minh Duc, et al.
Published: (2024)
Machine Learning Data Practices through a Data Curation Lens: An Evaluation Framework
by: Bhardwaj, Eshta, et al.
Published: (2024)
by: Bhardwaj, Eshta, et al.
Published: (2024)
LangLingual: A Personalised, Exercise-oriented English Language Learning Tool Leveraging Large Language Models
by: Gupta, Sammriddh, et al.
Published: (2025)
by: Gupta, Sammriddh, et al.
Published: (2025)
Machines in the Crowd? Measuring the Footprint of Machine-Generated Text on Reddit
by: La Cava, Lucio, et al.
Published: (2025)
by: La Cava, Lucio, et al.
Published: (2025)
Evaluating Machine Translation Datasets for Low-Web Data Languages: A Gendered Lens
by: Nigatu, Hellina Hailu, et al.
Published: (2025)
by: Nigatu, Hellina Hailu, et al.
Published: (2025)
Leveraging Prompts in LLMs to Overcome Imbalances in Complex Educational Text Data
by: McClure, Jeanne, et al.
Published: (2024)
by: McClure, Jeanne, et al.
Published: (2024)
Leveraging Large Language Models for Predictive Analysis of Human Misery
by: Seal, Bishanka, et al.
Published: (2025)
by: Seal, Bishanka, et al.
Published: (2025)
Prompt Design Matters for Computational Social Science Tasks but in Unpredictable Ways
by: Atreja, Shubham, et al.
Published: (2024)
by: Atreja, Shubham, et al.
Published: (2024)
"HOT" ChatGPT: The promise of ChatGPT in detecting and discriminating hateful, offensive, and toxic comments on social media
by: Li, Lingyao, et al.
Published: (2023)
by: Li, Lingyao, et al.
Published: (2023)
Is Contrasting All You Need? Contrastive Learning for the Detection and Attribution of AI-generated Text
by: La Cava, Lucio, et al.
Published: (2024)
by: La Cava, Lucio, et al.
Published: (2024)
The ProLiFIC dataset: Leveraging LLMs to Unveil the Italian Lawmaking Process
by: Contestabile, Matilde, et al.
Published: (2025)
by: Contestabile, Matilde, et al.
Published: (2025)
Leveraging Large Language Models for Actionable Course Evaluation Student Feedback to Lecturers
by: Zhang, Mike, et al.
Published: (2024)
by: Zhang, Mike, et al.
Published: (2024)
Large Language Models Require Curated Context for Reliable Political Fact-Checking -- Even with Reasoning and Web Search
by: DeVerna, Matthew R., et al.
Published: (2025)
by: DeVerna, Matthew R., et al.
Published: (2025)
Mental Health Equity in LLMs: Leveraging Multi-Hop Question Answering to Detect Amplified and Silenced Perspectives
by: Haider, Batool, et al.
Published: (2025)
by: Haider, Batool, et al.
Published: (2025)
AIMSCheck: Leveraging LLMs for AI-Assisted Review of Modern Slavery Statements Across Jurisdictions
by: Bora, Adriana Eufrosina, et al.
Published: (2025)
by: Bora, Adriana Eufrosina, et al.
Published: (2025)
Leveraging LLMs to Assess Tutor Moves in Real-Life Dialogues: A Feasibility Study
by: Thomas, Danielle R., et al.
Published: (2025)
by: Thomas, Danielle R., et al.
Published: (2025)
A Big Data-empowered System for Real-time Detection of Regional Discriminatory Comments on Vietnamese Social Media
by: Huynh, An Nghiep, et al.
Published: (2024)
by: Huynh, An Nghiep, et al.
Published: (2024)
Detecting a Proxy for Potential Comorbid ADHD in People Reporting Anxiety Symptoms from Social Media Data
by: Lee, Claire S., et al.
Published: (2024)
by: Lee, Claire S., et al.
Published: (2024)
People Make Better Edits: Measuring the Efficacy of LLM-Generated Counterfactually Augmented Data for Harmful Language Detection
by: Sen, Indira, et al.
Published: (2023)
by: Sen, Indira, et al.
Published: (2023)
DetectAnyLLM: Towards Generalizable and Robust Detection of Machine-Generated Text Across Domains and Models
by: Fu, Jiachen, et al.
Published: (2025)
by: Fu, Jiachen, et al.
Published: (2025)
Similar Items
-
Evaluating how LLM annotations represent diverse views on contentious topics
by: Brown, Megan A., et al.
Published: (2025) -
The Effects of Moral Framing on Online Fundraising Outcomes: Evidence from GoFundMe Campaigns
by: Kim, Ji Eun, et al.
Published: (2025) -
War and Peace (WarAgent): Large Language Model-based Multi-Agent Simulation of World Wars
by: Hua, Wenyue, et al.
Published: (2023) -
Detecting AI-Generated Text in Educational Content: Leveraging Machine Learning and Explainable AI for Academic Integrity
by: Najjar, Ayat A., et al.
Published: (2025) -
Crowdsourced reviews reveal substantial disparities in public perceptions of parking
by: Li, Lingyao, et al.
Published: (2024)