Can Unconfident LLM Annotations Be Used for Confident Conclusions?
Fuente:
arXiv
Saved in:
| Main Authors: | Gligorić, Kristina, Zrnic, Tijana, Lee, Cinoo, Candès, Emmanuel J., Jurafsky, Dan |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
NLP Systems That Can't Tell Use from Mention Censor Counterspeech, but Teaching the Distinction Helps
by: Gligoric, Kristina, et al.
Published: (2024)
by: Gligoric, Kristina, et al.
Published: (2024)
Humans overrely on overconfident language models, across languages
by: Rathi, Neil, et al.
Published: (2025)
by: Rathi, Neil, et al.
Published: (2025)
Grounding Gaps in Language Model Generations
by: Shaikh, Omar, et al.
Published: (2023)
by: Shaikh, Omar, et al.
Published: (2023)
Rel-A.I.: An Interaction-Centered Approach To Measuring Human-LM Reliance
by: Zhou, Kaitlyn, et al.
Published: (2024)
by: Zhou, Kaitlyn, et al.
Published: (2024)
Human and LLM Biases in Hate Speech Annotations: A Socio-Demographic Analysis of Annotators and Targets
by: Giorgi, Tommaso, et al.
Published: (2024)
by: Giorgi, Tommaso, et al.
Published: (2024)
Othering and low status framing of immigrant cuisines in US restaurant reviews and large language models
by: Luo, Yiwei, et al.
Published: (2023)
by: Luo, Yiwei, et al.
Published: (2023)
The Alternative Annotator Test for LLM-as-a-Judge: How to Statistically Justify Replacing Human Annotators with LLMs
by: Calderon, Nitay, et al.
Published: (2025)
by: Calderon, Nitay, et al.
Published: (2025)
Can Large Language Models be Used to Provide Psychological Counselling? An Analysis of GPT-4-Generated Responses Using Role-play Dialogues
by: Inaba, Michimasa, et al.
Published: (2024)
by: Inaba, Michimasa, et al.
Published: (2024)
LLM Can be a Dangerous Persuader: Empirical Study of Persuasion Safety in Large Language Models
by: Liu, Minqian, et al.
Published: (2025)
by: Liu, Minqian, et al.
Published: (2025)
VeriLA: A Human-Centered Evaluation Framework for Interpretable Verification of LLM Agent Failures
by: Sung, Yoo Yeon, et al.
Published: (2025)
by: Sung, Yoo Yeon, et al.
Published: (2025)
Can LLMs Assist Annotators in Identifying Morality Frames? -- Case Study on Vaccination Debate on Social Media
by: Islam, Tunazzina, et al.
Published: (2025)
by: Islam, Tunazzina, et al.
Published: (2025)
AnthroScore: A Computational Linguistic Measure of Anthropomorphism
by: Cheng, Myra, et al.
Published: (2024)
by: Cheng, Myra, et al.
Published: (2024)
Fakes of Varying Shades: How Warning Affects Human Perception and Engagement Regarding LLM Hallucinations
by: Nahar, Mahjabin, et al.
Published: (2024)
by: Nahar, Mahjabin, et al.
Published: (2024)
A Generalized LLM-Augmented BIM Framework: Application to a Speech-to-BIM system
by: Lee, Ghang, et al.
Published: (2024)
by: Lee, Ghang, et al.
Published: (2024)
Investigating Low-Cost LLM Annotation for~Spoken Dialogue Understanding Datasets
by: Druart, Lucas, et al.
Published: (2024)
by: Druart, Lucas, et al.
Published: (2024)
QACP: An Annotated Question Answering Dataset for Assisting Chinese Python Programming Learners
by: Xiao, Rui, et al.
Published: (2024)
by: Xiao, Rui, et al.
Published: (2024)
Explore, Select, Derive, and Recall: Augmenting LLM with Human-like Memory for Mobile Task Automation
by: Lee, Sunjae, et al.
Published: (2023)
by: Lee, Sunjae, et al.
Published: (2023)
Can Large Language Models Detect Verbal Indicators of Romantic Attraction?
by: Matz, Sandra C., et al.
Published: (2024)
by: Matz, Sandra C., et al.
Published: (2024)
Can Large Language Models generalize analogy solving like children can?
by: Stevenson, Claire E., et al.
Published: (2024)
by: Stevenson, Claire E., et al.
Published: (2024)
Unsupervised Word-level Quality Estimation for Machine Translation Through the Lens of Annotators (Dis)agreement
by: Sarti, Gabriele, et al.
Published: (2025)
by: Sarti, Gabriele, et al.
Published: (2025)
Can LLMs Generate Visualizations with Dataless Prompts?
by: Coelho, Darius, et al.
Published: (2024)
by: Coelho, Darius, et al.
Published: (2024)
Evaluating Behavioral Alignment in Conflict Dialogue: A Multi-Dimensional Comparison of LLM Agents and Humans
by: Kwon, Deuksin, et al.
Published: (2025)
by: Kwon, Deuksin, et al.
Published: (2025)
Creativity in LLM-based Multi-Agent Systems: A Survey
by: Lin, Yi-Cheng, et al.
Published: (2025)
by: Lin, Yi-Cheng, et al.
Published: (2025)
Attention to Non-Adopters
by: Zhou, Kaitlyn, et al.
Published: (2025)
by: Zhou, Kaitlyn, et al.
Published: (2025)
The Generative AI Paradox on Evaluation: What It Can Solve, It May Not Evaluate
by: Oh, Juhyun, et al.
Published: (2024)
by: Oh, Juhyun, et al.
Published: (2024)
Can Large Language Model Agents Simulate Human Trust Behavior?
by: Xie, Chengxing, et al.
Published: (2024)
by: Xie, Chengxing, et al.
Published: (2024)
CiteME: Can Language Models Accurately Cite Scientific Claims?
by: Press, Ori, et al.
Published: (2024)
by: Press, Ori, et al.
Published: (2024)
Impacts of Anthropomorphizing Large Language Models in Learning Environments
by: Schaaff, Kristina, et al.
Published: (2024)
by: Schaaff, Kristina, et al.
Published: (2024)
How Can I Get It Right? Using GPT to Rephrase Incorrect Trainee Responses
by: Lin, Jionghao, et al.
Published: (2024)
by: Lin, Jionghao, et al.
Published: (2024)
Can LLMs Model Incorrect Student Reasoning? A Case Study on Distractor Generation
by: Zengaffinen, Yanick, et al.
Published: (2026)
by: Zengaffinen, Yanick, et al.
Published: (2026)
User Perceptions vs. Proxy LLM Judges: Privacy and Helpfulness in LLM Responses to Privacy-Sensitive Scenarios
by: Wu, Xiaoyuan, et al.
Published: (2025)
by: Wu, Xiaoyuan, et al.
Published: (2025)
Game Development as Human-LLM Interaction
by: Hong, Jiale, et al.
Published: (2024)
by: Hong, Jiale, et al.
Published: (2024)
Benchmarking LLM Tool-Use in the Wild
by: Yu, Peijie, et al.
Published: (2026)
by: Yu, Peijie, et al.
Published: (2026)
Grading Scale Impact on LLM-as-a-Judge: Human-LLM Alignment Is Highest on 0-5 Grading Scale
by: Li, Weiyue, et al.
Published: (2026)
by: Li, Weiyue, et al.
Published: (2026)
Large Language Models Can Solve Real-World Planning Rigorously with Formal Verification Tools
by: Hao, Yilun, et al.
Published: (2024)
by: Hao, Yilun, et al.
Published: (2024)
Interaction Techniques that Encourage Longer Prompts Can Improve Psychological Ownership when Writing with AI
by: Joshi, Nikhita, et al.
Published: (2025)
by: Joshi, Nikhita, et al.
Published: (2025)
Sycophantic AI makes human interaction feel more effortful and less satisfying over time
by: Ibrahim, Lujain, et al.
Published: (2026)
by: Ibrahim, Lujain, et al.
Published: (2026)
How Can I Improve? Using GPT to Highlight the Desired and Undesired Parts of Open-ended Responses
by: Lin, Jionghao, et al.
Published: (2024)
by: Lin, Jionghao, et al.
Published: (2024)
Mediating Modes of Thought: LLM's for design scripting
by: Rietschel, Moritz, et al.
Published: (2024)
by: Rietschel, Moritz, et al.
Published: (2024)
BADGE: BADminton report Generation and Evaluation with LLM
by: Chiang, Shang-Hsuan, et al.
Published: (2024)
by: Chiang, Shang-Hsuan, et al.
Published: (2024)
Similar Items
-
NLP Systems That Can't Tell Use from Mention Censor Counterspeech, but Teaching the Distinction Helps
by: Gligoric, Kristina, et al.
Published: (2024) -
Humans overrely on overconfident language models, across languages
by: Rathi, Neil, et al.
Published: (2025) -
Grounding Gaps in Language Model Generations
by: Shaikh, Omar, et al.
Published: (2023) -
Rel-A.I.: An Interaction-Centered Approach To Measuring Human-LM Reliance
by: Zhou, Kaitlyn, et al.
Published: (2024) -
Human and LLM Biases in Hate Speech Annotations: A Socio-Demographic Analysis of Annotators and Targets
by: Giorgi, Tommaso, et al.
Published: (2024)