Humans Hallucinate Too: Language Models Identify and Correct Subjective Annotation Errors With Label-in-a-Haystack Prompts
Fuente:
arXiv
Guardado en:
| Autores principales: | Chochlakis, Georgios, Wu, Peter, Bedi, Arjun, Ma, Marcus, Lerman, Kristina, Narayanan, Shrikanth |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Aggregation Artifacts in Subjective Tasks Collapse Large Language Models' Posteriors
por: Chochlakis, Georgios, et al.
Publicado: (2024)
por: Chochlakis, Georgios, et al.
Publicado: (2024)
Larger Language Models Don't Care How You Think: Why Chain-of-Thought Prompting Fails in Subjective Tasks
por: Chochlakis, Georgios, et al.
Publicado: (2024)
por: Chochlakis, Georgios, et al.
Publicado: (2024)
The Strong Pull of Prior Knowledge in Large Language Models and Its Impact on Emotion Recognition
por: Chochlakis, Georgios, et al.
Publicado: (2024)
por: Chochlakis, Georgios, et al.
Publicado: (2024)
Large Language Models Do Multi-Label Classification Differently
por: Ma, Marcus, et al.
Publicado: (2025)
por: Ma, Marcus, et al.
Publicado: (2025)
Authors Should Label Their Own Documents
por: Ma, Marcus, et al.
Publicado: (2025)
por: Ma, Marcus, et al.
Publicado: (2025)
Intelligence Requires Grounding But Not Embodiment
por: Ma, Marcus, et al.
Publicado: (2026)
por: Ma, Marcus, et al.
Publicado: (2026)
Capturing Perspectives of Crowdsourced Annotators in Subjective Learning Tasks
por: Mokhberian, Negar, et al.
Publicado: (2023)
por: Mokhberian, Negar, et al.
Publicado: (2023)
Large Language Models based ASR Error Correction for Child Conversations
por: Xu, Anfeng, et al.
Publicado: (2025)
por: Xu, Anfeng, et al.
Publicado: (2025)
CHATTER: A Character Attribution Dataset for Narrative Understanding
por: Baruah, Sabyasachee, et al.
Publicado: (2024)
por: Baruah, Sabyasachee, et al.
Publicado: (2024)
How to Retrieve Examples in In-context Learning to Improve Conversational Emotion Recognition using Large Language Models?
por: Wang, Mengqi, et al.
Publicado: (2025)
por: Wang, Mengqi, et al.
Publicado: (2025)
Semantic F1 Scores: Fair Evaluation Under Fuzzy Class Boundaries
por: Chochlakis, Georgios, et al.
Publicado: (2025)
por: Chochlakis, Georgios, et al.
Publicado: (2025)
CPL-NoViD: Context-Aware Prompt-based Learning for Norm Violation Detection in Online Communities
por: He, Zihao, et al.
Publicado: (2023)
por: He, Zihao, et al.
Publicado: (2023)
Don't Blame the Data, Blame the Model: Understanding Noise and Bias When Learning from Subjective Annotations
por: Anand, Abhishek, et al.
Publicado: (2024)
por: Anand, Abhishek, et al.
Publicado: (2024)
Needle in the Haystack for Memory Based Large Language Models
por: Nelson, Elliot, et al.
Publicado: (2024)
por: Nelson, Elliot, et al.
Publicado: (2024)
Smoothing Out Hallucinations: Mitigating LLM Hallucination with Smoothed Knowledge Distillation
por: Nguyen, Hieu, et al.
Publicado: (2025)
por: Nguyen, Hieu, et al.
Publicado: (2025)
VariErr NLI: Separating Annotation Error from Human Label Variation
por: Weber-Genzel, Leon, et al.
Publicado: (2024)
por: Weber-Genzel, Leon, et al.
Publicado: (2024)
Uncovering Gaps in How Humans and LLMs Interpret Subjective Language
por: Jones, Erik, et al.
Publicado: (2025)
por: Jones, Erik, et al.
Publicado: (2025)
Prompting Large Language Models with Human Error Markings for Self-Correcting Machine Translation
por: Berger, Nathaniel, et al.
Publicado: (2024)
por: Berger, Nathaniel, et al.
Publicado: (2024)
Multi-Perspective LLM Annotations for Valid Analyses in Subjective Tasks
por: Mehrotra, Navya, et al.
Publicado: (2026)
por: Mehrotra, Navya, et al.
Publicado: (2026)
The Subjectivity of Respect in Police Traffic Stops: Modeling Community Perspectives in Body-Worn Camera Footage
por: Golazizian, Preni, et al.
Publicado: (2026)
por: Golazizian, Preni, et al.
Publicado: (2026)
Will Annotators Disagree? Identifying Subjectivity in Value-Laden Arguments
por: Homayounirad, Amir, et al.
Publicado: (2025)
por: Homayounirad, Amir, et al.
Publicado: (2025)
Large Language Models Reveal Information Operation Goals, Tactics, and Narrative Frames
por: Burghardt, Keith, et al.
Publicado: (2024)
por: Burghardt, Keith, et al.
Publicado: (2024)
Can Memory-Augmented Language Models Generalize on Reasoning-in-a-Haystack Tasks?
por: Das, Payel, et al.
Publicado: (2025)
por: Das, Payel, et al.
Publicado: (2025)
STEER-BENCH: A Benchmark for Evaluating the Steerability of Large Language Models
por: Chen, Kai, et al.
Publicado: (2025)
por: Chen, Kai, et al.
Publicado: (2025)
Evaluating Prompting Strategies for Grammatical Error Correction Based on Language Proficiency
por: Zeng, Min, et al.
Publicado: (2024)
por: Zeng, Min, et al.
Publicado: (2024)
DecoPrompt : Decoding Prompts Reduces Hallucinations when Large Language Models Meet False Premises
por: Xu, Nan, et al.
Publicado: (2024)
por: Xu, Nan, et al.
Publicado: (2024)
Harmful Speech Detection by Language Models Exhibits Gender-Queer Dialect Bias
por: Dorn, Rebecca, et al.
Publicado: (2024)
por: Dorn, Rebecca, et al.
Publicado: (2024)
Large Language Models Struggle to Describe the Haystack without Human Help: Human-in-the-loop Evaluation of Topic Models
por: Li, Zongxia, et al.
Publicado: (2025)
por: Li, Zongxia, et al.
Publicado: (2025)
Jailbreaking in the Haystack
por: Shah, Rishi Rajesh, et al.
Publicado: (2025)
por: Shah, Rishi Rajesh, et al.
Publicado: (2025)
From Haystack to Needle: Label Space Reduction for Zero-shot Classification
por: Vandemoortele, Nathan, et al.
Publicado: (2025)
por: Vandemoortele, Nathan, et al.
Publicado: (2025)
CLEANANERCorp: Identifying and Correcting Incorrect Labels in the ANERcorp Dataset
por: Al-Duwais, Mashael, et al.
Publicado: (2024)
por: Al-Duwais, Mashael, et al.
Publicado: (2024)
ErAConD : Error Annotated Conversational Dialog Dataset for Grammatical Error Correction
por: Yuan, Xun, et al.
Publicado: (2021)
por: Yuan, Xun, et al.
Publicado: (2021)
Whose Emotions and Moral Sentiments Do Language Models Reflect?
por: He, Zihao, et al.
Publicado: (2024)
por: He, Zihao, et al.
Publicado: (2024)
Hallucinations in Neural Automatic Speech Recognition: Identifying Errors and Hallucinatory Models
por: Frieske, Rita, et al.
Publicado: (2024)
por: Frieske, Rita, et al.
Publicado: (2024)
Correct Like Humans: Progressive Learning Framework for Chinese Text Error Correction
por: Li, Yinghui, et al.
Publicado: (2023)
por: Li, Yinghui, et al.
Publicado: (2023)
ANAH: Analytical Annotation of Hallucinations in Large Language Models
por: Ji, Ziwei, et al.
Publicado: (2024)
por: Ji, Ziwei, et al.
Publicado: (2024)
Labels have Human Values: Value Calibration of Subjective Tasks
por: Parappan, Mohammed Fayiz, et al.
Publicado: (2026)
por: Parappan, Mohammed Fayiz, et al.
Publicado: (2026)
TI-ASU: Toward Robust Automatic Speech Understanding through Text-to-speech Imputation Against Missing Speech Modality
por: Feng, Tiantian, et al.
Publicado: (2024)
por: Feng, Tiantian, et al.
Publicado: (2024)
Functional Misalignment in Human-AI Interactions on Digital Platforms
por: Lerman, Kristina
Publicado: (2026)
por: Lerman, Kristina
Publicado: (2026)
COMMUNITY-CROSS-INSTRUCT: Unsupervised Instruction Generation for Aligning Large Language Models to Online Communities
por: He, Zihao, et al.
Publicado: (2024)
por: He, Zihao, et al.
Publicado: (2024)
Ejemplares similares
-
Aggregation Artifacts in Subjective Tasks Collapse Large Language Models' Posteriors
por: Chochlakis, Georgios, et al.
Publicado: (2024) -
Larger Language Models Don't Care How You Think: Why Chain-of-Thought Prompting Fails in Subjective Tasks
por: Chochlakis, Georgios, et al.
Publicado: (2024) -
The Strong Pull of Prior Knowledge in Large Language Models and Its Impact on Emotion Recognition
por: Chochlakis, Georgios, et al.
Publicado: (2024) -
Large Language Models Do Multi-Label Classification Differently
por: Ma, Marcus, et al.
Publicado: (2025) -
Authors Should Label Their Own Documents
por: Ma, Marcus, et al.
Publicado: (2025)