Annotation alignment: Comparing LLM and human annotations of conversational safety
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Movva, Rajiv, Koh, Pang Wei, Pierson, Emma |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
What's In My Human Feedback? Learning Interpretable Descriptions of Preference Data
von: Movva, Rajiv, et al.
Veröffentlicht: (2025)
von: Movva, Rajiv, et al.
Veröffentlicht: (2025)
Sparse Autoencoders for Hypothesis Generation
von: Movva, Rajiv, et al.
Veröffentlicht: (2025)
von: Movva, Rajiv, et al.
Veröffentlicht: (2025)
Use Sparse Autoencoders to Discover Unknown Concepts, Not to Act on Known Concepts
von: Peng, Kenny, et al.
Veröffentlicht: (2025)
von: Peng, Kenny, et al.
Veröffentlicht: (2025)
Topics, Authors, and Institutions in Large Language Model Research: Trends from 17K arXiv Papers
von: Movva, Rajiv, et al.
Veröffentlicht: (2023)
von: Movva, Rajiv, et al.
Veröffentlicht: (2023)
MediQ: Question-Asking LLMs and a Benchmark for Reliable Interactive Clinical Reasoning
von: Li, Shuyue Stella, et al.
Veröffentlicht: (2024)
von: Li, Shuyue Stella, et al.
Veröffentlicht: (2024)
In your own words: computationally identifying interpretable themes in free-text survey data
von: Wang, Jenny S, et al.
Veröffentlicht: (2026)
von: Wang, Jenny S, et al.
Veröffentlicht: (2026)
Generative AI in Medicine
von: Shanmugam, Divya, et al.
Veröffentlicht: (2024)
von: Shanmugam, Divya, et al.
Veröffentlicht: (2024)
Turn-taking annotation for quantitative and qualitative analyses of conversation
von: Kelterer, Anneliese, et al.
Veröffentlicht: (2025)
von: Kelterer, Anneliese, et al.
Veröffentlicht: (2025)
Are aligned neural networks adversarially aligned?
von: Carlini, Nicholas, et al.
Veröffentlicht: (2023)
von: Carlini, Nicholas, et al.
Veröffentlicht: (2023)
LLM_annotate: A Python package for annotating and analyzing fiction characters
von: Rosenbusch, Hannes
Veröffentlicht: (2025)
von: Rosenbusch, Hannes
Veröffentlicht: (2025)
"You are an expert annotator": Automatic Best-Worst-Scaling Annotations for Emotion Intensity Modeling
von: Bagdon, Christopher, et al.
Veröffentlicht: (2024)
von: Bagdon, Christopher, et al.
Veröffentlicht: (2024)
Counting on Consensus: Selecting the Right Inter-annotator Agreement Metric for NLP Annotation and Evaluation
von: James, Joseph
Veröffentlicht: (2026)
von: James, Joseph
Veröffentlicht: (2026)
Comparing human and LLM politeness strategies in free production
von: Zhao, Haoran, et al.
Veröffentlicht: (2025)
von: Zhao, Haoran, et al.
Veröffentlicht: (2025)
Evolution and compression in LLMs: On the emergence of human-aligned categorization
von: Imel, Nathaniel, et al.
Veröffentlicht: (2025)
von: Imel, Nathaniel, et al.
Veröffentlicht: (2025)
Efficient RL for optimizing conversation level outcomes with an LLM-based tutor
von: Nam, Hyunji, et al.
Veröffentlicht: (2025)
von: Nam, Hyunji, et al.
Veröffentlicht: (2025)
Parser agreement and disagreement in L2 Korean UD: Implications for human-in-the-loop annotation
von: Sung, Hakyung, et al.
Veröffentlicht: (2026)
von: Sung, Hakyung, et al.
Veröffentlicht: (2026)
Refining and Reusing Annotation Guidelines for LLM Annotation
von: Kim, Kon Woo, et al.
Veröffentlicht: (2026)
von: Kim, Kon Woo, et al.
Veröffentlicht: (2026)
A Comparative Study on Annotation Quality of Crowdsourcing and LLM via Label Aggregation
von: Li, Jiyi
Veröffentlicht: (2024)
von: Li, Jiyi
Veröffentlicht: (2024)
How desirable is alignment between LLMs and linguistically diverse human users?
von: Knoeferle, Pia, et al.
Veröffentlicht: (2025)
von: Knoeferle, Pia, et al.
Veröffentlicht: (2025)
Inroads to a Structured Data Natural Language Bijection and the role of LLM annotation
von: Vente, Blake
Veröffentlicht: (2024)
von: Vente, Blake
Veröffentlicht: (2024)
AgoraSpeech: A multi-annotated comprehensive dataset of political discourse through the lens of humans and AI
von: Sermpezis, Pavlos, et al.
Veröffentlicht: (2025)
von: Sermpezis, Pavlos, et al.
Veröffentlicht: (2025)
Can External Validation Tools Improve Annotation Quality for LLM-as-a-Judge?
von: Findeis, Arduin, et al.
Veröffentlicht: (2025)
von: Findeis, Arduin, et al.
Veröffentlicht: (2025)
LLM-as-an-Annotator: Training Lightweight Models with LLM-Annotated Examples for Aspect Sentiment Tuple Prediction
von: Hellwig, Nils Constantin, et al.
Veröffentlicht: (2026)
von: Hellwig, Nils Constantin, et al.
Veröffentlicht: (2026)
Alignment-Weighted DPO: A principled reasoning approach to improve safety alignment
von: Hu, Mengxuan, et al.
Veröffentlicht: (2026)
von: Hu, Mengxuan, et al.
Veröffentlicht: (2026)
Identifying and interpreting non-aligned human conceptual representations using language modeling
von: Bao, Wanqian, et al.
Veröffentlicht: (2024)
von: Bao, Wanqian, et al.
Veröffentlicht: (2024)
Comparing human and LLM proofreading in L2 writing: Impact on lexical and syntactic features
von: Sung, Hakyung, et al.
Veröffentlicht: (2025)
von: Sung, Hakyung, et al.
Veröffentlicht: (2025)
Comparing LLM-generated and human-authored news text using formal syntactic theory
von: Zamaraeva, Olga, et al.
Veröffentlicht: (2025)
von: Zamaraeva, Olga, et al.
Veröffentlicht: (2025)
Large-Scale Data Selection for Instruction Tuning
von: Ivison, Hamish, et al.
Veröffentlicht: (2025)
von: Ivison, Hamish, et al.
Veröffentlicht: (2025)
Exploring How Generative MLLMs Perceive More Than CLIP with the Same Vision Encoder
von: Li, Siting, et al.
Veröffentlicht: (2024)
von: Li, Siting, et al.
Veröffentlicht: (2024)
To Err Is Human; To Annotate, SILICON? Toward Robust Reproducibility in LLM Annotation
von: Cheng, Xiang, et al.
Veröffentlicht: (2024)
von: Cheng, Xiang, et al.
Veröffentlicht: (2024)
EvalTree: Profiling Language Model Weaknesses via Hierarchical Capability Trees
von: Zeng, Zhiyuan, et al.
Veröffentlicht: (2025)
von: Zeng, Zhiyuan, et al.
Veröffentlicht: (2025)
Strong and weak alignment of large language models with human values
von: Khamassi, Mehdi, et al.
Veröffentlicht: (2024)
von: Khamassi, Mehdi, et al.
Veröffentlicht: (2024)
Language models align with human judgments on key grammatical constructions
von: Hu, Jennifer, et al.
Veröffentlicht: (2024)
von: Hu, Jennifer, et al.
Veröffentlicht: (2024)
Researchers waste 80% of LLM annotation costs by classifying one text at a time
von: Pipal, Christian, et al.
Veröffentlicht: (2026)
von: Pipal, Christian, et al.
Veröffentlicht: (2026)
Evaluating how LLM annotations represent diverse views on contentious topics
von: Brown, Megan A., et al.
Veröffentlicht: (2025)
von: Brown, Megan A., et al.
Veröffentlicht: (2025)
Hybrid Annotation for Propaganda Detection: Integrating LLM Pre-Annotations with Human Intelligence
von: Sahitaj, Ariana, et al.
Veröffentlicht: (2025)
von: Sahitaj, Ariana, et al.
Veröffentlicht: (2025)
Exposing propaganda: an analysis of stylistic cues comparing human annotations and machine classification
von: Faye, Géraud, et al.
Veröffentlicht: (2024)
von: Faye, Géraud, et al.
Veröffentlicht: (2024)
How Much Annotation is Needed to Compare Summarization Models?
von: Shaib, Chantal, et al.
Veröffentlicht: (2024)
von: Shaib, Chantal, et al.
Veröffentlicht: (2024)
JPEG-LM: LLMs as Image Generators with Canonical Codec Representations
von: Han, Xiaochuang, et al.
Veröffentlicht: (2024)
von: Han, Xiaochuang, et al.
Veröffentlicht: (2024)
Cross-lingual robustness of LLM-brain alignment and its computational roots
von: Yang, Ni, et al.
Veröffentlicht: (2026)
von: Yang, Ni, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
What's In My Human Feedback? Learning Interpretable Descriptions of Preference Data
von: Movva, Rajiv, et al.
Veröffentlicht: (2025) -
Sparse Autoencoders for Hypothesis Generation
von: Movva, Rajiv, et al.
Veröffentlicht: (2025) -
Use Sparse Autoencoders to Discover Unknown Concepts, Not to Act on Known Concepts
von: Peng, Kenny, et al.
Veröffentlicht: (2025) -
Topics, Authors, and Institutions in Large Language Model Research: Trends from 17K arXiv Papers
von: Movva, Rajiv, et al.
Veröffentlicht: (2023) -
MediQ: Question-Asking LLMs and a Benchmark for Reliable Interactive Clinical Reasoning
von: Li, Shuyue Stella, et al.
Veröffentlicht: (2024)