The Effectiveness of LLMs as Annotators: A Comparative Overview and Empirical Analysis of Direct Representation
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Pavlovic, Maja, Poesio, Massimo |
|---|---|
| Format: | Preprint |
| Publié: |
2024
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
An Assessment of Human vs. Model Uncertainty in Soft-Label Learning and Calibration
par: Pavlovic, Maja, et autres
Publié: (2026)
par: Pavlovic, Maja, et autres
Publié: (2026)
Understanding The Effect Of Temperature On Alignment With Human Opinions
par: Pavlovic, Maja, et autres
Publié: (2024)
par: Pavlovic, Maja, et autres
Publié: (2024)
Extending Activation Steering to Broad Skills and Multiple Behaviours
par: van der Weij, Teun, et autres
Publié: (2024)
par: van der Weij, Teun, et autres
Publié: (2024)
Grounded Misunderstandings in Asymmetric Dialogue: A Perspectivist Annotation Scheme for MapTask
par: Li, Nan, et autres
Publié: (2025)
par: Li, Nan, et autres
Publié: (2025)
Understanding Model Calibration -- A gentle introduction and visual exploration of calibration and the expected calibration error (ECE)
par: Pavlovic, Maja
Publié: (2025)
par: Pavlovic, Maja
Publié: (2025)
Spread Preference Annotation: Direct Preference Judgment for Efficient LLM Alignment
par: Kim, Dongyoung, et autres
Publié: (2024)
par: Kim, Dongyoung, et autres
Publié: (2024)
Anomaly Detection of Tabular Data Using LLMs
par: Li, Aodong, et autres
Publié: (2024)
par: Li, Aodong, et autres
Publié: (2024)
An Empirical Study of Data Ability Boundary in LLMs' Math Reasoning
par: Chen, Zui, et autres
Publié: (2024)
par: Chen, Zui, et autres
Publié: (2024)
Cost-Effective Hallucination Detection for LLMs
par: Valentin, Simon, et autres
Publié: (2024)
par: Valentin, Simon, et autres
Publié: (2024)
Semantic Refinement with LLMs for Graph Representations
par: Thapaliya, Safal, et autres
Publié: (2025)
par: Thapaliya, Safal, et autres
Publié: (2025)
Math-Shepherd: Verify and Reinforce LLMs Step-by-step without Human Annotations
par: Wang, Peiyi, et autres
Publié: (2023)
par: Wang, Peiyi, et autres
Publié: (2023)
Relative Bias: A Comparative Framework for Quantifying Bias in LLMs
par: Arbabi, Alireza, et autres
Publié: (2025)
par: Arbabi, Alireza, et autres
Publié: (2025)
Large Language Model Hacking: Quantifying the Hidden Risks of Using LLMs for Text Annotation
par: Baumann, Joachim, et autres
Publié: (2025)
par: Baumann, Joachim, et autres
Publié: (2025)
Referential ambiguity and clarification requests: comparing human and LLM behaviour
par: Madge, Chris, et autres
Publié: (2025)
par: Madge, Chris, et autres
Publié: (2025)
Direct Behavior Optimization: Unlocking the Potential of Lightweight LLMs
par: Yang, Hongming, et autres
Publié: (2025)
par: Yang, Hongming, et autres
Publié: (2025)
Is Child-Directed Speech Effective Training Data for Language Models?
par: Feng, Steven Y., et autres
Publié: (2024)
par: Feng, Steven Y., et autres
Publié: (2024)
Fine-Tuning or Retrieval? Comparing Knowledge Injection in LLMs
par: Ovadia, Oded, et autres
Publié: (2023)
par: Ovadia, Oded, et autres
Publié: (2023)
Picky LLMs and Unreliable RMs: An Empirical Study on Safety Alignment after Instruction Tuning
par: Li, Guanlin, et autres
Publié: (2025)
par: Li, Guanlin, et autres
Publié: (2025)
Agents Thinking Fast and Slow: A Talker-Reasoner Architecture
par: Christakopoulou, Konstantina, et autres
Publié: (2024)
par: Christakopoulou, Konstantina, et autres
Publié: (2024)
A Unified Framework with Novel Metrics for Evaluating the Effectiveness of XAI Techniques in LLMs
par: Mersha, Melkamu Abay, et autres
Publié: (2025)
par: Mersha, Melkamu Abay, et autres
Publié: (2025)
EBFT: Effective and Block-Wise Fine-Tuning for Sparse LLMs
par: Guo, Song, et autres
Publié: (2024)
par: Guo, Song, et autres
Publié: (2024)
MolReasoner: Toward Effective and Interpretable Reasoning for Molecular LLMs
par: Zhao, Guojiang, et autres
Publié: (2025)
par: Zhao, Guojiang, et autres
Publié: (2025)
TopicTag: Automatic Annotation of NMF Topic Models Using Chain of Thought and Prompt Tuning with LLMs
par: Wanna, Selma, et autres
Publié: (2024)
par: Wanna, Selma, et autres
Publié: (2024)
SteeringSafety: A Systematic Safety Evaluation Framework of Representation Steering in LLMs
par: Siu, Vincent, et autres
Publié: (2025)
par: Siu, Vincent, et autres
Publié: (2025)
Does The Way You Plan Matter? An Empirical Study of Planning Representations for LLM Web Agents
par: Zambrano, Alejandra, et autres
Publié: (2026)
par: Zambrano, Alejandra, et autres
Publié: (2026)
Unintended Harms of Value-Aligned LLMs: Psychological and Empirical Insights
par: Choi, Sooyung, et autres
Publié: (2025)
par: Choi, Sooyung, et autres
Publié: (2025)
Shared Lexical Task Representations Explain Behavioral Variability In LLMs
par: Yang, Zhuonan, et autres
Publié: (2026)
par: Yang, Zhuonan, et autres
Publié: (2026)
From Variance to Invariance: Qualitative Content Analysis for Narrative Graph Annotation
par: Huang, Junbo, et autres
Publié: (2026)
par: Huang, Junbo, et autres
Publié: (2026)
Can LLMs Help Uncover Insights about LLMs? A Large-Scale, Evolving Literature Analysis of Frontier LLMs
par: Park, Jungsoo, et autres
Publié: (2025)
par: Park, Jungsoo, et autres
Publié: (2025)
Split, Unlearn, Merge: Leveraging Data Attributes for More Effective Unlearning in LLMs
par: Kadhe, Swanand Ravindra, et autres
Publié: (2024)
par: Kadhe, Swanand Ravindra, et autres
Publié: (2024)
Predicting Training Re-evaluation Curves Enables Effective Data Curriculums for LLMs
par: Bergsma, Shane, et autres
Publié: (2025)
par: Bergsma, Shane, et autres
Publié: (2025)
Direct Alignment of Draft Model for Speculative Decoding with Chat-Fine-Tuned LLMs
par: Goel, Raghavv, et autres
Publié: (2024)
par: Goel, Raghavv, et autres
Publié: (2024)
Direct-Inverse Prompting: Analyzing LLMs' Discriminative Capacity in Self-Improving Generation
par: Ahn, Jihyun Janice, et autres
Publié: (2024)
par: Ahn, Jihyun Janice, et autres
Publié: (2024)
An Overview of Large Language Models for Statisticians
par: Ji, Wenlong, et autres
Publié: (2025)
par: Ji, Wenlong, et autres
Publié: (2025)
From Associations to Activations: Comparing Behavioral and Hidden-State Semantic Geometry in LLMs
par: Schiekiera, Louis, et autres
Publié: (2026)
par: Schiekiera, Louis, et autres
Publié: (2026)
From Isolated Conversations to Hierarchical Schemas: Dynamic Tree Memory Representation for LLMs
par: Rezazadeh, Alireza, et autres
Publié: (2024)
par: Rezazadeh, Alireza, et autres
Publié: (2024)
False Fixed Points: Kantian Feedback, Stable Miscalibration, and Representational Compression in LLMs
par: Okutomi, Akira
Publié: (2025)
par: Okutomi, Akira
Publié: (2025)
RadAnnotate: Large Language Models for Efficient and Reliable Radiology Report Annotation
par: Shetty, Saisha Pradeep, et autres
Publié: (2026)
par: Shetty, Saisha Pradeep, et autres
Publié: (2026)
LLMs as Zero-shot Graph Learners: Alignment of GNN Representations with LLM Token Embeddings
par: Wang, Duo, et autres
Publié: (2024)
par: Wang, Duo, et autres
Publié: (2024)
Arithmetic Control of LLMs for Diverse User Preferences: Directional Preference Alignment with Multi-Objective Rewards
par: Wang, Haoxiang, et autres
Publié: (2024)
par: Wang, Haoxiang, et autres
Publié: (2024)
Documents similaires
-
An Assessment of Human vs. Model Uncertainty in Soft-Label Learning and Calibration
par: Pavlovic, Maja, et autres
Publié: (2026) -
Understanding The Effect Of Temperature On Alignment With Human Opinions
par: Pavlovic, Maja, et autres
Publié: (2024) -
Extending Activation Steering to Broad Skills and Multiple Behaviours
par: van der Weij, Teun, et autres
Publié: (2024) -
Grounded Misunderstandings in Asymmetric Dialogue: A Perspectivist Annotation Scheme for MapTask
par: Li, Nan, et autres
Publié: (2025) -
Understanding Model Calibration -- A gentle introduction and visual exploration of calibration and the expected calibration error (ECE)
par: Pavlovic, Maja
Publié: (2025)