AI Annotation Orchestration: Evaluating LLM verifiers to Improve the Quality of LLM Annotations in Learning Analytics
Fuente:
arXiv
Saved in:
| Main Authors: | Ahtisham, Bakhtawar, Vanacore, Kirk, Lee, Jinsook, Zhou, Zhuqian, Pietrzak, Doug, Kizilcec, Rene F. |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Domain-Adapted Retrieval for In-Context Annotation of Pedagogical Dialogue Acts
by: Lee, Jinsook, et al.
Published: (2026)
by: Lee, Jinsook, et al.
Published: (2026)
Codebook-Injected Dialogue Segmentation for Multi-Utterance Constructs Annotation: LLM-Assisted and Gold-Label-Free Evaluation
by: Lee, Jinsook, et al.
Published: (2026)
by: Lee, Jinsook, et al.
Published: (2026)
Optimizing LLM Annotation of Classroom Discourse through Multi-Agent Orchestration
by: Ahtisham, Bakhtawar, et al.
Published: (2026)
by: Ahtisham, Bakhtawar, et al.
Published: (2026)
LLM Reasoning Predicts When Models Are Right: Evidence from Coding Classroom Discourse
by: Ahtisham, Bakhtawar, et al.
Published: (2026)
by: Ahtisham, Bakhtawar, et al.
Published: (2026)
Sandpiper: Orchestrated AI-Annotation for Educational Discourse at Scale
by: Hedley, Daryl, et al.
Published: (2026)
by: Hedley, Daryl, et al.
Published: (2026)
Utility-Preserving De-Identification for Math Tutoring: Investigating Numeric Ambiguity in the MathEd-PII Benchmark Dataset
by: Zhou, Zhuqian, et al.
Published: (2026)
by: Zhou, Zhuqian, et al.
Published: (2026)
Million Tutoring Moves (MTM): An Open Multimodal Dataset for the Science of Tutoring
by: Kizilcec, René, et al.
Published: (2026)
by: Kizilcec, René, et al.
Published: (2026)
Tutor Move Taxonomy: A Theory-Aligned Framework for Analyzing Instructional Moves in Tutoring
by: Zhou, Zhuqian, et al.
Published: (2026)
by: Zhou, Zhuqian, et al.
Published: (2026)
How well do Large Language Models Recognize Instructional Moves? Establishing Baselines for Foundation Models in Educational Discourse
by: Vanacore, Kirk, et al.
Published: (2025)
by: Vanacore, Kirk, et al.
Published: (2025)
Can External Validation Tools Improve Annotation Quality for LLM-as-a-Judge?
by: Findeis, Arduin, et al.
Published: (2025)
by: Findeis, Arduin, et al.
Published: (2025)
Evaluating the Impact of LLM-Assisted Annotation in a Perspectivized Setting: the Case of FrameNet Annotation
by: Belcavello, Frederico, et al.
Published: (2025)
by: Belcavello, Frederico, et al.
Published: (2025)
Can Machines Learn the True Probabilities?
by: Kim, Jinsook
Published: (2024)
by: Kim, Jinsook
Published: (2024)
Repurposing Annotation Guidelines to Instruct LLM Annotators: A Case Study
by: Kim, Kon Woo, et al.
Published: (2025)
by: Kim, Kon Woo, et al.
Published: (2025)
Prompt Segmentation and Annotation Optimisation: Controlling LLM Behaviour via Optimised Segment-Level Annotations
by: Prasad, Devika, et al.
Published: (2026)
by: Prasad, Devika, et al.
Published: (2026)
AFaCTA: Assisting the Annotation of Factual Claim Detection with Reliable LLM Annotators
by: Ni, Jingwei, et al.
Published: (2024)
by: Ni, Jingwei, et al.
Published: (2024)
Users as Annotators: LLM Preference Learning from Comparison Mode
by: Cai, Zhongze, et al.
Published: (2025)
by: Cai, Zhongze, et al.
Published: (2025)
Emergent Convergence in Multi-Agent LLM Annotation
by: Parfenova, Angelina, et al.
Published: (2025)
by: Parfenova, Angelina, et al.
Published: (2025)
Can Unconfident LLM Annotations Be Used for Confident Conclusions?
by: Gligorić, Kristina, et al.
Published: (2024)
by: Gligorić, Kristina, et al.
Published: (2024)
Enhancing LLM-Based Data Annotation with Error Decomposition
by: Xu, Zhen, et al.
Published: (2026)
by: Xu, Zhen, et al.
Published: (2026)
Enhancing Annotated Bibliography Generation with LLM Ensembles
by: Bermejo, Sergio
Published: (2024)
by: Bermejo, Sergio
Published: (2024)
Spread Preference Annotation: Direct Preference Judgment for Efficient LLM Alignment
by: Kim, Dongyoung, et al.
Published: (2024)
by: Kim, Dongyoung, et al.
Published: (2024)
TAACKIT: Track Annotation and Analytics with Continuous Knowledge Integration Tool
by: Lee, Lily, et al.
Published: (2024)
by: Lee, Lily, et al.
Published: (2024)
SelectLLM: Can LLMs Select Important Instructions to Annotate?
by: Parkar, Ritik Sachin, et al.
Published: (2024)
by: Parkar, Ritik Sachin, et al.
Published: (2024)
Can LLMs Evaluate What They Cannot Annotate? Revisiting LLM Reliability in Hate Speech Detection
by: Piot, Paloma, et al.
Published: (2025)
by: Piot, Paloma, et al.
Published: (2025)
Human and LLM Biases in Hate Speech Annotations: A Socio-Demographic Analysis of Annotators and Targets
by: Giorgi, Tommaso, et al.
Published: (2024)
by: Giorgi, Tommaso, et al.
Published: (2024)
From Human Annotation to Automation: LLM-in-the-Loop Active Learning for Arabic Sentiment Analysis
by: Refai, Dania, et al.
Published: (2025)
by: Refai, Dania, et al.
Published: (2025)
Modernizing Ground Truth: Four Shifts Toward Improving Reliability and Validity in AI in Education
by: Thomas, Danielle R., et al.
Published: (2026)
by: Thomas, Danielle R., et al.
Published: (2026)
Student Engagement in AI Assisted Complex Problem Solving: A Pilot Study of Human AI Rubik's Cube Collaboration
by: Vanacore, Kirk, et al.
Published: (2025)
by: Vanacore, Kirk, et al.
Published: (2025)
Improving Reproducibility in Evaluation through Multi-Level Annotator Modeling
by: Pandita, Deepak, et al.
Published: (2026)
by: Pandita, Deepak, et al.
Published: (2026)
Is LLM an Overconfident Judge? Unveiling the Capabilities of LLMs in Detecting Offensive Language with Annotation Disagreement
by: Lu, Junyu, et al.
Published: (2025)
by: Lu, Junyu, et al.
Published: (2025)
Grid-Orch: An LLM-Powered Orchestrator for Distribution Grid Simulation and Analytics
by: Liu, Boming, et al.
Published: (2026)
by: Liu, Boming, et al.
Published: (2026)
The Alternative Annotator Test for LLM-as-a-Judge: How to Statistically Justify Replacing Human Annotators with LLMs
by: Calderon, Nitay, et al.
Published: (2025)
by: Calderon, Nitay, et al.
Published: (2025)
Quality Assured: Rethinking Annotation Strategies in Imaging AI
by: Rädsch, Tim, et al.
Published: (2024)
by: Rädsch, Tim, et al.
Published: (2024)
Text-to-SQL Domain Adaptation via Human-LLM Collaborative Data Annotation
by: Tian, Yuan, et al.
Published: (2025)
by: Tian, Yuan, et al.
Published: (2025)
CoALFake: Collaborative Active Learning with Human-LLM Co-Annotation for Cross-Domain Fake News Detection
by: Aïmeur, Esma, et al.
Published: (2026)
by: Aïmeur, Esma, et al.
Published: (2026)
LLMind: Orchestrating AI and IoT with LLM for Complex Task Execution
by: Cui, Hongwei, et al.
Published: (2023)
by: Cui, Hongwei, et al.
Published: (2023)
Utility-Focused LLM Annotation for Retrieval and Retrieval-Augmented Generation
by: Zhang, Hengran, et al.
Published: (2025)
by: Zhang, Hengran, et al.
Published: (2025)
ANAH: Analytical Annotation of Hallucinations in Large Language Models
by: Ji, Ziwei, et al.
Published: (2024)
by: Ji, Ziwei, et al.
Published: (2024)
Confidence-Orchestrated Self-Evolution against Uncertain LLM Feedback
by: Wei, Bowen, et al.
Published: (2026)
by: Wei, Bowen, et al.
Published: (2026)
OLAF: Towards Robust LLM-Based Annotation Framework in Empirical Software Engineering
by: Imran, Mia Mohammad, et al.
Published: (2025)
by: Imran, Mia Mohammad, et al.
Published: (2025)
Similar Items
-
Domain-Adapted Retrieval for In-Context Annotation of Pedagogical Dialogue Acts
by: Lee, Jinsook, et al.
Published: (2026) -
Codebook-Injected Dialogue Segmentation for Multi-Utterance Constructs Annotation: LLM-Assisted and Gold-Label-Free Evaluation
by: Lee, Jinsook, et al.
Published: (2026) -
Optimizing LLM Annotation of Classroom Discourse through Multi-Agent Orchestration
by: Ahtisham, Bakhtawar, et al.
Published: (2026) -
LLM Reasoning Predicts When Models Are Right: Evidence from Coding Classroom Discourse
by: Ahtisham, Bakhtawar, et al.
Published: (2026) -
Sandpiper: Orchestrated AI-Annotation for Educational Discourse at Scale
by: Hedley, Daryl, et al.
Published: (2026)