Sandpiper: Orchestrated AI-Annotation for Educational Discourse at Scale
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866911568307945472 |
|---|---|
| author | Hedley, Daryl Pietrzak, Doug Dias, Jorge Burden, Ian Ahtisham, Bakhtawar Zhou, Zhuqian Vanacore, Kirk Marland, Josh Slama, Rachel Reich, Justin Koedinger, Kenneth Kizilcec, René |
| author_facet | Hedley, Daryl Pietrzak, Doug Dias, Jorge Burden, Ian Ahtisham, Bakhtawar Zhou, Zhuqian Vanacore, Kirk Marland, Josh Slama, Rachel Reich, Justin Koedinger, Kenneth Kizilcec, René |
| contents | Digital educational environments are expanding toward complex AI and human discourse, providing researchers with an abundance of data that offers deep insights into learning and instructional processes. However, traditional qualitative analysis remains a labor-intensive bottleneck, severely limiting the scale at which this research can be conducted. We present Sandpiper, a mixed-initiative system designed to serve as a bridge between high-volume conversational data and human qualitative expertise. By tightly coupling interactive researcher dashboards with agentic Large Language Model (LLM) engines, the platform enables scalable analysis without sacrificing methodological rigor. Sandpiper addresses critical barriers to AI adoption in education by implementing context-aware, automated de-identification workflows supported by secure, university-housed infrastructure to ensure data privacy. Furthermore, the system employs schema-constrained orchestration to eliminate LLM hallucinations and enforces strict adherence to qualitative codebooks. An integrated evaluations engine allows for the continuous benchmarking of AI performance against human labels, fostering an iterative approach to model refinement and validation. We propose a user study to evaluate the system's efficacy in improving research efficiency, inter-rater reliability, and researcher trust in AI-assisted qualitative workflows. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2603_08406 |
| institution | arXiv |
| publishDate | 2026 |
| record_format | arxiv |
| spellingShingle | Sandpiper: Orchestrated AI-Annotation for Educational Discourse at Scale Hedley, Daryl Pietrzak, Doug Dias, Jorge Burden, Ian Ahtisham, Bakhtawar Zhou, Zhuqian Vanacore, Kirk Marland, Josh Slama, Rachel Reich, Justin Koedinger, Kenneth Kizilcec, René Human-Computer Interaction Computation and Language Digital educational environments are expanding toward complex AI and human discourse, providing researchers with an abundance of data that offers deep insights into learning and instructional processes. However, traditional qualitative analysis remains a labor-intensive bottleneck, severely limiting the scale at which this research can be conducted. We present Sandpiper, a mixed-initiative system designed to serve as a bridge between high-volume conversational data and human qualitative expertise. By tightly coupling interactive researcher dashboards with agentic Large Language Model (LLM) engines, the platform enables scalable analysis without sacrificing methodological rigor. Sandpiper addresses critical barriers to AI adoption in education by implementing context-aware, automated de-identification workflows supported by secure, university-housed infrastructure to ensure data privacy. Furthermore, the system employs schema-constrained orchestration to eliminate LLM hallucinations and enforces strict adherence to qualitative codebooks. An integrated evaluations engine allows for the continuous benchmarking of AI performance against human labels, fostering an iterative approach to model refinement and validation. We propose a user study to evaluate the system's efficacy in improving research efficiency, inter-rater reliability, and researcher trust in AI-assisted qualitative workflows. |
| title | Sandpiper: Orchestrated AI-Annotation for Educational Discourse at Scale |
| topic | Human-Computer Interaction Computation and Language |
| url | https://arxiv.org/abs/2603.08406 |