Building Scaffolding Dialogue Data with LLM-Simulated Novices

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Chen, Si, Molnar, Izzy, Hua, Ting, Li, Peiyu, Khiem, Le Huy, Ambrose, G. Alex, Lang, Jim, Metoyer, Ronald, Chawla, Nitesh V.
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908810692526080
author Chen, Si
Molnar, Izzy
Hua, Ting
Li, Peiyu
Khiem, Le Huy
Ambrose, G. Alex
Lang, Jim
Metoyer, Ronald
Chawla, Nitesh V.
author_facet Chen, Si
Molnar, Izzy
Hua, Ting
Li, Peiyu
Khiem, Le Huy
Ambrose, G. Alex
Lang, Jim
Metoyer, Ronald
Chawla, Nitesh V.
contents High-quality, multi-turn instructional dialogues between novices and experts are essential for developing AI systems that support teaching, learning, and decision-making. These dialogues often involve scaffolding -- the process by which an expert supports a novice's thinking through questions, feedback, and step-by-step guidance. However, such data are scarce due to privacy concerns in recording and the vulnerability inherent in help-seeking. We present SimInstruct, a scalable, expert-in-the-loop tool for collecting scaffolding dialogues. Using teaching development coaching as an example domain, SimInstruct simulates novice instructors via LLMs, varying their teaching challenges and LLM's persona traits, while human experts provide multi-turn feedback, reasoning, and instructional support. This design enables the creation of realistic, pedagogically rich dialogues without requiring real novice participants. Our results reveal that persona traits, such as extroversion and introversion, meaningfully influence how experts engage. Compared to real mentoring recordings, SimInstruct dialogues demonstrate comparable pedagogical relevance and cognitive depth. Experts also reported the process as engaging and reflective, improving both data quality and their own professional insight. We further fine-tuned a LLaMA model to be an expert model using the augmented dataset, which outperformed GPT-4o in instructional quality. Our analysis highlights GPT-4o's limitations in weak reflective questioning, overuse of generic praise, a condescending tone, and a tendency to overwhelm novices with excessive suggestions.
format Preprint
id arxiv_https___arxiv_org_abs_2508_04428
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Building Scaffolding Dialogue Data with LLM-Simulated Novices
Chen, Si
Molnar, Izzy
Hua, Ting
Li, Peiyu
Khiem, Le Huy
Ambrose, G. Alex
Lang, Jim
Metoyer, Ronald
Chawla, Nitesh V.
Artificial Intelligence
High-quality, multi-turn instructional dialogues between novices and experts are essential for developing AI systems that support teaching, learning, and decision-making. These dialogues often involve scaffolding -- the process by which an expert supports a novice's thinking through questions, feedback, and step-by-step guidance. However, such data are scarce due to privacy concerns in recording and the vulnerability inherent in help-seeking. We present SimInstruct, a scalable, expert-in-the-loop tool for collecting scaffolding dialogues. Using teaching development coaching as an example domain, SimInstruct simulates novice instructors via LLMs, varying their teaching challenges and LLM's persona traits, while human experts provide multi-turn feedback, reasoning, and instructional support. This design enables the creation of realistic, pedagogically rich dialogues without requiring real novice participants. Our results reveal that persona traits, such as extroversion and introversion, meaningfully influence how experts engage. Compared to real mentoring recordings, SimInstruct dialogues demonstrate comparable pedagogical relevance and cognitive depth. Experts also reported the process as engaging and reflective, improving both data quality and their own professional insight. We further fine-tuned a LLaMA model to be an expert model using the augmented dataset, which outperformed GPT-4o in instructional quality. Our analysis highlights GPT-4o's limitations in weak reflective questioning, overuse of generic praise, a condescending tone, and a tendency to overwhelm novices with excessive suggestions.
title Building Scaffolding Dialogue Data with LLM-Simulated Novices
topic Artificial Intelligence
url https://arxiv.org/abs/2508.04428