STaD: Scaffolded Task Design for Identifying Compositional Skill Gaps in LLMs

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: An, Sungeun, Kadhe, Swanand Ravindra, Thakur, Shailja, DeLuca, Chad, Patel, Hima
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908982273114112
author An, Sungeun
Kadhe, Swanand Ravindra
Thakur, Shailja
DeLuca, Chad
Patel, Hima
author_facet An, Sungeun
Kadhe, Swanand Ravindra
Thakur, Shailja
DeLuca, Chad
Patel, Hima
contents Benchmarks are often used as a standard to understand LLM capabilities in different domains. However, aggregate benchmark scores provide limited insight into compositional skill gaps of LLMs and how to improve them. To make these weaknesses visible, we propose Scaffolded Task Design (STaD) framework. STaD generates controlled variations of benchmark tasks based on the concept of scaffolding, which introduces structured, incremental support in a step-by-step manner. Rather than inspecting failures individually, this approach enables systematic and scalable probing of model behavior by identifying the specific reasoning skill compositions they lack. Treating the LLM as a black box, our experiments on six models of varying sizes reveal multiple failure points in three reasoning benchmarks and highlight each model's unique and distinct skill gaps.
format Preprint
id arxiv_https___arxiv_org_abs_2604_18177
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle STaD: Scaffolded Task Design for Identifying Compositional Skill Gaps in LLMs
An, Sungeun
Kadhe, Swanand Ravindra
Thakur, Shailja
DeLuca, Chad
Patel, Hima
Computation and Language
Artificial Intelligence
Benchmarks are often used as a standard to understand LLM capabilities in different domains. However, aggregate benchmark scores provide limited insight into compositional skill gaps of LLMs and how to improve them. To make these weaknesses visible, we propose Scaffolded Task Design (STaD) framework. STaD generates controlled variations of benchmark tasks based on the concept of scaffolding, which introduces structured, incremental support in a step-by-step manner. Rather than inspecting failures individually, this approach enables systematic and scalable probing of model behavior by identifying the specific reasoning skill compositions they lack. Treating the LLM as a black box, our experiments on six models of varying sizes reveal multiple failure points in three reasoning benchmarks and highlight each model's unique and distinct skill gaps.
title STaD: Scaffolded Task Design for Identifying Compositional Skill Gaps in LLMs
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2604.18177