STaD: Scaffolded Task Design for Identifying Compositional Skill Gaps in LLMs
Fuente:
arXiv
Saved in:
| Main Authors: | An, Sungeun, Kadhe, Swanand Ravindra, Thakur, Shailja, DeLuca, Chad, Patel, Hima |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Split, Unlearn, Merge: Leveraging Data Attributes for More Effective Unlearning in LLMs
by: Kadhe, Swanand Ravindra, et al.
Published: (2024)
by: Kadhe, Swanand Ravindra, et al.
Published: (2024)
Runtime-Structured Task Decomposition for Agentic Coding Systems
by: Asthana, Shubhi, et al.
Published: (2026)
by: Asthana, Shubhi, et al.
Published: (2026)
SafeMERGE: Preserving Safety Alignment in Fine-Tuned Large Language Models via Selective Layer-Wise Model Merging
by: Djuhera, Aladin, et al.
Published: (2025)
by: Djuhera, Aladin, et al.
Published: (2025)
Fixing It in Post: A Comparative Study of LLM Post-Training Data Quality and Model Performance
by: Djuhera, Aladin, et al.
Published: (2025)
by: Djuhera, Aladin, et al.
Published: (2025)
When Data is the Algorithm: A Systematic Study and Curation of Preference Optimization Datasets
by: Djuhera, Aladin, et al.
Published: (2025)
by: Djuhera, Aladin, et al.
Published: (2025)
TSR: Trajectory-Search Rollouts for Multi-Turn RL of LLM Agents
by: Djuhera, Aladin, et al.
Published: (2026)
by: Djuhera, Aladin, et al.
Published: (2026)
STRIDE: A Systematic Framework for Selecting AI Modalities -- Agentic AI, AI Assistants, or LLM Calls
by: Asthana, Shubhi, et al.
Published: (2025)
by: Asthana, Shubhi, et al.
Published: (2025)
A Systematic Approach for Large Language Models Debugging
by: Shbita, Basel, et al.
Published: (2026)
by: Shbita, Basel, et al.
Published: (2026)
Generating Verifiable Chain of Thoughts from Exection-Traces
by: Thakur, Shailja, et al.
Published: (2025)
by: Thakur, Shailja, et al.
Published: (2025)
LogitScope: A Framework for Analyzing LLM Uncertainty Through Information Metrics
by: Ahmed, Farhan, et al.
Published: (2026)
by: Ahmed, Farhan, et al.
Published: (2026)
Backprompting: Leveraging Synthetic Production Data for Health Advice Guardrails
by: Cheng, Kellen Tan, et al.
Published: (2025)
by: Cheng, Kellen Tan, et al.
Published: (2025)
Evaluating the Dynamics of Membership Privacy in Deep Learning
by: Chen, Yuetian, et al.
Published: (2025)
by: Chen, Yuetian, et al.
Published: (2025)
Kwai-STaR: Transform LLMs into State-Transition Reasoners
by: Lu, Xingyu, et al.
Published: (2024)
by: Lu, Xingyu, et al.
Published: (2024)
Turning Generative Models Degenerate: The Power of Data Poisoning Attacks
by: Jiang, Shuli, et al.
Published: (2024)
by: Jiang, Shuli, et al.
Published: (2024)
Protecting Users From Themselves: Safeguarding Contextual Privacy in Interactions with Conversational Agents
by: Ngong, Ivoline, et al.
Published: (2025)
by: Ngong, Ivoline, et al.
Published: (2025)
SafeCOMM: A Study on Safety Degradation in Fine-Tuned Telecom Large Language Models
by: Djuhera, Aladin, et al.
Published: (2025)
by: Djuhera, Aladin, et al.
Published: (2025)
MermaidSeqBench: An Evaluation Benchmark for NL-to-Mermaid Sequence Diagram Generation
by: Shbita, Basel, et al.
Published: (2025)
by: Shbita, Basel, et al.
Published: (2025)
LLMON: An LLM-native Markup Language to Leverage Structure and Semantics at the LLM Interface
by: Hind, Michael, et al.
Published: (2026)
by: Hind, Michael, et al.
Published: (2026)
Reading between the Lines: Can LLMs Identify Cross-Cultural Communication Gaps?
by: Saha, Sougata, et al.
Published: (2025)
by: Saha, Sougata, et al.
Published: (2025)
Automatically Improving LLM-based Verilog Generation using EDA Tool Feedback
by: Blocklove, Jason, et al.
Published: (2024)
by: Blocklove, Jason, et al.
Published: (2024)
Identifying Good and Bad Neurons for Task-Level Controllable LLMs
by: Li, Wenjie, et al.
Published: (2026)
by: Li, Wenjie, et al.
Published: (2026)
A Comprehensive Survey of Bias in LLMs: Current Landscape and Future Directions
by: Ranjan, Rajesh, et al.
Published: (2024)
by: Ranjan, Rajesh, et al.
Published: (2024)
STaR-GATE: Teaching Language Models to Ask Clarifying Questions
by: Andukuri, Chinmaya, et al.
Published: (2024)
by: Andukuri, Chinmaya, et al.
Published: (2024)
Towards Compositional Generalization of LLMs via Skill Taxonomy Guided Data Synthesis
by: Wei, Yifan, et al.
Published: (2026)
by: Wei, Yifan, et al.
Published: (2026)
V-STaR: Training Verifiers for Self-Taught Reasoners
by: Hosseini, Arian, et al.
Published: (2024)
by: Hosseini, Arian, et al.
Published: (2024)
OneShield -- the Next Generation of LLM Guardrails
by: DeLuca, Chad, et al.
Published: (2025)
by: DeLuca, Chad, et al.
Published: (2025)
An Automatic Prompt Generation System for Tabular Data Tasks
by: Akella, Ashlesha, et al.
Published: (2024)
by: Akella, Ashlesha, et al.
Published: (2024)
Deploying Privacy Guardrails for LLMs: A Comparative Analysis of Real-World Applications
by: Asthana, Shubhi, et al.
Published: (2025)
by: Asthana, Shubhi, et al.
Published: (2025)
STaRR: Spatial-Temporal Token-Dynamics-Aware Responsive Remasking for Diffusion Language Models
by: Sun, Xinhao, et al.
Published: (2025)
by: Sun, Xinhao, et al.
Published: (2025)
GneissWeb: Preparing High Quality Data for LLMs at Scale
by: Gohari, Hajar Emami, et al.
Published: (2025)
by: Gohari, Hajar Emami, et al.
Published: (2025)
A Comparative Study of Task Adaptation Techniques of Large Language Models for Identifying Sustainable Development Goals
by: Cadeddu, Andrea, et al.
Published: (2025)
by: Cadeddu, Andrea, et al.
Published: (2025)
AdaSTaR: Adaptive Data Sampling for Training Self-Taught Reasoners
by: Koh, Woosung, et al.
Published: (2025)
by: Koh, Woosung, et al.
Published: (2025)
B-STaR: Monitoring and Balancing Exploration and Exploitation in Self-Taught Reasoners
by: Zeng, Weihao, et al.
Published: (2024)
by: Zeng, Weihao, et al.
Published: (2024)
Reasoning Isn't Enough: Examining Truth-Bias and Sycophancy in LLMs
by: Barkett, Emilio, et al.
Published: (2025)
by: Barkett, Emilio, et al.
Published: (2025)
Teaching and Evaluating LLMs to Reason About Polymer Design Related Tasks
by: Mohanty, Dikshya, et al.
Published: (2026)
by: Mohanty, Dikshya, et al.
Published: (2026)
Quiet-STaR: Language Models Can Teach Themselves to Think Before Speaking
by: Zelikman, Eric, et al.
Published: (2024)
by: Zelikman, Eric, et al.
Published: (2024)
Negotiating with LLMS: Prompt Hacks, Skill Gaps, and Reasoning Deficits
by: Schneider, Johannes, et al.
Published: (2023)
by: Schneider, Johannes, et al.
Published: (2023)
A Comprehensive Survey of Retrieval-Augmented Generation (RAG): Evolution, Current Landscape and Future Directions
by: Gupta, Shailja, et al.
Published: (2024)
by: Gupta, Shailja, et al.
Published: (2024)
NC-Bench: An LLM Benchmark for Evaluating Conversational Competence
by: Moore, Robert J., et al.
Published: (2026)
by: Moore, Robert J., et al.
Published: (2026)
Alignment-Constrained Dynamic Pruning for LLMs: Identifying and Preserving Alignment-Critical Circuits
by: Patel, Dev, et al.
Published: (2025)
by: Patel, Dev, et al.
Published: (2025)
Similar Items
-
Split, Unlearn, Merge: Leveraging Data Attributes for More Effective Unlearning in LLMs
by: Kadhe, Swanand Ravindra, et al.
Published: (2024) -
Runtime-Structured Task Decomposition for Agentic Coding Systems
by: Asthana, Shubhi, et al.
Published: (2026) -
SafeMERGE: Preserving Safety Alignment in Fine-Tuned Large Language Models via Selective Layer-Wise Model Merging
by: Djuhera, Aladin, et al.
Published: (2025) -
Fixing It in Post: A Comparative Study of LLM Post-Training Data Quality and Model Performance
by: Djuhera, Aladin, et al.
Published: (2025) -
When Data is the Algorithm: A Systematic Study and Curation of Preference Optimization Datasets
by: Djuhera, Aladin, et al.
Published: (2025)