Towards Compositional Generalization of LLMs via Skill Taxonomy Guided Data Synthesis

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Wei, Yifan, Du, Li, Yu, Xiaoyan, Feng, Yang, Li, Angsheng
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908751091466240
author Wei, Yifan
Du, Li
Yu, Xiaoyan
Feng, Yang
Li, Angsheng
author_facet Wei, Yifan
Du, Li
Yu, Xiaoyan
Feng, Yang
Li, Angsheng
contents Large Language Models (LLMs) and agent-based systems often struggle with compositional generalization due to a data bottleneck in which complex skill combinations follow a long-tailed, power-law distribution, limiting both instruction-following performance and generalization in agent-centric tasks. To address this challenge, we propose STEPS, a Skill Taxonomy guided Entropy-based Post-training data Synthesis framework for generating compositionally challenging data. STEPS explicitly targets compositional generalization by uncovering latent relationships among skills and organizing them into an interpretable, hierarchical skill taxonomy using structural information theory. Building on this taxonomy, we formulate data synthesis as a constrained information maximization problem, selecting skill combinations that maximize marginal structural information within the hierarchy while preserving semantic coherence. Experiments on challenging instruction-following benchmarks show that STEPS outperforms existing data synthesis baselines, while also yielding improved compositional generalization in downstream agent-based evaluations.
format Preprint
id arxiv_https___arxiv_org_abs_2601_03676
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Towards Compositional Generalization of LLMs via Skill Taxonomy Guided Data Synthesis
Wei, Yifan
Du, Li
Yu, Xiaoyan
Feng, Yang
Li, Angsheng
Computation and Language
Artificial Intelligence
Large Language Models (LLMs) and agent-based systems often struggle with compositional generalization due to a data bottleneck in which complex skill combinations follow a long-tailed, power-law distribution, limiting both instruction-following performance and generalization in agent-centric tasks. To address this challenge, we propose STEPS, a Skill Taxonomy guided Entropy-based Post-training data Synthesis framework for generating compositionally challenging data. STEPS explicitly targets compositional generalization by uncovering latent relationships among skills and organizing them into an interpretable, hierarchical skill taxonomy using structural information theory. Building on this taxonomy, we formulate data synthesis as a constrained information maximization problem, selecting skill combinations that maximize marginal structural information within the hierarchy while preserving semantic coherence. Experiments on challenging instruction-following benchmarks show that STEPS outperforms existing data synthesis baselines, while also yielding improved compositional generalization in downstream agent-based evaluations.
title Towards Compositional Generalization of LLMs via Skill Taxonomy Guided Data Synthesis
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2601.03676