Summarize-Exemplify-Reflect: Data-driven Insight Distillation Empowers LLMs for Few-shot Tabular Classification

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Yuan, Yifei, Li, Jiatong, Zhang, Weijia, Aliannejadi, Mohammad, Kanoulas, Evangelos, Hu, Renjun
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866918132630683648
author Yuan, Yifei
Li, Jiatong
Zhang, Weijia
Aliannejadi, Mohammad
Kanoulas, Evangelos
Hu, Renjun
author_facet Yuan, Yifei
Li, Jiatong
Zhang, Weijia
Aliannejadi, Mohammad
Kanoulas, Evangelos
Hu, Renjun
contents Recent studies show the promise of large language models (LLMs) for few-shot tabular classification but highlight challenges due to the variability in structured data. To address this, we propose distilling data into actionable insights to enable robust and effective classification by LLMs. Drawing inspiration from human learning processes, we introduce InsightTab, an insight distillation framework guided by principles of divide-and-conquer, easy-first, and reflective learning. Our approach integrates rule summarization, strategic exemplification, and insight reflection through deep collaboration between LLMs and data modeling techniques. The obtained insights enable LLMs to better align their general knowledge and capabilities with the particular requirements of specific tabular tasks. We extensively evaluate InsightTab on nine datasets. The results demonstrate consistent improvement over state-of-the-art methods. Ablation studies further validate the principle-guided distillation process, while analyses emphasize InsightTab's effectiveness in leveraging labeled data and managing bias.
format Preprint
id arxiv_https___arxiv_org_abs_2508_21561
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Summarize-Exemplify-Reflect: Data-driven Insight Distillation Empowers LLMs for Few-shot Tabular Classification
Yuan, Yifei
Li, Jiatong
Zhang, Weijia
Aliannejadi, Mohammad
Kanoulas, Evangelos
Hu, Renjun
Machine Learning
Computation and Language
Recent studies show the promise of large language models (LLMs) for few-shot tabular classification but highlight challenges due to the variability in structured data. To address this, we propose distilling data into actionable insights to enable robust and effective classification by LLMs. Drawing inspiration from human learning processes, we introduce InsightTab, an insight distillation framework guided by principles of divide-and-conquer, easy-first, and reflective learning. Our approach integrates rule summarization, strategic exemplification, and insight reflection through deep collaboration between LLMs and data modeling techniques. The obtained insights enable LLMs to better align their general knowledge and capabilities with the particular requirements of specific tabular tasks. We extensively evaluate InsightTab on nine datasets. The results demonstrate consistent improvement over state-of-the-art methods. Ablation studies further validate the principle-guided distillation process, while analyses emphasize InsightTab's effectiveness in leveraging labeled data and managing bias.
title Summarize-Exemplify-Reflect: Data-driven Insight Distillation Empowers LLMs for Few-shot Tabular Classification
topic Machine Learning
Computation and Language
url https://arxiv.org/abs/2508.21561