Curate-Train-Refine: A Closed-Loop Agentic Framework for Zero Shot Classification

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Maheshwari, Gaurav, Haddad, Kevin El
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911393543880704
author Maheshwari, Gaurav
Haddad, Kevin El
author_facet Maheshwari, Gaurav
Haddad, Kevin El
contents Large language models (LLMs) and high-capacity encoders have advanced zero and few-shot classification, but their inference cost and latency limit practical deployment. We propose training lightweight text classifiers using dynamically generated supervision from an LLM. Our method employs an iterative, agentic loop in which the LLM curates training data, analyzes model successes and failures, and synthesizes targeted examples to address observed errors. This closed-loop generation and evaluation process progressively improves data quality and adapts it to the downstream classifier and task. Across four widely used benchmarks, our approach consistently outperforms standard zero and few-shot baselines. These results indicate that LLMs can serve effectively as data curators, enabling accurate and efficient classification without the operational cost of large-model deployment.
format Preprint
id arxiv_https___arxiv_org_abs_2601_16530
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Curate-Train-Refine: A Closed-Loop Agentic Framework for Zero Shot Classification
Maheshwari, Gaurav
Haddad, Kevin El
Computation and Language
Machine Learning
Large language models (LLMs) and high-capacity encoders have advanced zero and few-shot classification, but their inference cost and latency limit practical deployment. We propose training lightweight text classifiers using dynamically generated supervision from an LLM. Our method employs an iterative, agentic loop in which the LLM curates training data, analyzes model successes and failures, and synthesizes targeted examples to address observed errors. This closed-loop generation and evaluation process progressively improves data quality and adapts it to the downstream classifier and task. Across four widely used benchmarks, our approach consistently outperforms standard zero and few-shot baselines. These results indicate that LLMs can serve effectively as data curators, enabling accurate and efficient classification without the operational cost of large-model deployment.
title Curate-Train-Refine: A Closed-Loop Agentic Framework for Zero Shot Classification
topic Computation and Language
Machine Learning
url https://arxiv.org/abs/2601.16530