Curate-Train-Refine: A Closed-Loop Agentic Framework for Zero Shot Classification
Fuente:
arXiv
Saved in:
| Main Authors: | , |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866911393543880704 |
|---|---|
| author | Maheshwari, Gaurav Haddad, Kevin El |
| author_facet | Maheshwari, Gaurav Haddad, Kevin El |
| contents | Large language models (LLMs) and high-capacity encoders have advanced zero and few-shot classification, but their inference cost and latency limit practical deployment. We propose training lightweight text classifiers using dynamically generated supervision from an LLM. Our method employs an iterative, agentic loop in which the LLM curates training data, analyzes model successes and failures, and synthesizes targeted examples to address observed errors. This closed-loop generation and evaluation process progressively improves data quality and adapts it to the downstream classifier and task. Across four widely used benchmarks, our approach consistently outperforms standard zero and few-shot baselines. These results indicate that LLMs can serve effectively as data curators, enabling accurate and efficient classification without the operational cost of large-model deployment. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2601_16530 |
| institution | arXiv |
| publishDate | 2026 |
| record_format | arxiv |
| spellingShingle | Curate-Train-Refine: A Closed-Loop Agentic Framework for Zero Shot Classification Maheshwari, Gaurav Haddad, Kevin El Computation and Language Machine Learning Large language models (LLMs) and high-capacity encoders have advanced zero and few-shot classification, but their inference cost and latency limit practical deployment. We propose training lightweight text classifiers using dynamically generated supervision from an LLM. Our method employs an iterative, agentic loop in which the LLM curates training data, analyzes model successes and failures, and synthesizes targeted examples to address observed errors. This closed-loop generation and evaluation process progressively improves data quality and adapts it to the downstream classifier and task. Across four widely used benchmarks, our approach consistently outperforms standard zero and few-shot baselines. These results indicate that LLMs can serve effectively as data curators, enabling accurate and efficient classification without the operational cost of large-model deployment. |
| title | Curate-Train-Refine: A Closed-Loop Agentic Framework for Zero Shot Classification |
| topic | Computation and Language Machine Learning |
| url | https://arxiv.org/abs/2601.16530 |