Bandit Guided Submodular Curriculum for Adaptive Subset Selection

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Chanda, Prateek, Agrawal, Prayas, Sureka, Saral, Polu, Lokesh Reddy, Kshirsagar, Atharv, Ramakrishnan, Ganesh
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866918221839335424
author Chanda, Prateek
Agrawal, Prayas
Sureka, Saral
Polu, Lokesh Reddy
Kshirsagar, Atharv
Ramakrishnan, Ganesh
author_facet Chanda, Prateek
Agrawal, Prayas
Sureka, Saral
Polu, Lokesh Reddy
Kshirsagar, Atharv
Ramakrishnan, Ganesh
contents Traditional curriculum learning proceeds from easy to hard samples, yet defining a reliable notion of difficulty remains elusive. Prior work has used submodular functions to induce difficulty scores in curriculum learning. We reinterpret adaptive subset selection and formulate it as a multi-armed bandit problem, where each arm corresponds to a submodular function guiding sample selection. We introduce ONLINESUBMOD, a novel online greedy policy that optimizes a utility-driven reward and provably achieves no-regret performance under various sampling regimes. Empirically, ONLINESUBMOD outperforms both traditional curriculum learning and bi-level optimization approaches across vision and language datasets, showing superior accuracy-efficiency tradeoffs. More broadly, we show that validationdriven reward metrics offer a principled way to guide the curriculum schedule.
format Preprint
id arxiv_https___arxiv_org_abs_2511_22944
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Bandit Guided Submodular Curriculum for Adaptive Subset Selection
Chanda, Prateek
Agrawal, Prayas
Sureka, Saral
Polu, Lokesh Reddy
Kshirsagar, Atharv
Ramakrishnan, Ganesh
Machine Learning
Artificial Intelligence
Traditional curriculum learning proceeds from easy to hard samples, yet defining a reliable notion of difficulty remains elusive. Prior work has used submodular functions to induce difficulty scores in curriculum learning. We reinterpret adaptive subset selection and formulate it as a multi-armed bandit problem, where each arm corresponds to a submodular function guiding sample selection. We introduce ONLINESUBMOD, a novel online greedy policy that optimizes a utility-driven reward and provably achieves no-regret performance under various sampling regimes. Empirically, ONLINESUBMOD outperforms both traditional curriculum learning and bi-level optimization approaches across vision and language datasets, showing superior accuracy-efficiency tradeoffs. More broadly, we show that validationdriven reward metrics offer a principled way to guide the curriculum schedule.
title Bandit Guided Submodular Curriculum for Adaptive Subset Selection
topic Machine Learning
Artificial Intelligence
url https://arxiv.org/abs/2511.22944