Finding the Muses: Identifying Coresets through Loss Trajectories

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Nagaraj, Manish, Ravikumar, Deepak, Soufleri, Efstathia, Roy, Kaushik
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866929757106470912
author Nagaraj, Manish
Ravikumar, Deepak
Soufleri, Efstathia
Roy, Kaushik
author_facet Nagaraj, Manish
Ravikumar, Deepak
Soufleri, Efstathia
Roy, Kaushik
contents Deep learning models achieve state-of-the-art performance across domains but face scalability challenges in real-time or resource-constrained scenarios. To address this, we propose Loss Trajectory Correlation (LTC), a novel metric for coreset selection that identifies critical training samples driving generalization. $LTC$ quantifies the alignment between training sample loss trajectories and validation set loss trajectories, enabling the construction of compact, representative subsets. Unlike traditional methods with computational and storage overheads that are infeasible to scale to large datasets, $LTC$ achieves superior efficiency as it can be computed as a byproduct of training. Our results on CIFAR-100 and ImageNet-1k show that $LTC$ consistently achieves accuracy on par with or surpassing state-of-the-art coreset selection methods, with any differences remaining under 1%. LTC also effectively transfers across various architectures, including ResNet, VGG, DenseNet, and Swin Transformer, with minimal performance degradation (<2%). Additionally, LTC offers insights into training dynamics, such as identifying aligned and conflicting sample behaviors, at a fraction of the computational cost of traditional methods. This framework paves the way for scalable coreset selection and efficient dataset optimization.
format Preprint
id arxiv_https___arxiv_org_abs_2503_09721
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Finding the Muses: Identifying Coresets through Loss Trajectories
Nagaraj, Manish
Ravikumar, Deepak
Soufleri, Efstathia
Roy, Kaushik
Machine Learning
Artificial Intelligence
Deep learning models achieve state-of-the-art performance across domains but face scalability challenges in real-time or resource-constrained scenarios. To address this, we propose Loss Trajectory Correlation (LTC), a novel metric for coreset selection that identifies critical training samples driving generalization. $LTC$ quantifies the alignment between training sample loss trajectories and validation set loss trajectories, enabling the construction of compact, representative subsets. Unlike traditional methods with computational and storage overheads that are infeasible to scale to large datasets, $LTC$ achieves superior efficiency as it can be computed as a byproduct of training. Our results on CIFAR-100 and ImageNet-1k show that $LTC$ consistently achieves accuracy on par with or surpassing state-of-the-art coreset selection methods, with any differences remaining under 1%. LTC also effectively transfers across various architectures, including ResNet, VGG, DenseNet, and Swin Transformer, with minimal performance degradation (<2%). Additionally, LTC offers insights into training dynamics, such as identifying aligned and conflicting sample behaviors, at a fraction of the computational cost of traditional methods. This framework paves the way for scalable coreset selection and efficient dataset optimization.
title Finding the Muses: Identifying Coresets through Loss Trajectories
topic Machine Learning
Artificial Intelligence
url https://arxiv.org/abs/2503.09721