Saved in:
Bibliographic Details
Main Authors: Kislay, Kaustubh, Singh, Shlok, Joshi, Soham, Dutta, Rohan, Shim, Jay, Flint, George, Zhu, Kevin
Format: Preprint
Published: 2024
Subjects:
Online Access:https://arxiv.org/abs/2410.21896
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866918075546206208
author Kislay, Kaustubh
Singh, Shlok
Joshi, Soham
Dutta, Rohan
Shim, Jay
Flint, George
Zhu, Kevin
author_facet Kislay, Kaustubh
Singh, Shlok
Joshi, Soham
Dutta, Rohan
Shim, Jay
Flint, George
Zhu, Kevin
contents Symbolic Regression remains an NP-Hard problem, with extensive research focusing on AI models for this task. Transformer models have shown promise in Symbolic Regression, but performance suffers with smaller datasets. We propose applying k-fold cross-validation to a transformer-based symbolic regression model trained on a significantly reduced dataset (15,000 data points, down from 500,000). This technique partitions the training data into multiple subsets (folds), iteratively training on some while validating on others. Our aim is to provide an estimate of model generalization and mitigate overfitting issues associated with smaller datasets. Results show that this process improves the model's output consistency and generalization by a relative improvement in validation loss of 53.31%. Potentially enabling more efficient and accessible symbolic regression in resource-constrained environments.
format Preprint
id arxiv_https___arxiv_org_abs_2410_21896
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Evaluating K-Fold Cross Validation for Transformer Based Symbolic Regression Models
Kislay, Kaustubh
Singh, Shlok
Joshi, Soham
Dutta, Rohan
Shim, Jay
Flint, George
Zhu, Kevin
Machine Learning
Computation and Language
Symbolic Regression remains an NP-Hard problem, with extensive research focusing on AI models for this task. Transformer models have shown promise in Symbolic Regression, but performance suffers with smaller datasets. We propose applying k-fold cross-validation to a transformer-based symbolic regression model trained on a significantly reduced dataset (15,000 data points, down from 500,000). This technique partitions the training data into multiple subsets (folds), iteratively training on some while validating on others. Our aim is to provide an estimate of model generalization and mitigate overfitting issues associated with smaller datasets. Results show that this process improves the model's output consistency and generalization by a relative improvement in validation loss of 53.31%. Potentially enabling more efficient and accessible symbolic regression in resource-constrained environments.
title Evaluating K-Fold Cross Validation for Transformer Based Symbolic Regression Models
topic Machine Learning
Computation and Language
url https://arxiv.org/abs/2410.21896