Saved in:
Bibliographic Details
Main Authors: Anantheswaran, Ujjwala, Gupta, Himanshu, Scaria, Kevin, Verma, Shreyas, Baral, Chitta, Mishra, Swaroop
Format: Preprint
Published: 2024
Subjects:
Online Access:https://arxiv.org/abs/2406.15444
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911155788709888
author Anantheswaran, Ujjwala
Gupta, Himanshu
Scaria, Kevin
Verma, Shreyas
Baral, Chitta
Mishra, Swaroop
author_facet Anantheswaran, Ujjwala
Gupta, Himanshu
Scaria, Kevin
Verma, Shreyas
Baral, Chitta
Mishra, Swaroop
contents Large Language Models (LLMs) excel at various tasks, including solving math word problems (MWPs), but struggle with real-world problems containing irrelevant information. To address this, we propose a prompting framework that generates adversarial variants of MWPs by adding irrelevant variables. We introduce a dataset, PROBLEMATHIC, containing both adversarial and non-adversarial MWPs. Our experiments reveal that LLMs are susceptible to distraction by numerical noise, resulting in an average relative performance drop of ~26% on adversarial MWPs. To mitigate this, we fine-tune LLMs (Llama-2, Mistral) on the adversarial samples from our dataset. Fine-tuning on adversarial training instances improves performance on adversarial MWPs by ~8%, indicating increased robustness to noise and improved ability to identify relevant data for reasoning. Finally, to assess the generalizability of our prompting framework, we introduce GSM-8K-Adv, an adversarial variant of the GSM-8K benchmark. LLMs continue to struggle when faced with adversarial information, reducing performance by up to 6%.
format Preprint
id arxiv_https___arxiv_org_abs_2406_15444
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Cutting Through the Noise: Boosting LLM Performance on Math Word Problems
Anantheswaran, Ujjwala
Gupta, Himanshu
Scaria, Kevin
Verma, Shreyas
Baral, Chitta
Mishra, Swaroop
Computation and Language
Large Language Models (LLMs) excel at various tasks, including solving math word problems (MWPs), but struggle with real-world problems containing irrelevant information. To address this, we propose a prompting framework that generates adversarial variants of MWPs by adding irrelevant variables. We introduce a dataset, PROBLEMATHIC, containing both adversarial and non-adversarial MWPs. Our experiments reveal that LLMs are susceptible to distraction by numerical noise, resulting in an average relative performance drop of ~26% on adversarial MWPs. To mitigate this, we fine-tune LLMs (Llama-2, Mistral) on the adversarial samples from our dataset. Fine-tuning on adversarial training instances improves performance on adversarial MWPs by ~8%, indicating increased robustness to noise and improved ability to identify relevant data for reasoning. Finally, to assess the generalizability of our prompting framework, we introduce GSM-8K-Adv, an adversarial variant of the GSM-8K benchmark. LLMs continue to struggle when faced with adversarial information, reducing performance by up to 6%.
title Cutting Through the Noise: Boosting LLM Performance on Math Word Problems
topic Computation and Language
url https://arxiv.org/abs/2406.15444