CABINET: Content Relevance based Noise Reduction for Table Question Answering

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Patnaik, Sohan, Changwal, Heril, Aggarwal, Milan, Bhatia, Sumit, Kumar, Yaman, Krishnamurthy, Balaji
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866929241666355200
author Patnaik, Sohan
Changwal, Heril
Aggarwal, Milan
Bhatia, Sumit
Kumar, Yaman
Krishnamurthy, Balaji
author_facet Patnaik, Sohan
Changwal, Heril
Aggarwal, Milan
Bhatia, Sumit
Kumar, Yaman
Krishnamurthy, Balaji
contents Table understanding capability of Large Language Models (LLMs) has been extensively studied through the task of question-answering (QA) over tables. Typically, only a small part of the whole table is relevant to derive the answer for a given question. The irrelevant parts act as noise and are distracting information, resulting in sub-optimal performance due to the vulnerability of LLMs to noise. To mitigate this, we propose CABINET (Content RelevAnce-Based NoIse ReductioN for TablE QuesTion-Answering) - a framework to enable LLMs to focus on relevant tabular data by suppressing extraneous information. CABINET comprises an Unsupervised Relevance Scorer (URS), trained differentially with the QA LLM, that weighs the table content based on its relevance to the input question before feeding it to the question-answering LLM (QA LLM). To further aid the relevance scorer, CABINET employs a weakly supervised module that generates a parsing statement describing the criteria of rows and columns relevant to the question and highlights the content of corresponding table cells. CABINET significantly outperforms various tabular LLM baselines, as well as GPT3-based in-context learning methods, is more robust to noise, maintains outperformance on tables of varying sizes, and establishes new SoTA performance on WikiTQ, FeTaQA, and WikiSQL datasets. We release our code and datasets at https://github.com/Sohanpatnaik106/CABINET_QA.
format Preprint
id arxiv_https___arxiv_org_abs_2402_01155
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle CABINET: Content Relevance based Noise Reduction for Table Question Answering
Patnaik, Sohan
Changwal, Heril
Aggarwal, Milan
Bhatia, Sumit
Kumar, Yaman
Krishnamurthy, Balaji
Computation and Language
Table understanding capability of Large Language Models (LLMs) has been extensively studied through the task of question-answering (QA) over tables. Typically, only a small part of the whole table is relevant to derive the answer for a given question. The irrelevant parts act as noise and are distracting information, resulting in sub-optimal performance due to the vulnerability of LLMs to noise. To mitigate this, we propose CABINET (Content RelevAnce-Based NoIse ReductioN for TablE QuesTion-Answering) - a framework to enable LLMs to focus on relevant tabular data by suppressing extraneous information. CABINET comprises an Unsupervised Relevance Scorer (URS), trained differentially with the QA LLM, that weighs the table content based on its relevance to the input question before feeding it to the question-answering LLM (QA LLM). To further aid the relevance scorer, CABINET employs a weakly supervised module that generates a parsing statement describing the criteria of rows and columns relevant to the question and highlights the content of corresponding table cells. CABINET significantly outperforms various tabular LLM baselines, as well as GPT3-based in-context learning methods, is more robust to noise, maintains outperformance on tables of varying sizes, and establishes new SoTA performance on WikiTQ, FeTaQA, and WikiSQL datasets. We release our code and datasets at https://github.com/Sohanpatnaik106/CABINET_QA.
title CABINET: Content Relevance based Noise Reduction for Table Question Answering
topic Computation and Language
url https://arxiv.org/abs/2402.01155