Abstract Meaning Representation-Based Logic-Driven Data Augmentation for Logical Reasoning

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Bao, Qiming, Peng, Alex Yuxuan, Deng, Zhenyun, Zhong, Wanjun, Gendron, Gael, Pistotti, Timothy, Tan, Neset, Young, Nathan, Chen, Yang, Zhu, Yonghua, Denny, Paul, Witbrock, Michael, Liu, Jiamou
Natura: Preprint
Pubblicazione: 2023
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866912332870844416
author Bao, Qiming
Peng, Alex Yuxuan
Deng, Zhenyun
Zhong, Wanjun
Gendron, Gael
Pistotti, Timothy
Tan, Neset
Young, Nathan
Chen, Yang
Zhu, Yonghua
Denny, Paul
Witbrock, Michael
Liu, Jiamou
author_facet Bao, Qiming
Peng, Alex Yuxuan
Deng, Zhenyun
Zhong, Wanjun
Gendron, Gael
Pistotti, Timothy
Tan, Neset
Young, Nathan
Chen, Yang
Zhu, Yonghua
Denny, Paul
Witbrock, Michael
Liu, Jiamou
contents Combining large language models with logical reasoning enhances their capacity to address problems in a robust and reliable manner. Nevertheless, the intricate nature of logical reasoning poses challenges when gathering reliable data from the web to build comprehensive training datasets, subsequently affecting performance on downstream tasks. To address this, we introduce a novel logic-driven data augmentation approach, AMR-LDA. AMR-LDA converts the original text into an Abstract Meaning Representation (AMR) graph, a structured semantic representation that encapsulates the logical structure of the sentence, upon which operations are performed to generate logically modified AMR graphs. The modified AMR graphs are subsequently converted back into text to create augmented data. Notably, our methodology is architecture-agnostic and enhances both generative large language models, such as GPT-3.5 and GPT-4, through prompt augmentation, and discriminative large language models through contrastive learning with logic-driven data augmentation. Empirical evidence underscores the efficacy of our proposed method with improvement in performance across seven downstream tasks, such as reading comprehension requiring logical reasoning, textual entailment, and natural language inference. Furthermore, our method leads on the ReClor leaderboard at https://eval.ai/web/challenges/challenge-page/503/leaderboard/1347. The source code and data are publicly available at https://github.com/Strong-AI-Lab/Logical-Equivalence-driven-AMR-Data-Augmentation-for-Representation-Learning.
format Preprint
id arxiv_https___arxiv_org_abs_2305_12599
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle Abstract Meaning Representation-Based Logic-Driven Data Augmentation for Logical Reasoning
Bao, Qiming
Peng, Alex Yuxuan
Deng, Zhenyun
Zhong, Wanjun
Gendron, Gael
Pistotti, Timothy
Tan, Neset
Young, Nathan
Chen, Yang
Zhu, Yonghua
Denny, Paul
Witbrock, Michael
Liu, Jiamou
Computation and Language
Artificial Intelligence
Combining large language models with logical reasoning enhances their capacity to address problems in a robust and reliable manner. Nevertheless, the intricate nature of logical reasoning poses challenges when gathering reliable data from the web to build comprehensive training datasets, subsequently affecting performance on downstream tasks. To address this, we introduce a novel logic-driven data augmentation approach, AMR-LDA. AMR-LDA converts the original text into an Abstract Meaning Representation (AMR) graph, a structured semantic representation that encapsulates the logical structure of the sentence, upon which operations are performed to generate logically modified AMR graphs. The modified AMR graphs are subsequently converted back into text to create augmented data. Notably, our methodology is architecture-agnostic and enhances both generative large language models, such as GPT-3.5 and GPT-4, through prompt augmentation, and discriminative large language models through contrastive learning with logic-driven data augmentation. Empirical evidence underscores the efficacy of our proposed method with improvement in performance across seven downstream tasks, such as reading comprehension requiring logical reasoning, textual entailment, and natural language inference. Furthermore, our method leads on the ReClor leaderboard at https://eval.ai/web/challenges/challenge-page/503/leaderboard/1347. The source code and data are publicly available at https://github.com/Strong-AI-Lab/Logical-Equivalence-driven-AMR-Data-Augmentation-for-Representation-Learning.
title Abstract Meaning Representation-Based Logic-Driven Data Augmentation for Logical Reasoning
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2305.12599