Enhancing Transformers for Generalizable First-Order Logical Entailment

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zheng, Tianshi, Wang, Jiazheng, Wang, Zihao, Bai, Jiaxin, Yin, Hang, Deng, Zheye, Song, Yangqiu, Li, Jianxin
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908443801026560
author Zheng, Tianshi
Wang, Jiazheng
Wang, Zihao
Bai, Jiaxin
Yin, Hang
Deng, Zheye
Song, Yangqiu
Li, Jianxin
author_facet Zheng, Tianshi
Wang, Jiazheng
Wang, Zihao
Bai, Jiaxin
Yin, Hang
Deng, Zheye
Song, Yangqiu
Li, Jianxin
contents Transformers, as the fundamental deep learning architecture, have demonstrated great capability in reasoning. This paper studies the generalizable first-order logical reasoning ability of transformers with their parameterized knowledge and how to improve it. Transformers' capability of first-order reasoning is further captured by whether they can conduct first-order logical entailment, which is quantitatively measured by their performance in answering knowledge graph queries. We establish the connections between (1) two types of distribution shifts studied in out-of-distribution generalization and (2) unseen knowledge and query settings discussed in the task of knowledge graph query answering, which makes it possible to characterize the fine-grained generalizability. Results on our comprehensive dataset showed that transformers \textit{outperform} previous methods designed particularly for this task and provided detailed empirical evidence about the impact of the input query syntax, token embedding, and transformer architectures on their reasoning capability. Interestingly, our results revealed the mismatch of positional encoding and other design choices of transformer architectures in previous practices. Motivated by this, we propose TEGA, a logic-aware architecture that significantly improves the performance in generalizable first-order logical entailment.
format Preprint
id arxiv_https___arxiv_org_abs_2501_00759
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Enhancing Transformers for Generalizable First-Order Logical Entailment
Zheng, Tianshi
Wang, Jiazheng
Wang, Zihao
Bai, Jiaxin
Yin, Hang
Deng, Zheye
Song, Yangqiu
Li, Jianxin
Computation and Language
Artificial Intelligence
Transformers, as the fundamental deep learning architecture, have demonstrated great capability in reasoning. This paper studies the generalizable first-order logical reasoning ability of transformers with their parameterized knowledge and how to improve it. Transformers' capability of first-order reasoning is further captured by whether they can conduct first-order logical entailment, which is quantitatively measured by their performance in answering knowledge graph queries. We establish the connections between (1) two types of distribution shifts studied in out-of-distribution generalization and (2) unseen knowledge and query settings discussed in the task of knowledge graph query answering, which makes it possible to characterize the fine-grained generalizability. Results on our comprehensive dataset showed that transformers \textit{outperform} previous methods designed particularly for this task and provided detailed empirical evidence about the impact of the input query syntax, token embedding, and transformer architectures on their reasoning capability. Interestingly, our results revealed the mismatch of positional encoding and other design choices of transformer architectures in previous practices. Motivated by this, we propose TEGA, a logic-aware architecture that significantly improves the performance in generalizable first-order logical entailment.
title Enhancing Transformers for Generalizable First-Order Logical Entailment
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2501.00759