Concept-Reversed Winograd Schema Challenge: Evaluating and Improving Robust Reasoning in Large Language Models via Abstraction
Fuente:
arXiv
Saved in:
| Main Authors: | Han, Kaiqiao, Fang, Tianqing, Wang, Zhaowei, Song, Yangqiu, Steedman, Mark |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
ConstraintChecker: A Plugin for Large Language Models to Reason on Commonsense Knowledge Bases
by: Do, Quyet V., et al.
Published: (2024)
by: Do, Quyet V., et al.
Published: (2024)
AbsPyramid: Benchmarking the Abstraction Ability of Language Models with a Unified Entailment Graph
by: Wang, Zhaowei, et al.
Published: (2023)
by: Wang, Zhaowei, et al.
Published: (2023)
The Cognitive Bandwidth Bottleneck: Shifting Long-Horizon Agent from Planning with Actions to Planning with Schemas
by: Xu, Baixuan, et al.
Published: (2025)
by: Xu, Baixuan, et al.
Published: (2025)
Getting Sick After Seeing a Doctor? Diagnosing and Mitigating Knowledge Conflicts in Event Temporal Reasoning
by: Fang, Tianqing, et al.
Published: (2023)
by: Fang, Tianqing, et al.
Published: (2023)
Acquiring and Modelling Abstract Commonsense Knowledge via Conceptualization
by: He, Mutian, et al.
Published: (2022)
by: He, Mutian, et al.
Published: (2022)
Complex Reasoning over Logical Queries on Commonsense Knowledge Graphs
by: Fang, Tianqing, et al.
Published: (2024)
by: Fang, Tianqing, et al.
Published: (2024)
ConKE: Conceptualization-Augmented Knowledge Editing in Large Language Models for Commonsense Reasoning
by: Zhang, Liyu, et al.
Published: (2024)
by: Zhang, Liyu, et al.
Published: (2024)
Picturing Ambiguity: A Visual Twist on the Winograd Schema Challenge
by: Park, Brendan, et al.
Published: (2024)
by: Park, Brendan, et al.
Published: (2024)
Policy Frameworks for Transparent Chain-of-Thought Reasoning in Large Language Models
by: Chen, Yihang, et al.
Published: (2025)
by: Chen, Yihang, et al.
Published: (2025)
LLMs are Frequency Pattern Learners in Natural Language Inference
by: Cheng, Liang, et al.
Published: (2025)
by: Cheng, Liang, et al.
Published: (2025)
EvoGrad: A Dynamic Take on the Winograd Schema Challenge with Human Adversaries
by: Sun, Jing Han, et al.
Published: (2024)
by: Sun, Jing Han, et al.
Published: (2024)
CodeGraph: Enhancing Graph Reasoning of LLMs with Code
by: Cai, Qiaolong, et al.
Published: (2024)
by: Cai, Qiaolong, et al.
Published: (2024)
CLR-Fact: Evaluating the Complex Logical Reasoning Capability of Large Language Models over Factual Knowledge
by: Zheng, Tianshi, et al.
Published: (2024)
by: Zheng, Tianshi, et al.
Published: (2024)
Thai Winograd Schemas: A Benchmark for Thai Commonsense Reasoning
by: Artkaew, Phakphum
Published: (2024)
by: Artkaew, Phakphum
Published: (2024)
On-the-fly Denoising for Data Augmentation in Natural Language Understanding
by: Fang, Tianqing, et al.
Published: (2022)
by: Fang, Tianqing, et al.
Published: (2022)
S2LPP: Small-to-Large Prompt Prediction across LLMs
by: Cheng, Liang, et al.
Published: (2025)
by: Cheng, Liang, et al.
Published: (2025)
KNOWCOMP POKEMON Team at DialAM-2024: A Two-Stage Pipeline for Detecting Relations in Dialogical Argument Mining
by: Zheng, Zihao, et al.
Published: (2024)
by: Zheng, Zihao, et al.
Published: (2024)
WSC+: Enhancing The Winograd Schema Challenge Using Tree-of-Experts
by: Zahraei, Pardis Sadat, et al.
Published: (2024)
by: Zahraei, Pardis Sadat, et al.
Published: (2024)
CANDLE: Iterative Conceptualization and Instantiation Distillation from Large Language Models for Commonsense Reasoning
by: Wang, Weiqi, et al.
Published: (2024)
by: Wang, Weiqi, et al.
Published: (2024)
A Language-agnostic Model of Child Language Acquisition
by: Mahon, Louis, et al.
Published: (2024)
by: Mahon, Louis, et al.
Published: (2024)
CKBP v2: Better Annotation and Reasoning for Commonsense Knowledge Base Population
by: Fang, Tianqing, et al.
Published: (2023)
by: Fang, Tianqing, et al.
Published: (2023)
Neutralizing Bias in LLM Reasoning using Entailment Graphs
by: Cheng, Liang, et al.
Published: (2025)
by: Cheng, Liang, et al.
Published: (2025)
AbsInstruct: Eliciting Abstraction Ability from LLMs through Explanation Tuning with Plausibility Estimation
by: Wang, Zhaowei, et al.
Published: (2024)
by: Wang, Zhaowei, et al.
Published: (2024)
Take a Step Back: Evoking Reasoning via Abstraction in Large Language Models
by: Zheng, Huaixiu Steven, et al.
Published: (2023)
by: Zheng, Huaixiu Steven, et al.
Published: (2023)
DivScene: Towards Open-Vocabulary Object Navigation with Large Vision Language Models in Diverse Scenes
by: Wang, Zhaowei, et al.
Published: (2024)
by: Wang, Zhaowei, et al.
Published: (2024)
MARS: Benchmarking the Metaphysical Reasoning Abilities of Language Models with a Multi-task Evaluation Dataset
by: Wang, Weiqi, et al.
Published: (2024)
by: Wang, Weiqi, et al.
Published: (2024)
SciResearcher: Scaling Deep Research Agents for Frontier Scientific Reasoning
by: Zheng, Tianshi, et al.
Published: (2026)
by: Zheng, Tianshi, et al.
Published: (2026)
Solving the Challenge Set without Solving the Task: On Winograd Schemas as a Test of Pronominal Coreference Resolution
by: Porada, Ian, et al.
Published: (2024)
by: Porada, Ian, et al.
Published: (2024)
Abstraction-of-Thought Makes Language Models Better Reasoners
by: Hong, Ruixin, et al.
Published: (2024)
by: Hong, Ruixin, et al.
Published: (2024)
ComparisonQA: Evaluating Factuality Robustness of LLMs Through Knowledge Frequency Control and Uncertainty
by: Zong, Qing, et al.
Published: (2024)
by: Zong, Qing, et al.
Published: (2024)
Reasoning Abilities of Large Language Models: In-Depth Analysis on the Abstraction and Reasoning Corpus
by: Lee, Seungpil, et al.
Published: (2024)
by: Lee, Seungpil, et al.
Published: (2024)
Evaluating the Robustness of Analogical Reasoning in Large Language Models
by: Lewis, Martha, et al.
Published: (2024)
by: Lewis, Martha, et al.
Published: (2024)
Explicit Inductive Inference using Large Language Models
by: Liu, Tianyang, et al.
Published: (2024)
by: Liu, Tianyang, et al.
Published: (2024)
Understanding Inter-Session Intentions via Complex Logical Reasoning
by: Bai, Jiaxin, et al.
Published: (2023)
by: Bai, Jiaxin, et al.
Published: (2023)
Large Language Model Reasoning Failures
by: Song, Peiyang, et al.
Published: (2026)
by: Song, Peiyang, et al.
Published: (2026)
Meta-Judging with Large Language Models: Concepts, Methods, and Challenges
by: Silva, Hugo, et al.
Published: (2026)
by: Silva, Hugo, et al.
Published: (2026)
Mathematical Reasoning in Large Language Models: Benchmarks, Architectures, Evaluation, and Open Challenges
by: Amjad, Husnain, et al.
Published: (2026)
by: Amjad, Husnain, et al.
Published: (2026)
Reversal of Thought: Enhancing Large Language Models with Preference-Guided Reverse Reasoning Warm-up
by: Yuan, Jiahao, et al.
Published: (2024)
by: Yuan, Jiahao, et al.
Published: (2024)
Evaluating and Improving Robustness in Large Language Models: A Survey and Future Directions
by: Zhang, Kun, et al.
Published: (2025)
by: Zhang, Kun, et al.
Published: (2025)
Constrained Reasoning Chains for Enhancing Theory-of-Mind in Large Language Models
by: Lin, Zizheng, et al.
Published: (2024)
by: Lin, Zizheng, et al.
Published: (2024)
Similar Items
-
ConstraintChecker: A Plugin for Large Language Models to Reason on Commonsense Knowledge Bases
by: Do, Quyet V., et al.
Published: (2024) -
AbsPyramid: Benchmarking the Abstraction Ability of Language Models with a Unified Entailment Graph
by: Wang, Zhaowei, et al.
Published: (2023) -
The Cognitive Bandwidth Bottleneck: Shifting Long-Horizon Agent from Planning with Actions to Planning with Schemas
by: Xu, Baixuan, et al.
Published: (2025) -
Getting Sick After Seeing a Doctor? Diagnosing and Mitigating Knowledge Conflicts in Event Temporal Reasoning
by: Fang, Tianqing, et al.
Published: (2023) -
Acquiring and Modelling Abstract Commonsense Knowledge via Conceptualization
by: He, Mutian, et al.
Published: (2022)