RuozhiBench: Evaluating LLMs with Logical Fallacies and Misleading Premises
Fuente:
arXiv
Saved in:
| Main Authors: | Zhai, Zenan, Li, Hao, Han, Xudong, Zhang, Zhenxuan, Zhang, Yixuan, Baldwin, Timothy, Li, Haonan |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
A Chinese Dataset for Evaluating the Safeguards in Large Language Models
by: Wang, Yuxia, et al.
Published: (2024)
by: Wang, Yuxia, et al.
Published: (2024)
Learning From Failure: Integrating Negative Examples when Fine-tuning Large Language Models as Agents
by: Wang, Renxi, et al.
Published: (2024)
by: Wang, Renxi, et al.
Published: (2024)
Loki: An Open-Source Tool for Fact Verification
by: Li, Haonan, et al.
Published: (2024)
by: Li, Haonan, et al.
Published: (2024)
Against The Achilles' Heel: A Survey on Red Teaming for Generative Models
by: Lin, Lizhi, et al.
Published: (2024)
by: Lin, Lizhi, et al.
Published: (2024)
How susceptible are LLMs to Logical Fallacies?
by: Payandeh, Amirreza, et al.
Published: (2023)
by: Payandeh, Amirreza, et al.
Published: (2023)
From Leaky Thoughts to Private Reasoning: Controlling What LRMs Say to Themselves
by: Puerto, Haritz, et al.
Published: (2026)
by: Puerto, Haritz, et al.
Published: (2026)
Boosting Logical Fallacy Reasoning in LLMs via Logical Structure Tree
by: Lei, Yuanyuan, et al.
Published: (2024)
by: Lei, Yuanyuan, et al.
Published: (2024)
Detecting Fallacies in Climate Misinformation: A Technocognitive Approach to Identifying Misleading Argumentation
by: Zanartu, Francisco, et al.
Published: (2024)
by: Zanartu, Francisco, et al.
Published: (2024)
Location Aware Modular Biencoder for Tourism Question Answering
by: Li, Haonan, et al.
Published: (2024)
by: Li, Haonan, et al.
Published: (2024)
ToolGen: Unified Tool Retrieval and Calling via Generation
by: Wang, Renxi, et al.
Published: (2024)
by: Wang, Renxi, et al.
Published: (2024)
Reason from Fallacy: Enhancing Large Language Models' Logical Reasoning through Logical Fallacy Understanding
by: Li, Yanda, et al.
Published: (2024)
by: Li, Yanda, et al.
Published: (2024)
Don't Take the Premise for Granted: Evaluating the Premise Critique Ability of Large Language Models
by: Li, Jinzhe, et al.
Published: (2025)
by: Li, Jinzhe, et al.
Published: (2025)
Demystifying Instruction Mixing for Fine-tuning Large Language Models
by: Wang, Renxi, et al.
Published: (2023)
by: Wang, Renxi, et al.
Published: (2023)
Are LLMs Good Zero-Shot Fallacy Classifiers?
by: Pan, Fengjun, et al.
Published: (2024)
by: Pan, Fengjun, et al.
Published: (2024)
Benchmarking Gender and Political Bias in Large Language Models
by: Yang, Jinrui, et al.
Published: (2025)
by: Yang, Jinrui, et al.
Published: (2025)
SCALAR: Scientific Citation-based Live Assessment of Long-context Academic Reasoning
by: Wang, Renxi, et al.
Published: (2025)
by: Wang, Renxi, et al.
Published: (2025)
CMMLU: Measuring massive multitask language understanding in Chinese
by: Li, Haonan, et al.
Published: (2023)
by: Li, Haonan, et al.
Published: (2023)
FB-Bench: A Fine-Grained Multi-Task Benchmark for Evaluating LLMs' Responsiveness to Human Feedback
by: Li, Youquan, et al.
Published: (2024)
by: Li, Youquan, et al.
Published: (2024)
Control Illusion: The Failure of Instruction Hierarchies in Large Language Models
by: Geng, Yilin, et al.
Published: (2025)
by: Geng, Yilin, et al.
Published: (2025)
From Hypothesis to Premises: LLM-based Backward Logical Reasoning with Selective Symbolic Translation
by: Li, Qingchuan, et al.
Published: (2025)
by: Li, Qingchuan, et al.
Published: (2025)
Making Implicit Premises Explicit in Logical Understanding of Enthymemes
by: Feng, Xuyao, et al.
Published: (2026)
by: Feng, Xuyao, et al.
Published: (2026)
PAVE: Premise-Aware Validation and Editing for Retrieval-Augmented LLMs
by: Huang, Tianyi, et al.
Published: (2026)
by: Huang, Tianyi, et al.
Published: (2026)
LLMs Are Prone to Fallacies in Causal Inference
by: Joshi, Nitish, et al.
Published: (2024)
by: Joshi, Nitish, et al.
Published: (2024)
Don't Let It Hallucinate: Premise Verification via Retrieval-Augmented Logical Reasoning
by: Qin, Yuehan, et al.
Published: (2025)
by: Qin, Yuehan, et al.
Published: (2025)
Token-Level Precise Attack on RAG: Searching for the Best Alternatives to Mislead Generation
by: Li, Zizhong, et al.
Published: (2025)
by: Li, Zizhong, et al.
Published: (2025)
A Logical Fallacy-Informed Framework for Argument Generation
by: Mouchel, Luca, et al.
Published: (2024)
by: Mouchel, Luca, et al.
Published: (2024)
When Personalization Misleads: Understanding and Mitigating Hallucinations in Personalized LLMs
by: Sun, Zhongxiang, et al.
Published: (2026)
by: Sun, Zhongxiang, et al.
Published: (2026)
When LLMs Meet Cunning Texts: A Fallacy Understanding Benchmark for Large Language Models
by: Li, Yinghui, et al.
Published: (2024)
by: Li, Yinghui, et al.
Published: (2024)
CoCoLoFa: A Dataset of News Comments with Common Logical Fallacies Written by LLM-Assisted Crowds
by: Yeh, Min-Hsuan, et al.
Published: (2024)
by: Yeh, Min-Hsuan, et al.
Published: (2024)
JailNewsBench: Multi-Lingual and Regional Benchmark for Fake News Generation under Jailbreak Attacks
by: Kaneko, Masahiro, et al.
Published: (2026)
by: Kaneko, Masahiro, et al.
Published: (2026)
Evaluating Evidence Attribution in Generated Fact Checking Explanations
by: Xing, Rui, et al.
Published: (2024)
by: Xing, Rui, et al.
Published: (2024)
E-Bench: Towards Evaluating the Ease-of-Use of Large Language Models
by: Zhang, Zhenyu, et al.
Published: (2024)
by: Zhang, Zhenyu, et al.
Published: (2024)
Flee the Flaw: Annotating the Underlying Logic of Fallacious Arguments Through Templates and Slot-filling
by: Robbani, Irfan, et al.
Published: (2024)
by: Robbani, Irfan, et al.
Published: (2024)
FinResearchBench: A Logic Tree based Agent-as-a-Judge Evaluation Framework for Financial Research Agents
by: Sun, Rui, et al.
Published: (2025)
by: Sun, Rui, et al.
Published: (2025)
AgentBench: Evaluating LLMs as Agents
by: Liu, Xiao, et al.
Published: (2023)
by: Liu, Xiao, et al.
Published: (2023)
MoralBench: Moral Evaluation of LLMs
by: Ji, Jianchao, et al.
Published: (2024)
by: Ji, Jianchao, et al.
Published: (2024)
MolViBench: Evaluating LLMs on Molecular Vibe Coding
by: Li, Jiatong, et al.
Published: (2026)
by: Li, Jiatong, et al.
Published: (2026)
KG-FPQ: Evaluating Factuality Hallucination in LLMs with Knowledge Graph-based False Premise Questions
by: Zhu, Yanxu, et al.
Published: (2024)
by: Zhu, Yanxu, et al.
Published: (2024)
Bits Leaked per Query: Information-Theoretic Bounds on Adversarial Attacks against LLMs
by: Kaneko, Masahiro, et al.
Published: (2025)
by: Kaneko, Masahiro, et al.
Published: (2025)
Emotionally Charged, Logically Blurred: AI-driven Emotional Framing Impairs Human Fallacy Detection
by: Chen, Yanran, et al.
Published: (2025)
by: Chen, Yanran, et al.
Published: (2025)
Similar Items
-
A Chinese Dataset for Evaluating the Safeguards in Large Language Models
by: Wang, Yuxia, et al.
Published: (2024) -
Learning From Failure: Integrating Negative Examples when Fine-tuning Large Language Models as Agents
by: Wang, Renxi, et al.
Published: (2024) -
Loki: An Open-Source Tool for Fact Verification
by: Li, Haonan, et al.
Published: (2024) -
Against The Achilles' Heel: A Survey on Red Teaming for Generative Models
by: Lin, Lizhi, et al.
Published: (2024) -
How susceptible are LLMs to Logical Fallacies?
by: Payandeh, Amirreza, et al.
Published: (2023)