ZebraLogic: On the Scaling Limits of LLMs for Logical Reasoning
Fuente:
arXiv
Saved in:
| Main Authors: | Lin, Bill Yuchen, Bras, Ronan Le, Richardson, Kyle, Sabharwal, Ashish, Poovendran, Radha, Clark, Peter, Choi, Yejin |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Understanding the Logic of Direct Preference Alignment through Logic
by: Richardson, Kyle, et al.
Published: (2024)
by: Richardson, Kyle, et al.
Published: (2024)
Magpie: Alignment Data Synthesis from Scratch by Prompting Aligned LLMs with Nothing
by: Xu, Zhangchen, et al.
Published: (2024)
by: Xu, Zhangchen, et al.
Published: (2024)
MultiZebraLogic: A Multilingual Logical Reasoning Benchmark
by: Bruun, Sofie Helene, et al.
Published: (2025)
by: Bruun, Sofie Helene, et al.
Published: (2025)
WildBench: Benchmarking LLMs with Challenging Tasks from Real Users in the Wild
by: Lin, Bill Yuchen, et al.
Published: (2024)
by: Lin, Bill Yuchen, et al.
Published: (2024)
SimpleToM: Exposing the Gap between Explicit ToM Inference and Implicit ToM Application in LLMs
by: Gu, Yuling, et al.
Published: (2024)
by: Gu, Yuling, et al.
Published: (2024)
Stronger Models are NOT Stronger Teachers for Instruction Tuning
by: Xu, Zhangchen, et al.
Published: (2024)
by: Xu, Zhangchen, et al.
Published: (2024)
TinyV: Reducing False Negatives in Verification Improves RL for LLM Reasoning
by: Xu, Zhangchen, et al.
Published: (2025)
by: Xu, Zhangchen, et al.
Published: (2025)
VisualSphinx: Large-Scale Synthetic Vision Logic Puzzles for RL
by: Feng, Yichen, et al.
Published: (2025)
by: Feng, Yichen, et al.
Published: (2025)
SafeChain: Safety of Language Models with Long Chain-of-Thought Reasoning Capabilities
by: Jiang, Fengqing, et al.
Published: (2025)
by: Jiang, Fengqing, et al.
Published: (2025)
SafeDecoding: Defending against Jailbreak Attacks via Safety-Aware Decoding
by: Xu, Zhangchen, et al.
Published: (2024)
by: Xu, Zhangchen, et al.
Published: (2024)
On the Reasoning Abilities of Masked Diffusion Language Models
by: Svete, Anej, et al.
Published: (2025)
by: Svete, Anej, et al.
Published: (2025)
Can LLMs Reason with Rules? Logic Scaffolding for Stress-Testing and Improving LLMs
by: Wang, Siyuan, et al.
Published: (2024)
by: Wang, Siyuan, et al.
Published: (2024)
ChatBug: A Common Vulnerability of Aligned LLMs Induced by Chat Templates
by: Jiang, Fengqing, et al.
Published: (2024)
by: Jiang, Fengqing, et al.
Published: (2024)
Temporal Sampling for Forgotten Reasoning in LLMs
by: Li, Yuetai, et al.
Published: (2025)
by: Li, Yuetai, et al.
Published: (2025)
SUPER: Evaluating Agents on Setting Up and Executing Tasks from Research Repositories
by: Bogin, Ben, et al.
Published: (2024)
by: Bogin, Ben, et al.
Published: (2024)
From Dogwhistles to Bullhorns: Unveiling Coded Rhetoric with Language Models
by: Mendelsohn, Julia, et al.
Published: (2023)
by: Mendelsohn, Julia, et al.
Published: (2023)
Empowering LLMs with Logical Reasoning: A Comprehensive Survey
by: Cheng, Fengxiang, et al.
Published: (2025)
by: Cheng, Fengxiang, et al.
Published: (2025)
SynLogic: Synthesizing Verifiable Reasoning Data at Scale for Learning Logical Reasoning and Beyond
by: Liu, Junteng, et al.
Published: (2025)
by: Liu, Junteng, et al.
Published: (2025)
WildHallucinations: Evaluating Long-form Factuality in LLMs with Real-World Entity Queries
by: Zhao, Wenting, et al.
Published: (2024)
by: Zhao, Wenting, et al.
Published: (2024)
Logic-Regularized Verifier Elicits Reasoning from LLMs
by: Wang, Xinyu, et al.
Published: (2026)
by: Wang, Xinyu, et al.
Published: (2026)
ArtPrompt: ASCII Art-based Jailbreak Attacks against Aligned LLMs
by: Jiang, Fengqing, et al.
Published: (2024)
by: Jiang, Fengqing, et al.
Published: (2024)
MacGyver: Are Large Language Models Creative Problem Solvers?
by: Tian, Yufei, et al.
Published: (2023)
by: Tian, Yufei, et al.
Published: (2023)
Logical Reasoning with Outcome Reward Models for Test-Time Scaling
by: Thatikonda, Ramya Keerthy, et al.
Published: (2025)
by: Thatikonda, Ramya Keerthy, et al.
Published: (2025)
From Query to Logic: Ontology-Driven Multi-Hop Reasoning in LLMs
by: Bian, Haonan, et al.
Published: (2025)
by: Bian, Haonan, et al.
Published: (2025)
Language Modeling by Language Models
by: Cheng, Junyan, et al.
Published: (2025)
by: Cheng, Junyan, et al.
Published: (2025)
On Memorization of Large Language Models in Logical Reasoning
by: Xie, Chulin, et al.
Published: (2024)
by: Xie, Chulin, et al.
Published: (2024)
Bias Runs Deep: Implicit Reasoning Biases in Persona-Assigned LLMs
by: Gupta, Shashank, et al.
Published: (2023)
by: Gupta, Shashank, et al.
Published: (2023)
Small Models Struggle to Learn from Strong Reasoners
by: Li, Yuetai, et al.
Published: (2025)
by: Li, Yuetai, et al.
Published: (2025)
Does Math Reasoning Improve General LLM Capabilities? Understanding Transferability of LLM Reasoning
by: Huan, Maggie, et al.
Published: (2025)
by: Huan, Maggie, et al.
Published: (2025)
Problem-Solving Logic Guided Curriculum In-Context Learning for LLMs Complex Reasoning
by: Ma, Xuetao, et al.
Published: (2025)
by: Ma, Xuetao, et al.
Published: (2025)
From Implicit to Explicit: Token-Efficient Logical Supervision for Mathematical Reasoning in LLMs
by: Wang, Shaojie, et al.
Published: (2026)
by: Wang, Shaojie, et al.
Published: (2026)
Abstract Meaning Representation-Based Logic-Driven Data Augmentation for Logical Reasoning
by: Bao, Qiming, et al.
Published: (2023)
by: Bao, Qiming, et al.
Published: (2023)
Reason from Fallacy: Enhancing Large Language Models' Logical Reasoning through Logical Fallacy Understanding
by: Li, Yanda, et al.
Published: (2024)
by: Li, Yanda, et al.
Published: (2024)
Do Large Language Models Excel in Complex Logical Reasoning with Formal Language?
by: Jiang, Jin, et al.
Published: (2025)
by: Jiang, Jin, et al.
Published: (2025)
On the Limits of Hierarchically Embedded Logic in Classical Neural Networks
by: Cochran, Bill
Published: (2025)
by: Cochran, Bill
Published: (2025)
Enigmata: Scaling Logical Reasoning in Large Language Models with Synthetic Verifiable Puzzles
by: Chen, Jiangjie, et al.
Published: (2025)
by: Chen, Jiangjie, et al.
Published: (2025)
A Implies B: Circuit Analysis in LLMs for Propositional Logical Reasoning
by: Hong, Guan Zhe, et al.
Published: (2024)
by: Hong, Guan Zhe, et al.
Published: (2024)
LogicBench: Towards Systematic Evaluation of Logical Reasoning Ability of Large Language Models
by: Parmar, Mihir, et al.
Published: (2024)
by: Parmar, Mihir, et al.
Published: (2024)
Lost in the Logic: An Evaluation of Large Language Models' Reasoning Capabilities on LSAT Logic Games
by: Malik, Saumya
Published: (2024)
by: Malik, Saumya
Published: (2024)
Data-driven Discovery with Large Generative Models
by: Majumder, Bodhisattwa Prasad, et al.
Published: (2024)
by: Majumder, Bodhisattwa Prasad, et al.
Published: (2024)
Similar Items
-
Understanding the Logic of Direct Preference Alignment through Logic
by: Richardson, Kyle, et al.
Published: (2024) -
Magpie: Alignment Data Synthesis from Scratch by Prompting Aligned LLMs with Nothing
by: Xu, Zhangchen, et al.
Published: (2024) -
MultiZebraLogic: A Multilingual Logical Reasoning Benchmark
by: Bruun, Sofie Helene, et al.
Published: (2025) -
WildBench: Benchmarking LLMs with Challenging Tasks from Real Users in the Wild
by: Lin, Bill Yuchen, et al.
Published: (2024) -
SimpleToM: Exposing the Gap between Explicit ToM Inference and Implicit ToM Application in LLMs
by: Gu, Yuling, et al.
Published: (2024)