RuleArena: A Benchmark for Rule-Guided Reasoning with LLMs in Real-World Scenarios
Fuente:
arXiv
Saved in:
| Main Authors: | Zhou, Ruiwen, Hua, Wenyue, Pan, Liangming, Cheng, Sitao, Wu, Xiaobao, Yu, En, Wang, William Yang |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Epistemic Context Learning: Building Trust the Right Way in LLM-Based Multi-Agent Systems
by: Zhou, Ruiwen, et al.
Published: (2026)
by: Zhou, Ruiwen, et al.
Published: (2026)
AntiLeakBench: Preventing Data Contamination by Automatically Constructing Benchmarks with Updated Real-World Knowledge
by: Wu, Xiaobao, et al.
Published: (2024)
by: Wu, Xiaobao, et al.
Published: (2024)
Atomic Skills are the Prerequisite: When Reinforcement Learning Synthesizes Compositional Reasoning, and When It Only Amplifies
by: Cheng, Sitao, et al.
Published: (2025)
by: Cheng, Sitao, et al.
Published: (2025)
InductionBench: LLMs Fail in the Simplest Complexity Class
by: Hua, Wenyue, et al.
Published: (2025)
by: Hua, Wenyue, et al.
Published: (2025)
Understanding the Interplay between Parametric and Contextual Knowledge for Large Language Models
by: Cheng, Sitao, et al.
Published: (2024)
by: Cheng, Sitao, et al.
Published: (2024)
AKEW: Assessing Knowledge Editing in the Wild
by: Wu, Xiaobao, et al.
Published: (2024)
by: Wu, Xiaobao, et al.
Published: (2024)
THOUGHTTERMINATOR: Benchmarking, Calibrating, and Mitigating Overthinking in Reasoning Models
by: Pu, Xiao, et al.
Published: (2025)
by: Pu, Xiao, et al.
Published: (2025)
MultiAgent Collaboration Attack: Investigating Adversarial Attacks in Large Language Model Collaborations via Debate
by: Amayuelas, Alfonso, et al.
Published: (2024)
by: Amayuelas, Alfonso, et al.
Published: (2024)
Disentangling Memory and Reasoning Ability in Large Language Models
by: Jin, Mingyu, et al.
Published: (2024)
by: Jin, Mingyu, et al.
Published: (2024)
DeonticBench: A Benchmark for Reasoning over Rules
by: Dou, Guangyao, et al.
Published: (2026)
by: Dou, Guangyao, et al.
Published: (2026)
TIME: A Multi-level Benchmark for Temporal Reasoning of LLMs in Real-World Scenarios
by: Wei, Shaohang, et al.
Published: (2025)
by: Wei, Shaohang, et al.
Published: (2025)
Rule-Guided Feedback: Enhancing Reasoning by Enforcing Rule Adherence in Large Language Models
by: Diallo, Aissatou, et al.
Published: (2025)
by: Diallo, Aissatou, et al.
Published: (2025)
ReasonTabQA: A Comprehensive Benchmark for Table Question Answering from Real World Industrial Scenarios
by: Pan, Changzai, et al.
Published: (2026)
by: Pan, Changzai, et al.
Published: (2026)
REALM: A Dataset of Real-World LLM Use Cases
by: Cheng, Jingwen, et al.
Published: (2025)
by: Cheng, Jingwen, et al.
Published: (2025)
Modeling Dynamic Topics in Chain-Free Fashion by Evolution-Tracking Contrastive Learning and Unassociated Word Exclusion
by: Wu, Xiaobao, et al.
Published: (2024)
by: Wu, Xiaobao, et al.
Published: (2024)
How Is LLM Reasoning Distracted by Irrelevant Context? An Analysis Using a Controlled Benchmark
by: Yang, Minglai, et al.
Published: (2025)
by: Yang, Minglai, et al.
Published: (2025)
Business as Rulesual: A Benchmark and Framework for Business Rule Flow Modeling with LLMs
by: Yang, Chen, et al.
Published: (2025)
by: Yang, Chen, et al.
Published: (2025)
ChatRule: Mining Logical Rules with Large Language Models for Knowledge Graph Reasoning
by: Luo, Linhao, et al.
Published: (2023)
by: Luo, Linhao, et al.
Published: (2023)
Can LLMs Reason with Rules? Logic Scaffolding for Stress-Testing and Improving LLMs
by: Wang, Siyuan, et al.
Published: (2024)
by: Wang, Siyuan, et al.
Published: (2024)
RuleReasoner: Reinforced Rule-based Reasoning via Domain-aware Dynamic Sampling
by: Liu, Yang, et al.
Published: (2025)
by: Liu, Yang, et al.
Published: (2025)
DV-World: Benchmarking Data Visualization Agents in Real-World Scenarios
by: Meng, Jinxiang, et al.
Published: (2026)
by: Meng, Jinxiang, et al.
Published: (2026)
Read as You See: Guiding Unimodal LLMs for Low-Resource Explainable Harmful Meme Detection
by: Pan, Fengjun, et al.
Published: (2025)
by: Pan, Fengjun, et al.
Published: (2025)
LogicGame: Benchmarking Rule-Based Reasoning Abilities of Large Language Models
by: Gui, Jiayi, et al.
Published: (2024)
by: Gui, Jiayi, et al.
Published: (2024)
InfoCTM: A Mutual Information Maximization Perspective of Cross-Lingual Topic Modeling
by: Wu, Xiaobao, et al.
Published: (2023)
by: Wu, Xiaobao, et al.
Published: (2023)
Towards a Mechanistic Understanding of Large Reasoning Models: A Survey of Training, Inference, and Failures
by: Hu, Yi, et al.
Published: (2026)
by: Hu, Yi, et al.
Published: (2026)
Dynamic Evaluation for Oversensitivity in LLMs
by: Pu, Sophia Xiao, et al.
Published: (2025)
by: Pu, Sophia Xiao, et al.
Published: (2025)
Planning, Creation, Usage: Benchmarking LLMs for Comprehensive Tool Utilization in Real-World Complex Scenarios
by: Huang, Shijue, et al.
Published: (2024)
by: Huang, Shijue, et al.
Published: (2024)
DetectRL: Benchmarking LLM-Generated Text Detection in Real-World Scenarios
by: Wu, Junchao, et al.
Published: (2024)
by: Wu, Junchao, et al.
Published: (2024)
COrAL: Order-Agnostic Language Modeling for Efficient Iterative Refinement
by: Xie, Yuxi, et al.
Published: (2024)
by: Xie, Yuxi, et al.
Published: (2024)
MetaRuleGPT: Recursive Numerical Reasoning of Language Models Trained with Simple Rules
by: Chen, Kejie, et al.
Published: (2024)
by: Chen, Kejie, et al.
Published: (2024)
LEDOM: Reverse Language Model
by: Yin, Xunjian, et al.
Published: (2025)
by: Yin, Xunjian, et al.
Published: (2025)
AgentCoMa: A Compositional Benchmark Mixing Commonsense and Mathematical Reasoning in Real-World Scenarios
by: Alazraki, Lisa, et al.
Published: (2025)
by: Alazraki, Lisa, et al.
Published: (2025)
Baba Is AI: Break the Rules to Beat the Benchmark
by: Cloos, Nathan, et al.
Published: (2024)
by: Cloos, Nathan, et al.
Published: (2024)
ConciseRL: Conciseness-Guided Reinforcement Learning for Efficient Reasoning Models
by: Dumitru, Razvan-Gabriel, et al.
Published: (2025)
by: Dumitru, Razvan-Gabriel, et al.
Published: (2025)
Understanding Reasoning Ability of Language Models From the Perspective of Reasoning Paths Aggregation
by: Wang, Xinyi, et al.
Published: (2024)
by: Wang, Xinyi, et al.
Published: (2024)
RAG-Zeval: Towards Robust and Interpretable Evaluation on RAG Responses through End-to-End Rule-Guided Reasoning
by: Li, Kun, et al.
Published: (2025)
by: Li, Kun, et al.
Published: (2025)
Are LLMs Good Zero-Shot Fallacy Classifiers?
by: Pan, Fengjun, et al.
Published: (2024)
by: Pan, Fengjun, et al.
Published: (2024)
What if LLMs Have Different World Views: Simulating Alien Civilizations with LLM-based Agents
by: Xue, Zhaoqian, et al.
Published: (2024)
by: Xue, Zhaoqian, et al.
Published: (2024)
Can LLMs Follow Simple Rules?
by: Mu, Norman, et al.
Published: (2023)
by: Mu, Norman, et al.
Published: (2023)
RuleRAG: Rule-Guided Retrieval-Augmented Generation with Language Models for Question Answering
by: Chen, Zhongwu, et al.
Published: (2024)
by: Chen, Zhongwu, et al.
Published: (2024)
Similar Items
-
Epistemic Context Learning: Building Trust the Right Way in LLM-Based Multi-Agent Systems
by: Zhou, Ruiwen, et al.
Published: (2026) -
AntiLeakBench: Preventing Data Contamination by Automatically Constructing Benchmarks with Updated Real-World Knowledge
by: Wu, Xiaobao, et al.
Published: (2024) -
Atomic Skills are the Prerequisite: When Reinforcement Learning Synthesizes Compositional Reasoning, and When It Only Amplifies
by: Cheng, Sitao, et al.
Published: (2025) -
InductionBench: LLMs Fail in the Simplest Complexity Class
by: Hua, Wenyue, et al.
Published: (2025) -
Understanding the Interplay between Parametric and Contextual Knowledge for Large Language Models
by: Cheng, Sitao, et al.
Published: (2024)