R-ConstraintBench: Evaluating LLMs on NP-Complete Scheduling
Fuente:
arXiv
Saved in:
| Main Authors: | Jain, Raj, Wetter, Marc |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
ConstraintBench: Benchmarking LLM Constraint Reasoning on Direct Optimization
by: Tso, Joseph, et al.
Published: (2026)
by: Tso, Joseph, et al.
Published: (2026)
Implicit Intelligence -- Evaluating Agents on What Users Don't Say
by: Sirdeshmukh, Ved, et al.
Published: (2026)
by: Sirdeshmukh, Ved, et al.
Published: (2026)
Intent Laundering: AI Safety Datasets Are Not What They Seem
by: Golchin, Shahriar, et al.
Published: (2026)
by: Golchin, Shahriar, et al.
Published: (2026)
DCP-Bench-Open: Evaluating LLMs for Constraint Modelling of Discrete Combinatorial Problems
by: Michailidis, Kostis, et al.
Published: (2025)
by: Michailidis, Kostis, et al.
Published: (2025)
MoralBench: Moral Evaluation of LLMs
by: Ji, Jianchao, et al.
Published: (2024)
by: Ji, Jianchao, et al.
Published: (2024)
TPS-Bench: Evaluating AI Agents' Tool Planning \& Scheduling Abilities in Compounding Tasks
by: Xu, Hanwen, et al.
Published: (2025)
by: Xu, Hanwen, et al.
Published: (2025)
Codev-Bench: How Do LLMs Understand Developer-Centric Code Completion?
by: Pan, Zhenyu, et al.
Published: (2024)
by: Pan, Zhenyu, et al.
Published: (2024)
CCR-Bench: A Comprehensive Benchmark for Evaluating LLMs on Complex Constraints, Control Flows, and Real-World Cases
by: Xue, Xiaona, et al.
Published: (2026)
by: Xue, Xiaona, et al.
Published: (2026)
Autonomous Code Evolution Meets NP-Completeness
by: Yu, Cunxi, et al.
Published: (2025)
by: Yu, Cunxi, et al.
Published: (2025)
LLMs can Schedule
by: Abgaryan, Henrik, et al.
Published: (2024)
by: Abgaryan, Henrik, et al.
Published: (2024)
EchoChain: A Full-Duplex Benchmark for State-Update Reasoning Under Interruptions
by: Modi, Smit Nautambhai, et al.
Published: (2026)
by: Modi, Smit Nautambhai, et al.
Published: (2026)
AgentBench: Evaluating LLMs as Agents
by: Liu, Xiao, et al.
Published: (2023)
by: Liu, Xiao, et al.
Published: (2023)
Forge: Quality-Aware Reinforcement Learning for NP-Hard Optimization in LLMs
by: Li, Xiaozhe, et al.
Published: (2026)
by: Li, Xiaozhe, et al.
Published: (2026)
MCJudgeBench: A Benchmark for Constraint-Level Judge Evaluation in Multi-Constraint Instruction Following
by: Lee, Jaeyun, et al.
Published: (2026)
by: Lee, Jaeyun, et al.
Published: (2026)
AudioMotionBench: Evaluating Auditory Motion Perception in Audio LLMs
by: Sun, Zhe, et al.
Published: (2025)
by: Sun, Zhe, et al.
Published: (2025)
IDA-Bench: Evaluating LLMs on Interactive Guided Data Analysis
by: Li, Hanyu, et al.
Published: (2025)
by: Li, Hanyu, et al.
Published: (2025)
CyberCertBench: Evaluating LLMs in Cybersecurity Certification Knowledge
by: Keppler, Gustav, et al.
Published: (2026)
by: Keppler, Gustav, et al.
Published: (2026)
DialogBench: Evaluating LLMs as Human-like Dialogue Systems
by: Ou, Jiao, et al.
Published: (2023)
by: Ou, Jiao, et al.
Published: (2023)
Magis-Bench: Evaluating LLMs on Magistrate-Level Legal Tasks
by: Pires, Ramon, et al.
Published: (2026)
by: Pires, Ramon, et al.
Published: (2026)
FullStack Bench: Evaluating LLMs as Full Stack Coders
by: Bytedance-Seed-Foundation-Code-Team, et al.
Published: (2024)
by: Bytedance-Seed-Foundation-Code-Team, et al.
Published: (2024)
CivBench: Progress-Based Evaluation for LLMs' Strategic Decision-Making in Civilization V
by: Chen, John, et al.
Published: (2026)
by: Chen, John, et al.
Published: (2026)
Machine Learning and Constraint Programming for Efficient Healthcare Scheduling
by: Said, Aymen Ben, et al.
Published: (2024)
by: Said, Aymen Ben, et al.
Published: (2024)
TRIAGE: Evaluating Prospective Metacognitive Control in LLMs under Resource Constraints
by: Nazi, Zabir Al, et al.
Published: (2026)
by: Nazi, Zabir Al, et al.
Published: (2026)
CricBench: A Multilingual Benchmark for Evaluating LLMs in Cricket Analytics
by: Agarwal, Parth, et al.
Published: (2025)
by: Agarwal, Parth, et al.
Published: (2025)
LiveMedBench: A Contamination-Free Medical Benchmark for LLMs with Automated Rubric Evaluation
by: Yan, Zhiling, et al.
Published: (2026)
by: Yan, Zhiling, et al.
Published: (2026)
FrontendBench: A Benchmark for Evaluating LLMs on Front-End Development via Automatic Evaluation
by: Zhu, Hongda, et al.
Published: (2025)
by: Zhu, Hongda, et al.
Published: (2025)
Which LLMs Get the Joke? Probing Non-STEM Reasoning Abilities with HumorBench
by: Narad, Reuben, et al.
Published: (2025)
by: Narad, Reuben, et al.
Published: (2025)
LongCodeBench: Evaluating Coding LLMs at 1M Context Windows
by: Rando, Stefano, et al.
Published: (2025)
by: Rando, Stefano, et al.
Published: (2025)
AthenaBench: A Dynamic Benchmark for Evaluating LLMs in Cyber Threat Intelligence
by: Alam, Md Tanvirul, et al.
Published: (2025)
by: Alam, Md Tanvirul, et al.
Published: (2025)
CalBench: Evaluating Coordination-Privacy Trade-offs in Multi-Agent LLMs
by: Zou, Chelsea, et al.
Published: (2026)
by: Zou, Chelsea, et al.
Published: (2026)
ProMoral-Bench: Evaluating Prompting Strategies for Moral Reasoning and Safety in LLMs
by: Thomas, Rohan Subramanian, et al.
Published: (2026)
by: Thomas, Rohan Subramanian, et al.
Published: (2026)
NP-Engine: Empowering Optimization Reasoning in Large Language Models with Verifiable Synthetic NP Problems
by: Li, Xiaozhe, et al.
Published: (2025)
by: Li, Xiaozhe, et al.
Published: (2025)
Bottleneck Identification in Resource-Constrained Project Scheduling via Constraint Relaxation
by: Nedbálek, Lukáš, et al.
Published: (2025)
by: Nedbálek, Lukáš, et al.
Published: (2025)
ClinDet-Bench: Beyond Abstention, Evaluating Judgment Determinability of LLMs in Clinical Decision-Making
by: Watanabe, Yusuke, et al.
Published: (2026)
by: Watanabe, Yusuke, et al.
Published: (2026)
AbstentionBench: Reasoning LLMs Fail on Unanswerable Questions
by: Kirichenko, Polina, et al.
Published: (2025)
by: Kirichenko, Polina, et al.
Published: (2025)
lmgame-Bench: How Good are LLMs at Playing Games?
by: Hu, Lanxiang, et al.
Published: (2025)
by: Hu, Lanxiang, et al.
Published: (2025)
ReEfBench: Quantifying the Reasoning Efficiency of LLMs
by: Fu, Zhizhang, et al.
Published: (2026)
by: Fu, Zhizhang, et al.
Published: (2026)
ConvexBench: Can LLMs Recognize Convex Functions?
by: Liu, Yepeng, et al.
Published: (2026)
by: Liu, Yepeng, et al.
Published: (2026)
IndustryBench: Probing the Industrial Knowledge Boundaries of LLMs
by: Bai, Songlin, et al.
Published: (2026)
by: Bai, Songlin, et al.
Published: (2026)
SKA-Bench: A Fine-Grained Benchmark for Evaluating Structured Knowledge Understanding of LLMs
by: Liu, Zhiqiang, et al.
Published: (2025)
by: Liu, Zhiqiang, et al.
Published: (2025)
Similar Items
-
ConstraintBench: Benchmarking LLM Constraint Reasoning on Direct Optimization
by: Tso, Joseph, et al.
Published: (2026) -
Implicit Intelligence -- Evaluating Agents on What Users Don't Say
by: Sirdeshmukh, Ved, et al.
Published: (2026) -
Intent Laundering: AI Safety Datasets Are Not What They Seem
by: Golchin, Shahriar, et al.
Published: (2026) -
DCP-Bench-Open: Evaluating LLMs for Constraint Modelling of Discrete Combinatorial Problems
by: Michailidis, Kostis, et al.
Published: (2025) -
MoralBench: Moral Evaluation of LLMs
by: Ji, Jianchao, et al.
Published: (2024)