LLMEval-Logic: A Solver-Verified Chinese Benchmark for Logical Reasoning of LLMs with Adversarial Hardening
Fuente:
arXiv
Saved in:
| Main Authors: | Zhang, Ming, Peng, Qiyuan, Wei, Yinxi, Shen, Yujiong, Tan, Kexin, Wang, Yuhui, Xiang, Zhenghao, Ye, Junjie, Yin, Zhangyue, Xi, Zhiheng, Dou, Shihan, Gui, Tao, Pan, Maxm, Yang, Ruizhi, Zhang, Qi, Huang, Xuanjing |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
LLMEval-Med: A Real-world Clinical Benchmark for Medical LLMs with Physician Validation
by: Zhang, Ming, et al.
Published: (2025)
by: Zhang, Ming, et al.
Published: (2025)
LLMEval-Fair: A Large-Scale Longitudinal Study on Robust and Fair Evaluation of Large Language Models
by: Zhang, Ming, et al.
Published: (2025)
by: Zhang, Ming, et al.
Published: (2025)
OpenNovelty: An LLM-powered Agentic System for Verifiable Scholarly Novelty Assessment
by: Zhang, Ming, et al.
Published: (2026)
by: Zhang, Ming, et al.
Published: (2026)
Can Deep Research Agents Retrieve and Organize? Evaluating the Synthesis Gap with Expert Taxonomies
by: Zhang, Ming, et al.
Published: (2026)
by: Zhang, Ming, et al.
Published: (2026)
JFTA-Bench: Evaluate LLM's Ability of Tracking and Analyzing Malfunctions Using Fault Trees
by: Wang, Yuhui, et al.
Published: (2026)
by: Wang, Yuhui, et al.
Published: (2026)
TransferTOD: A Generalizable Chinese Multi-Domain Task-Oriented Dialogue System with Transfer Capabilities
by: Zhang, Ming, et al.
Published: (2024)
by: Zhang, Ming, et al.
Published: (2024)
Beyond Scaling: Measuring and Predicting the Upper Bound of Knowledge Retention in Language Model Pre-Training
by: Jiang, Changhao, et al.
Published: (2025)
by: Jiang, Changhao, et al.
Published: (2025)
Steering LLMs via Scalable Interactive Oversight
by: Zhou, Enyu, et al.
Published: (2026)
by: Zhou, Enyu, et al.
Published: (2026)
Subspace Defense: Discarding Adversarial Perturbations by Learning a Subspace for Clean Signals
by: Zheng, Rui, et al.
Published: (2024)
by: Zheng, Rui, et al.
Published: (2024)
Logic-Regularized Verifier Elicits Reasoning from LLMs
by: Wang, Xinyu, et al.
Published: (2026)
by: Wang, Xinyu, et al.
Published: (2026)
Can Language Models Pretend Solvers? Logic Code Simulation with LLMs
by: Chen, Minyu, et al.
Published: (2024)
by: Chen, Minyu, et al.
Published: (2024)
PFDial: A Structured Dialogue Instruction Fine-tuning Method Based on UML Flowcharts
by: Zhang, Ming, et al.
Published: (2025)
by: Zhang, Ming, et al.
Published: (2025)
Enhancing LLM-based Search Agents via Contribution Weighted Group Relative Policy Optimization
by: Wang, Junzhe, et al.
Published: (2026)
by: Wang, Junzhe, et al.
Published: (2026)
SynLogic: Synthesizing Verifiable Reasoning Data at Scale for Learning Logical Reasoning and Beyond
by: Liu, Junteng, et al.
Published: (2025)
by: Liu, Junteng, et al.
Published: (2025)
SciAgentGym: Benchmarking Multi-Step Scientific Tool-use in LLM Agents
by: Shen, Yujiong, et al.
Published: (2026)
by: Shen, Yujiong, et al.
Published: (2026)
CodeChameleon: Personalized Encryption Framework for Jailbreaking Large Language Models
by: Lv, Huijie, et al.
Published: (2024)
by: Lv, Huijie, et al.
Published: (2024)
From Blind Solvers to Logical Thinkers: Benchmarking LLMs' Logical Integrity on Faulty Mathematical Problems
by: Rahman, A M Muntasir, et al.
Published: (2024)
by: Rahman, A M Muntasir, et al.
Published: (2024)
Logic.py: Bridging the Gap between LLMs and Constraint Solvers
by: Kesseli, Pascal, et al.
Published: (2025)
by: Kesseli, Pascal, et al.
Published: (2025)
Improving RL Exploration for LLM Reasoning through Retrospective Replay
by: Dou, Shihan, et al.
Published: (2025)
by: Dou, Shihan, et al.
Published: (2025)
Last Layer Logits to Logic: Empowering LLMs with Logic-Consistent Structured Knowledge Reasoning
by: Li, Songze, et al.
Published: (2025)
by: Li, Songze, et al.
Published: (2025)
Dsat: A Native SAT Solver for Discrete Logic
by: Zhang, Yaofang, et al.
Published: (2026)
by: Zhang, Yaofang, et al.
Published: (2026)
Distill Visual Chart Reasoning Ability from LLMs to MLLMs
by: He, Wei, et al.
Published: (2024)
by: He, Wei, et al.
Published: (2024)
RoCoIns: Enhancing Robustness of Large Language Models through Code-Style Instructions
by: Zhang, Yuansen, et al.
Published: (2024)
by: Zhang, Yuansen, et al.
Published: (2024)
Agentic Harness Engineering: Observability-Driven Automatic Evolution of Coding-Agent Harnesses
by: Lin, Jiahang, et al.
Published: (2026)
by: Lin, Jiahang, et al.
Published: (2026)
From Scores to Preferences: Redefining MOS Benchmarking for Speech Quality Reward Modeling
by: Cao, Yifei, et al.
Published: (2025)
by: Cao, Yifei, et al.
Published: (2025)
DFPO: Scaling Value Modeling via Distributional Flow towards Robust and Generalizable LLM Post-Training
by: Zhu, Dingwei, et al.
Published: (2026)
by: Zhu, Dingwei, et al.
Published: (2026)
SATQuest: A Verifier for Logical Reasoning Evaluation and Reinforcement Fine-Tuning of LLMs
by: Zhao, Yanxiao, et al.
Published: (2025)
by: Zhao, Yanxiao, et al.
Published: (2025)
Mitigating Object Hallucinations in MLLMs via Multi-Frequency Perturbations
by: Li, Shuo, et al.
Published: (2025)
by: Li, Shuo, et al.
Published: (2025)
TroLLoc: Logic Locking and Layout Hardening for IC Security Closure against Hardware Trojans
by: Wang, Fangzhou, et al.
Published: (2024)
by: Wang, Fangzhou, et al.
Published: (2024)
Instantiation-based Formalization of Logical Reasoning Tasks using Language Models and Logical Solvers
by: Raza, Mohammad, et al.
Published: (2025)
by: Raza, Mohammad, et al.
Published: (2025)
BehaVerify: Verifying Temporal Logic Specifications for Behavior Trees
by: Serbinowska, Serena S., et al.
Published: (2022)
by: Serbinowska, Serena S., et al.
Published: (2022)
AgentV-RL: Scaling Reward Modeling with Agentic Verifier
by: Zhang, Jiazheng, et al.
Published: (2026)
by: Zhang, Jiazheng, et al.
Published: (2026)
Measuring Data Diversity for Instruction Tuning: A Systematic Analysis and A Reliable Metric
by: Yang, Yuming, et al.
Published: (2025)
by: Yang, Yuming, et al.
Published: (2025)
CLOAQ: Combined Logic and Angle Obfuscation for Quantum Circuits
by: Langford, Vincent, et al.
Published: (2026)
by: Langford, Vincent, et al.
Published: (2026)
Logic-Parametric Neuro-Symbolic NLI: Controlling Logical Formalisms for Verifiable LLM Reasoning
by: Farjami, Ali, et al.
Published: (2026)
by: Farjami, Ali, et al.
Published: (2026)
CMDAR: A Chinese Multi-scene Dynamic Audio Reasoning Benchmark with Diverse Challenges
by: Li, Hui, et al.
Published: (2025)
by: Li, Hui, et al.
Published: (2025)
Verification Algorithms for Automated Separation Logic Verifiers
by: Eilers, Marco, et al.
Published: (2024)
by: Eilers, Marco, et al.
Published: (2024)
Categorical Construction of Logically Verifiable Neural Architectures
by: Nye, Logan
Published: (2025)
by: Nye, Logan
Published: (2025)
RMB: Comprehensively Benchmarking Reward Models in LLM Alignment
by: Zhou, Enyu, et al.
Published: (2024)
by: Zhou, Enyu, et al.
Published: (2024)
VRPO: Rethinking Value Modeling for Robust RL Training under Noisy Supervision
by: Zhu, Dingwei, et al.
Published: (2025)
by: Zhu, Dingwei, et al.
Published: (2025)
Similar Items
-
LLMEval-Med: A Real-world Clinical Benchmark for Medical LLMs with Physician Validation
by: Zhang, Ming, et al.
Published: (2025) -
LLMEval-Fair: A Large-Scale Longitudinal Study on Robust and Fair Evaluation of Large Language Models
by: Zhang, Ming, et al.
Published: (2025) -
OpenNovelty: An LLM-powered Agentic System for Verifiable Scholarly Novelty Assessment
by: Zhang, Ming, et al.
Published: (2026) -
Can Deep Research Agents Retrieve and Organize? Evaluating the Synthesis Gap with Expert Taxonomies
by: Zhang, Ming, et al.
Published: (2026) -
JFTA-Bench: Evaluate LLM's Ability of Tracking and Analyzing Malfunctions Using Fault Trees
by: Wang, Yuhui, et al.
Published: (2026)