NPHardEval: Dynamic Benchmark on Reasoning Ability of Large Language Models via Complexity Classes
Fuente:
arXiv
Guardado en:
| Autores principales: | Fan, Lizhou, Hua, Wenyue, Li, Lingyao, Ling, Haoyang, Zhang, Yongfeng |
|---|---|
| Formato: | Preprint |
| Publicado: |
2023
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
NPHardEval4V: Dynamic Evaluation of Large Vision-Language Models with Effects of Vision
por: Li, Xiang, et al.
Publicado: (2024)
por: Li, Xiang, et al.
Publicado: (2024)
War and Peace (WarAgent): Large Language Model-based Multi-Agent Simulation of World Wars
por: Hua, Wenyue, et al.
Publicado: (2023)
por: Hua, Wenyue, et al.
Publicado: (2023)
Disentangling Logic: The Role of Context in Large Language Model Reasoning Capabilities
por: Hua, Wenyue, et al.
Publicado: (2024)
por: Hua, Wenyue, et al.
Publicado: (2024)
Disentangling Memory and Reasoning Ability in Large Language Models
por: Jin, Mingyu, et al.
Publicado: (2024)
por: Jin, Mingyu, et al.
Publicado: (2024)
BattleAgent: Multi-modal Dynamic Emulation on Historical Battles to Complement Historical Analysis
por: Lin, Shuhang, et al.
Publicado: (2024)
por: Lin, Shuhang, et al.
Publicado: (2024)
The Impact of Reasoning Step Length on Large Language Models
por: Jin, Mingyu, et al.
Publicado: (2024)
por: Jin, Mingyu, et al.
Publicado: (2024)
Large Language Models in Biomedical and Health Informatics: A Review with Bibliometric Analysis
por: Yu, Huizi, et al.
Publicado: (2024)
por: Yu, Huizi, et al.
Publicado: (2024)
The Computational Limits of State-Space Models and Mamba via the Lens of Circuit Complexity
por: Chen, Yifang, et al.
Publicado: (2024)
por: Chen, Yifang, et al.
Publicado: (2024)
CMAT: A Multi-Agent Collaboration Tuning Framework for Enhancing Small Language Models
por: Liang, Xuechen, et al.
Publicado: (2024)
por: Liang, Xuechen, et al.
Publicado: (2024)
BigO(Bench) -- Can LLMs Generate Code with Controlled Time and Space Complexity?
por: Chambon, Pierre, et al.
Publicado: (2025)
por: Chambon, Pierre, et al.
Publicado: (2025)
Circuit Complexity Bounds for Visual Autoregressive Model
por: Ke, Yekun, et al.
Publicado: (2025)
por: Ke, Yekun, et al.
Publicado: (2025)
EmojiPrompt: Generative Prompt Obfuscation for Privacy-Preserving Communication with Cloud-based LLMs
por: Lin, Sam, et al.
Publicado: (2024)
por: Lin, Sam, et al.
Publicado: (2024)
Journalists, Emotions, and the Introduction of Generative AI Chatbots: A Large-Scale Analysis of Tweets Before and After the Launch of ChatGPT
por: Lewis, Seth C., et al.
Publicado: (2024)
por: Lewis, Seth C., et al.
Publicado: (2024)
"HOT" ChatGPT: The promise of ChatGPT in detecting and discriminating hateful, offensive, and toxic comments on social media
por: Li, Lingyao, et al.
Publicado: (2023)
por: Li, Lingyao, et al.
Publicado: (2023)
StrucText-Eval: Evaluating Large Language Model's Reasoning Ability in Structure-Rich Text
por: Gu, Zhouhong, et al.
Publicado: (2024)
por: Gu, Zhouhong, et al.
Publicado: (2024)
Neural Algorithmic Reasoning for Hypergraphs with Looped Transformers
por: Huang, Zekai, et al.
Publicado: (2025)
por: Huang, Zekai, et al.
Publicado: (2025)
AttackEval: How to Evaluate the Effectiveness of Jailbreak Attacking on Large Language Models
por: Shu, Dong, et al.
Publicado: (2024)
por: Shu, Dong, et al.
Publicado: (2024)
Multi-LogiEval: Towards Evaluating Multi-Step Logical Reasoning Ability of Large Language Models
por: Patel, Nisarg, et al.
Publicado: (2024)
por: Patel, Nisarg, et al.
Publicado: (2024)
Circuit Complexity Bounds for RoPE-based Transformer Architecture
por: Chen, Bo, et al.
Publicado: (2024)
por: Chen, Bo, et al.
Publicado: (2024)
On Fine-Grained I/O Complexity of Attention Backward Passes
por: Li, Xiaoyu, et al.
Publicado: (2024)
por: Li, Xiaoyu, et al.
Publicado: (2024)
A Measure-Theoretic Analysis of Reasoning: Structural Generalization and Approximation Limits
por: Zhang, Yuyang, et al.
Publicado: (2026)
por: Zhang, Yuyang, et al.
Publicado: (2026)
2-ASP(Q) programs with weak constraints: Complexity and efficient implementation
por: Cuteri, Andrea, et al.
Publicado: (2026)
por: Cuteri, Andrea, et al.
Publicado: (2026)
LogicGame: Benchmarking Rule-Based Reasoning Abilities of Large Language Models
por: Gui, Jiayi, et al.
Publicado: (2024)
por: Gui, Jiayi, et al.
Publicado: (2024)
Simulating Weighted Automata over Sequences and Trees with Transformers
por: Rizvi, Michael, et al.
Publicado: (2024)
por: Rizvi, Michael, et al.
Publicado: (2024)
Quantifying over Optimum Answer Sets
por: Mazzotta, Giuseppe, et al.
Publicado: (2024)
por: Mazzotta, Giuseppe, et al.
Publicado: (2024)
The Fine-Grained Complexity of Gradient Computation for Training Large Language Models
por: Alman, Josh, et al.
Publicado: (2024)
por: Alman, Josh, et al.
Publicado: (2024)
Toward Equitable Access: Leveraging Crowdsourced Reviews to Investigate Public Perceptions of Health Resource Accessibility
por: Xue, Zhaoqian, et al.
Publicado: (2025)
por: Xue, Zhaoqian, et al.
Publicado: (2025)
Have Large Language Models Learned to Reason? A Characterization via 3-SAT Phase Transition
por: Hazra, Rishi, et al.
Publicado: (2025)
por: Hazra, Rishi, et al.
Publicado: (2025)
Transformers Can Represent $n$-gram Language Models
por: Svete, Anej, et al.
Publicado: (2024)
por: Svete, Anej, et al.
Publicado: (2024)
Game-theoretic LLM: Agent Workflow for Negotiation Games
por: Hua, Wenyue, et al.
Publicado: (2024)
por: Hua, Wenyue, et al.
Publicado: (2024)
Characterizing Online Toxicity During the 2022 Mpox Outbreak: A Computational Analysis of Topical and Network Dynamics
por: Fan, Lizhou, et al.
Publicado: (2024)
por: Fan, Lizhou, et al.
Publicado: (2024)
AccessEval: Benchmarking Disability Bias in Large Language Models
por: Panda, Srikant, et al.
Publicado: (2025)
por: Panda, Srikant, et al.
Publicado: (2025)
AdEval: Alignment-based Dynamic Evaluation to Mitigate Data Contamination in Large Language Models
por: Fan, Yang
Publicado: (2025)
por: Fan, Yang
Publicado: (2025)
Exposing Numeracy Gaps: A Benchmark to Evaluate Fundamental Numerical Abilities in Large Language Models
por: Li, Haoyang, et al.
Publicado: (2025)
por: Li, Haoyang, et al.
Publicado: (2025)
Formal-LLM: Integrating Formal Language and Natural Language for Controllable LLM-based Agents
por: Li, Zelong, et al.
Publicado: (2024)
por: Li, Zelong, et al.
Publicado: (2024)
AutoFlow: Automated Workflow Generation for Large Language Model Agents
por: Li, Zelong, et al.
Publicado: (2024)
por: Li, Zelong, et al.
Publicado: (2024)
Eliciting Causal Abilities in Large Language Models for Reasoning Tasks
por: Wang, Yajing, et al.
Publicado: (2024)
por: Wang, Yajing, et al.
Publicado: (2024)
Parameterized Complexity Of Representing Models Of MSO Formulas
por: Kučera, Petr, et al.
Publicado: (2026)
por: Kučera, Petr, et al.
Publicado: (2026)
On the Complexity of Identification in Linear Structural Causal Models
por: Dörfler, Julian, et al.
Publicado: (2024)
por: Dörfler, Julian, et al.
Publicado: (2024)
DivLogicEval: A Framework for Benchmarking Logical Reasoning Evaluation in Large Language Models
por: Chung, Tsz Ting, et al.
Publicado: (2025)
por: Chung, Tsz Ting, et al.
Publicado: (2025)
Ejemplares similares
-
NPHardEval4V: Dynamic Evaluation of Large Vision-Language Models with Effects of Vision
por: Li, Xiang, et al.
Publicado: (2024) -
War and Peace (WarAgent): Large Language Model-based Multi-Agent Simulation of World Wars
por: Hua, Wenyue, et al.
Publicado: (2023) -
Disentangling Logic: The Role of Context in Large Language Model Reasoning Capabilities
por: Hua, Wenyue, et al.
Publicado: (2024) -
Disentangling Memory and Reasoning Ability in Large Language Models
por: Jin, Mingyu, et al.
Publicado: (2024) -
BattleAgent: Multi-modal Dynamic Emulation on Historical Battles to Complement Historical Analysis
por: Lin, Shuhang, et al.
Publicado: (2024)