Salvato in:
| Autori principali: | Agarwal, Shradha, Rajbhar, Deepak, J, Tariq |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | https://arxiv.org/abs/2605.16675 |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
AlgBench: To What Extent Do Large Reasoning Models Understand Algorithms?
di: Sun, Henan, et al.
Pubblicazione: (2026)
di: Sun, Henan, et al.
Pubblicazione: (2026)
Mathematical Proof as a Litmus Test: Revealing Failure Modes of Advanced Large Reasoning Models
di: Guo, Dadi, et al.
Pubblicazione: (2025)
di: Guo, Dadi, et al.
Pubblicazione: (2025)
CombiBench: Benchmarking LLM Capability for Combinatorial Mathematics
di: Liu, Junqi, et al.
Pubblicazione: (2025)
di: Liu, Junqi, et al.
Pubblicazione: (2025)
LinTree: Improving LLM Reasoning with Explicitly Structured Search Histories
di: Kang, Liwei, et al.
Pubblicazione: (2026)
di: Kang, Liwei, et al.
Pubblicazione: (2026)
AlgOS: Algorithm Operating System
di: Salt, Llewyn, et al.
Pubblicazione: (2025)
di: Salt, Llewyn, et al.
Pubblicazione: (2025)
Riemann-Bench: A Benchmark for Moonshot Mathematics
di: Garre, Suhaas, et al.
Pubblicazione: (2026)
di: Garre, Suhaas, et al.
Pubblicazione: (2026)
Revealing Interpretable Failure Modes of VLMs
di: Chaudhary, Isha, et al.
Pubblicazione: (2026)
di: Chaudhary, Isha, et al.
Pubblicazione: (2026)
ConstraintBench: Benchmarking LLM Constraint Reasoning on Direct Optimization
di: Tso, Joseph, et al.
Pubblicazione: (2026)
di: Tso, Joseph, et al.
Pubblicazione: (2026)
Style Outweighs Substance: Failure Modes of LLM Judges in Alignment Benchmarking
di: Feuer, Benjamin, et al.
Pubblicazione: (2024)
di: Feuer, Benjamin, et al.
Pubblicazione: (2024)
Large Language Models and Mathematical Reasoning Failures
di: Boye, Johan, et al.
Pubblicazione: (2025)
di: Boye, Johan, et al.
Pubblicazione: (2025)
BenchHub: A Unified Benchmark Suite for Holistic and Customizable LLM Evaluation
di: Kim, Eunsu, et al.
Pubblicazione: (2025)
di: Kim, Eunsu, et al.
Pubblicazione: (2025)
UGMathBench: A Diverse and Dynamic Benchmark for Undergraduate-Level Mathematical Reasoning with Large Language Models
di: Xu, Xin, et al.
Pubblicazione: (2025)
di: Xu, Xin, et al.
Pubblicazione: (2025)
HiBench: Benchmarking LLMs Capability on Hierarchical Structure Reasoning
di: Jiang, Zhuohang, et al.
Pubblicazione: (2025)
di: Jiang, Zhuohang, et al.
Pubblicazione: (2025)
CAM-Bench: A Benchmark for Computational and Applied Mathematics in Lean
di: Long, Wentao, et al.
Pubblicazione: (2026)
di: Long, Wentao, et al.
Pubblicazione: (2026)
HeroBench: A Benchmark for Long-Horizon Planning and Structured Reasoning in Virtual Worlds
di: Anokhin, Petr, et al.
Pubblicazione: (2025)
di: Anokhin, Petr, et al.
Pubblicazione: (2025)
T2S-Bench & Structure-of-Thought: Benchmarking and Prompting Comprehensive Text-to-Structure Reasoning
di: Wang, Qinsi, et al.
Pubblicazione: (2026)
di: Wang, Qinsi, et al.
Pubblicazione: (2026)
SMH-Bench: Benchmarking LLM Agents for Environment-Grounded Reasoning and Action in Smart Homes
di: Li, Kuan, et al.
Pubblicazione: (2026)
di: Li, Kuan, et al.
Pubblicazione: (2026)
Unmasking Reasoning Processes: A Process-aware Benchmark for Evaluating Structural Mathematical Reasoning in LLMs
di: Zheng, Xiang, et al.
Pubblicazione: (2026)
di: Zheng, Xiang, et al.
Pubblicazione: (2026)
Conversational Speech Reveals Structural Robustness Failures in SpeechLLM Backbones
di: Teleki, Maria, et al.
Pubblicazione: (2025)
di: Teleki, Maria, et al.
Pubblicazione: (2025)
KG-LLM-Bench: A Scalable Benchmark for Evaluating LLM Reasoning on Textualized Knowledge Graphs
di: Markowitz, Elan, et al.
Pubblicazione: (2025)
di: Markowitz, Elan, et al.
Pubblicazione: (2025)
MorphoBench: A Benchmark with Difficulty Adaptive to Model Reasoning
di: Wang, Xukai, et al.
Pubblicazione: (2025)
di: Wang, Xukai, et al.
Pubblicazione: (2025)
ProcessBench: Identifying Process Errors in Mathematical Reasoning
di: Zheng, Chujie, et al.
Pubblicazione: (2024)
di: Zheng, Chujie, et al.
Pubblicazione: (2024)
The Confidence Shortcut: A Reasoning Failure Mode of Masked Diffusion Models
di: Kim, Dueun, et al.
Pubblicazione: (2026)
di: Kim, Dueun, et al.
Pubblicazione: (2026)
IndiMathBench: Autoformalizing Mathematical Reasoning Problems with a Human Touch
di: Biyani, Param, et al.
Pubblicazione: (2025)
di: Biyani, Param, et al.
Pubblicazione: (2025)
MMR-Bench: A Comprehensive Benchmark for Multimodal LLM Routing
di: Ma, Haoxuan, et al.
Pubblicazione: (2026)
di: Ma, Haoxuan, et al.
Pubblicazione: (2026)
NC-Bench: An LLM Benchmark for Evaluating Conversational Competence
di: Moore, Robert J., et al.
Pubblicazione: (2026)
di: Moore, Robert J., et al.
Pubblicazione: (2026)
RoMath: A Mathematical Reasoning Benchmark in Romanian
di: Cosma, Adrian, et al.
Pubblicazione: (2024)
di: Cosma, Adrian, et al.
Pubblicazione: (2024)
GlobalDentBench: A Multinational Benchmark for Evaluating LLM Clinical Reasoning in Dentistry with Expert Calibration
di: Zhao, Junjie, et al.
Pubblicazione: (2026)
di: Zhao, Junjie, et al.
Pubblicazione: (2026)
Failure Modes in LLM Systems: A System-Level Taxonomy for Reliable AI Applications
di: Vinay, Vaishali
Pubblicazione: (2025)
di: Vinay, Vaishali
Pubblicazione: (2025)
CXReasonBench: A Benchmark for Evaluating Structured Diagnostic Reasoning in Chest X-rays
di: Lee, Hyungyung, et al.
Pubblicazione: (2025)
di: Lee, Hyungyung, et al.
Pubblicazione: (2025)
FeynmanBench: Benchmarking Multimodal LLMs on Diagrammatic Physics Reasoning
di: Wang, Zeyu, et al.
Pubblicazione: (2026)
di: Wang, Zeyu, et al.
Pubblicazione: (2026)
FrontierMath: A Benchmark for Evaluating Advanced Mathematical Reasoning in AI
di: Glazer, Elliot, et al.
Pubblicazione: (2024)
di: Glazer, Elliot, et al.
Pubblicazione: (2024)
AttuneBench: A Conversation-Based Benchmark for LLM Emotional Intelligence
di: Lubrano, Kate M., et al.
Pubblicazione: (2026)
di: Lubrano, Kate M., et al.
Pubblicazione: (2026)
LLMRouterBench: A Massive Benchmark and Unified Framework for LLM Routing
di: Li, Hao, et al.
Pubblicazione: (2026)
di: Li, Hao, et al.
Pubblicazione: (2026)
Real-Time Deadlines Reveal Temporal Awareness Failures in LLM Strategic Dialogues
di: Sehgal, Neil K. R., et al.
Pubblicazione: (2026)
di: Sehgal, Neil K. R., et al.
Pubblicazione: (2026)
OckBench: Measuring the Efficiency of LLM Reasoning
di: Du, Zheng, et al.
Pubblicazione: (2025)
di: Du, Zheng, et al.
Pubblicazione: (2025)
TopoBench: Benchmarking LLMs on Hard Topological Reasoning
di: Maniparambil, Mayug, et al.
Pubblicazione: (2026)
di: Maniparambil, Mayug, et al.
Pubblicazione: (2026)
FractalBench: Diagnosing Visual-Mathematical Reasoning Through Recursive Program Synthesis
di: Ondras, Jan, et al.
Pubblicazione: (2025)
di: Ondras, Jan, et al.
Pubblicazione: (2025)
ReactBench: A Benchmark for Topological Reasoning in MLLMs on Chemical Reaction Diagrams
di: Xu, Qiang, et al.
Pubblicazione: (2026)
di: Xu, Qiang, et al.
Pubblicazione: (2026)
FAM-Bench: A Multimodal Benchmark for Condition-Aware Food-as-Medicine Reasoning
di: Mao, Mingyang, et al.
Pubblicazione: (2026)
di: Mao, Mingyang, et al.
Pubblicazione: (2026)
Documenti analoghi
-
AlgBench: To What Extent Do Large Reasoning Models Understand Algorithms?
di: Sun, Henan, et al.
Pubblicazione: (2026) -
Mathematical Proof as a Litmus Test: Revealing Failure Modes of Advanced Large Reasoning Models
di: Guo, Dadi, et al.
Pubblicazione: (2025) -
CombiBench: Benchmarking LLM Capability for Combinatorial Mathematics
di: Liu, Junqi, et al.
Pubblicazione: (2025) -
LinTree: Improving LLM Reasoning with Explicitly Structured Search Histories
di: Kang, Liwei, et al.
Pubblicazione: (2026) -
AlgOS: Algorithm Operating System
di: Salt, Llewyn, et al.
Pubblicazione: (2025)