Functional Benchmarks for Robust Evaluation of Reasoning Performance, and the Reasoning Gap
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Srivastava, Saurabh, B, Annarose M, P V, Anto, Menon, Shashank, Sukumar, Ajay, T, Adwaith Samod, Philipose, Alan, Prince, Stevin, Thomas, Sooraj |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Neurosymbolic Language Reasoning as Satisfiability Modulo Theory
von: Oh, Hyunseok, et al.
Veröffentlicht: (2026)
von: Oh, Hyunseok, et al.
Veröffentlicht: (2026)
The Point of No Return: Counterfactual Localization of Deceptive Commitment in Language-Model Reasoning
von: Merrill, Scott, et al.
Veröffentlicht: (2026)
von: Merrill, Scott, et al.
Veröffentlicht: (2026)
DISCERN: Decoding Systematic Errors in Natural Language for Text Classifiers
von: Menon, Rakesh R., et al.
Veröffentlicht: (2024)
von: Menon, Rakesh R., et al.
Veröffentlicht: (2024)
An Objective Performance Evaluation of the LSTM Networks in Time Series Classification
von: Sunil, Sooraj, et al.
Veröffentlicht: (2026)
von: Sunil, Sooraj, et al.
Veröffentlicht: (2026)
Revisiting Prompt Optimization with Large Reasoning Models-A Case Study on Event Extraction
von: Srivastava, Saurabh, et al.
Veröffentlicht: (2025)
von: Srivastava, Saurabh, et al.
Veröffentlicht: (2025)
Benchmarking Spatiotemporal Reasoning in LLMs and Reasoning Models: Capabilities and Challenges
von: Quan, Pengrui, et al.
Veröffentlicht: (2025)
von: Quan, Pengrui, et al.
Veröffentlicht: (2025)
A Causal Lens for Evaluating Faithfulness Metrics
von: Zaman, Kerem, et al.
Veröffentlicht: (2025)
von: Zaman, Kerem, et al.
Veröffentlicht: (2025)
Voice Evaluation of Reasoning Ability: Diagnosing the Modality-Induced Performance Gap
von: Lin, Yueqian, et al.
Veröffentlicht: (2025)
von: Lin, Yueqian, et al.
Veröffentlicht: (2025)
Bridging the Arithmetic Gap: The Cognitive Complexity Benchmark and Financial-PoT for Robust Financial Reasoning
von: Zhao, Boxiang, et al.
Veröffentlicht: (2026)
von: Zhao, Boxiang, et al.
Veröffentlicht: (2026)
Evaluating Chain-of-Thought Reasoning through Reusability and Verifiability
von: Aggarwal, Shashank, et al.
Veröffentlicht: (2026)
von: Aggarwal, Shashank, et al.
Veröffentlicht: (2026)
Continuous Optimization for Decoding Errors
von: Srivastava, Shashank
Veröffentlicht: (2024)
von: Srivastava, Shashank
Veröffentlicht: (2024)
Improved List Size for Folded Reed-Solomon Codes
von: Srivastava, Shashank
Veröffentlicht: (2024)
von: Srivastava, Shashank
Veröffentlicht: (2024)
JudgeBoard: Benchmarking and Enhancing Small Language Models for Reasoning Evaluation
von: Bi, Zhenyu, et al.
Veröffentlicht: (2025)
von: Bi, Zhenyu, et al.
Veröffentlicht: (2025)
INTERACT: Enabling Interactive, Question-Driven Learning in Large Language Models
von: Kendapadi, Aum, et al.
Veröffentlicht: (2024)
von: Kendapadi, Aum, et al.
Veröffentlicht: (2024)
Real-Time Performance Benchmarking of TinyML Models in Embedded Systems (PICO: Performance of Inference, CPU, and Operations)
von: Dey, Abhishek, et al.
Veröffentlicht: (2025)
von: Dey, Abhishek, et al.
Veröffentlicht: (2025)
An Enigma of Artificial Reason: Investigating the Production-Evaluation Gap in Large Reasoning Models
von: Sun, Mingzhong, et al.
Veröffentlicht: (2026)
von: Sun, Mingzhong, et al.
Veröffentlicht: (2026)
Mind the Gap: Benchmarking Spatial Reasoning in Vision-Language Models
von: Stogiannidis, Ilias, et al.
Veröffentlicht: (2025)
von: Stogiannidis, Ilias, et al.
Veröffentlicht: (2025)
Cost Trade-offs of Reasoning and Non-Reasoning Large Language Models in Text-to-SQL
von: Deochake, Saurabh, et al.
Veröffentlicht: (2025)
von: Deochake, Saurabh, et al.
Veröffentlicht: (2025)
Integrating Performance Tools in Model Reasoning for GPU Kernel Optimization
von: Nichols, Daniel, et al.
Veröffentlicht: (2025)
von: Nichols, Daniel, et al.
Veröffentlicht: (2025)
RUPBench: Benchmarking Reasoning Under Perturbations for Robustness Evaluation in Large Language Models
von: Wang, Yuqing, et al.
Veröffentlicht: (2024)
von: Wang, Yuqing, et al.
Veröffentlicht: (2024)
ECG-Reasoning-Benchmark: A Benchmark for Evaluating Clinical Reasoning Capabilities in ECG Interpretation
von: Oh, Jungwoo, et al.
Veröffentlicht: (2026)
von: Oh, Jungwoo, et al.
Veröffentlicht: (2026)
Robustness and Reasoning Fidelity of Large Language Models in Long-Context Code Question Answering
von: Maharaj, Kishan, et al.
Veröffentlicht: (2026)
von: Maharaj, Kishan, et al.
Veröffentlicht: (2026)
Benchmarking Reasoning Robustness in Large Language Models
von: Yu, Tong, et al.
Veröffentlicht: (2025)
von: Yu, Tong, et al.
Veröffentlicht: (2025)
Mind the Gap: Evaluating the Representativeness of Quantitative Medical Language Reasoning LLM Benchmarks for African Disease Burdens
von: Mutisya, Fred, et al.
Veröffentlicht: (2025)
von: Mutisya, Fred, et al.
Veröffentlicht: (2025)
Enhancing Domain-Specific Retrieval-Augmented Generation: Synthetic Data Generation and Evaluation using Reasoning Models
von: Jadon, Aryan, et al.
Veröffentlicht: (2025)
von: Jadon, Aryan, et al.
Veröffentlicht: (2025)
LLMs Cannot Reliably Identify and Reason About Security Vulnerabilities (Yet?): A Comprehensive Evaluation, Framework, and Benchmarks
von: Ullah, Saad, et al.
Veröffentlicht: (2023)
von: Ullah, Saad, et al.
Veröffentlicht: (2023)
Enhancing the Diagnostic Evaluation of Thyroid Functionality Using Diffuse Reflectance Spectroscopy and Regression Models
von: W. Anto Win Shalini, et al.
Veröffentlicht: (2025)
von: W. Anto Win Shalini, et al.
Veröffentlicht: (2025)
H-ARC: A Robust Estimate of Human Performance on the Abstraction and Reasoning Corpus Benchmark
von: LeGris, Solim, et al.
Veröffentlicht: (2024)
von: LeGris, Solim, et al.
Veröffentlicht: (2024)
From Reasoning to Pixels: Benchmarking the Alignment Gap in Unified Multimodal Models
von: Yang, Cheng, et al.
Veröffentlicht: (2026)
von: Yang, Cheng, et al.
Veröffentlicht: (2026)
A Robust Placeability Metric for Model-Free Unified Pick-and-Place Reasoning
von: Wingender, Benno, et al.
Veröffentlicht: (2025)
von: Wingender, Benno, et al.
Veröffentlicht: (2025)
RE-IMAGINE: Symbolic Benchmark Synthesis for Reasoning Evaluation
von: Xu, Xinnuo, et al.
Veröffentlicht: (2025)
von: Xu, Xinnuo, et al.
Veröffentlicht: (2025)
Pharmacognostical and Preliminary phytochemical evaluation of Seed kernel of Chinchasthi (Tamarindus indica Linn)
von: MS Megha, et al.
Veröffentlicht: (2026)
von: MS Megha, et al.
Veröffentlicht: (2026)
The intersection of philosophy of language and artificial intelligence: Challenges in replicating human language understanding
von: Sooraj Kumar Maurya
Veröffentlicht: (2024)
von: Sooraj Kumar Maurya
Veröffentlicht: (2024)
Autonomous Evaluation of LLMs for Truth Maintenance and Reasoning Tasks
von: Karia, Rushang, et al.
Veröffentlicht: (2024)
von: Karia, Rushang, et al.
Veröffentlicht: (2024)
Ai-Powered Sales Demand Forecasting and Desicion Support system
von: M, Nithin, et al.
Veröffentlicht: (2026)
von: M, Nithin, et al.
Veröffentlicht: (2026)
MetaCluster: Enabling Deep Compression of Kolmogorov-Arnold Network
von: Raffel, Matthew, et al.
Veröffentlicht: (2025)
von: Raffel, Matthew, et al.
Veröffentlicht: (2025)
MastermindEval: A Simple But Scalable Reasoning Benchmark
von: Golde, Jonas, et al.
Veröffentlicht: (2025)
von: Golde, Jonas, et al.
Veröffentlicht: (2025)
Bullous Lung Disease in Turner Syndrome: An Underrecognized Comorbidity?
von: Stevin Lu, et al.
Veröffentlicht: (2024)
von: Stevin Lu, et al.
Veröffentlicht: (2024)
The Ouroboros of Benchmarking: Reasoning Evaluation in an Era of Saturation
von: Deveci, İbrahim Ethem, et al.
Veröffentlicht: (2025)
von: Deveci, İbrahim Ethem, et al.
Veröffentlicht: (2025)
Benchmarking and Confidence Evaluation of LALMs For Temporal Reasoning
von: Bhattacharya, Debarpan, et al.
Veröffentlicht: (2025)
von: Bhattacharya, Debarpan, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Neurosymbolic Language Reasoning as Satisfiability Modulo Theory
von: Oh, Hyunseok, et al.
Veröffentlicht: (2026) -
The Point of No Return: Counterfactual Localization of Deceptive Commitment in Language-Model Reasoning
von: Merrill, Scott, et al.
Veröffentlicht: (2026) -
DISCERN: Decoding Systematic Errors in Natural Language for Text Classifiers
von: Menon, Rakesh R., et al.
Veröffentlicht: (2024) -
An Objective Performance Evaluation of the LSTM Networks in Time Series Classification
von: Sunil, Sooraj, et al.
Veröffentlicht: (2026) -
Revisiting Prompt Optimization with Large Reasoning Models-A Case Study on Event Extraction
von: Srivastava, Saurabh, et al.
Veröffentlicht: (2025)