Measuring AI Reasoning: A Guide for Researchers
Fuente:
arXiv
Saved in:
| Main Authors: | Nwadike, Munachiso Samuel, Iklassov, Zangir, Ali, Kareem, Genadi, Rifo, Inui, Kentaro |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
The AI Data Scientist
by: Akimov, Farkhad, et al.
Published: (2025)
by: Akimov, Farkhad, et al.
Published: (2025)
Sycophancy Hides Linearly in the Attention Heads
by: Genadi, Rifo, et al.
Published: (2026)
by: Genadi, Rifo, et al.
Published: (2026)
RECALL: Library-Like Behavior In Language Models is Enhanced by Self-Referencing Causal Cycles
by: Nwadike, Munachiso, et al.
Published: (2025)
by: Nwadike, Munachiso, et al.
Published: (2025)
LLM-BABYBENCH: Understanding and Evaluating Grounded Planning and Reasoning in LLMs
by: Choukrani, Omar, et al.
Published: (2025)
by: Choukrani, Omar, et al.
Published: (2025)
Emergence of Primacy and Recency Effect in Mamba: A Mechanistic Point of View
by: Airlangga, Muhammad Cendekia, et al.
Published: (2025)
by: Airlangga, Muhammad Cendekia, et al.
Published: (2025)
Self-Guiding Exploration for Combinatorial Problems
by: Iklassov, Zangir, et al.
Published: (2024)
by: Iklassov, Zangir, et al.
Published: (2024)
Mechanistic Insights into Grokking from the Embedding Layer
by: AlquBoj, H. V., et al.
Published: (2025)
by: AlquBoj, H. V., et al.
Published: (2025)
Improving Personalisation in Valence and Arousal Prediction using Data Augmentation
by: Nwadike, Munachiso, et al.
Published: (2024)
by: Nwadike, Munachiso, et al.
Published: (2024)
Reinforcement Learning for Solving Stochastic Vehicle Routing Problem with Time Windows
by: Iklassov, Zangir, et al.
Published: (2024)
by: Iklassov, Zangir, et al.
Published: (2024)
SVRPBench: A Realistic Benchmark for Stochastic Vehicle Routing Problem
by: Heakl, Ahmed, et al.
Published: (2025)
by: Heakl, Ahmed, et al.
Published: (2025)
Culturally-Nuanced Story Generation for Reasoning in Low-Resource Languages: The Case of Javanese and Sundanese
by: Pranida, Salsabila Zahirah, et al.
Published: (2025)
by: Pranida, Salsabila Zahirah, et al.
Published: (2025)
Number Representations in LLMs: A Computational Parallel to Human Perception
by: AlquBoj, H. V., et al.
Published: (2025)
by: AlquBoj, H. V., et al.
Published: (2025)
LLMs Can Compensate for Deficiencies in Visual Representations
by: Takishita, Sho, et al.
Published: (2025)
by: Takishita, Sho, et al.
Published: (2025)
An Investigation of Warning Erroneous Chat Translations in Cross-lingual Communication
by: Li, Yunmeng, et al.
Published: (2024)
by: Li, Yunmeng, et al.
Published: (2024)
Measuring Reasoning Quality in LLMs: A Multi-Dimensional Behavioral Framework
by: Şenol, Ali, et al.
Published: (2026)
by: Şenol, Ali, et al.
Published: (2026)
Instruction-Guided Poetry Generation in Arabic and Its Dialects
by: Sadallah, Abdelrahman, et al.
Published: (2026)
by: Sadallah, Abdelrahman, et al.
Published: (2026)
ASR Under Noise: Exploring Robustness for Sundanese and Javanese
by: Pranida, Salsabila Zahirah, et al.
Published: (2025)
by: Pranida, Salsabila Zahirah, et al.
Published: (2025)
Abductive Reasoning with Syllogistic Forms in Large Language Models
by: Abe, Hirohiko, et al.
Published: (2026)
by: Abe, Hirohiko, et al.
Published: (2026)
Normative Reasoning in Large Language Models: A Comparative Benchmark from Logical and Modal Perspectives
by: Ozeki, Kentaro, et al.
Published: (2025)
by: Ozeki, Kentaro, et al.
Published: (2025)
Multilingual Idioms in Sentences and Conversations Across High-, Medium-, and Low-Resource Languages
by: Almheiri, Saeed, et al.
Published: (2026)
by: Almheiri, Saeed, et al.
Published: (2026)
GRACE: Discriminator-Guided Chain-of-Thought Reasoning
by: Khalifa, Muhammad, et al.
Published: (2023)
by: Khalifa, Muhammad, et al.
Published: (2023)
Measuring Iterative Temporal Reasoning with Time Puzzles
by: Wang, Zhengxiang, et al.
Published: (2026)
by: Wang, Zhengxiang, et al.
Published: (2026)
What Makes Good Multilingual Reasoning? Disentangling Reasoning Traces with Measurable Features
by: Ki, Dayeon, et al.
Published: (2026)
by: Ki, Dayeon, et al.
Published: (2026)
OckBench: Measuring the Efficiency of LLM Reasoning
by: Du, Zheng, et al.
Published: (2025)
by: Du, Zheng, et al.
Published: (2025)
Exploring Reasoning Biases in Large Language Models Through Syllogism: Insights from the NeuBAROCO Dataset
by: Ozeki, Kentaro, et al.
Published: (2024)
by: Ozeki, Kentaro, et al.
Published: (2024)
Recon: Reconstruction-Guided Reasoning Synthesis for User Modeling
by: Zhu, Alan, et al.
Published: (2026)
by: Zhu, Alan, et al.
Published: (2026)
Reward-Guided Speculative Decoding for Efficient LLM Reasoning
by: Liao, Baohao, et al.
Published: (2025)
by: Liao, Baohao, et al.
Published: (2025)
Feedback Forensics: A Toolkit to Measure AI Personality
by: Findeis, Arduin, et al.
Published: (2025)
by: Findeis, Arduin, et al.
Published: (2025)
Perspective Dial: Measuring Perspective of Text and Guiding LLM Outputs
by: Kim, Taejin, et al.
Published: (2025)
by: Kim, Taejin, et al.
Published: (2025)
RuleArena: A Benchmark for Rule-Guided Reasoning with LLMs in Real-World Scenarios
by: Zhou, Ruiwen, et al.
Published: (2024)
by: Zhou, Ruiwen, et al.
Published: (2024)
Certainty-Guided Reasoning in Large Language Models: A Dynamic Thinking Budget Approach
by: Nogueira, João Paulo, et al.
Published: (2025)
by: Nogueira, João Paulo, et al.
Published: (2025)
The Path of Least Resistance: Guiding LLM Reasoning Trajectories with Prefix Consensus
by: Jindal, Ishan, et al.
Published: (2026)
by: Jindal, Ishan, et al.
Published: (2026)
The Reasoning Lingua Franca: A Double-Edged Sword for Multilingual AI
by: Saji, Alan, et al.
Published: (2025)
by: Saji, Alan, et al.
Published: (2025)
RefTool: Reference-Guided Tool Creation for Knowledge-Intensive Reasoning
by: Liu, Xiao, et al.
Published: (2025)
by: Liu, Xiao, et al.
Published: (2025)
Guiding Clinical Reasoning with Large Language Models via Knowledge Seeds
by: WU, Jiageng, et al.
Published: (2024)
by: WU, Jiageng, et al.
Published: (2024)
Characterizing AlphaEarth Embedding Geometry for Agentic Environmental Reasoning
by: Rahman, Mashrekur, et al.
Published: (2026)
by: Rahman, Mashrekur, et al.
Published: (2026)
AI4Research: A Survey of Artificial Intelligence for Scientific Research
by: Chen, Qiguang, et al.
Published: (2025)
by: Chen, Qiguang, et al.
Published: (2025)
Is It Thinking or Cheating? Detecting Implicit Reward Hacking by Measuring Reasoning Effort
by: Wang, Xinpeng, et al.
Published: (2025)
by: Wang, Xinpeng, et al.
Published: (2025)
SciResearcher: Scaling Deep Research Agents for Frontier Scientific Reasoning
by: Zheng, Tianshi, et al.
Published: (2026)
by: Zheng, Tianshi, et al.
Published: (2026)
LLM Rationalis? Measuring Bargaining Capabilities of AI Negotiators
by: Shah, Cheril, et al.
Published: (2025)
by: Shah, Cheril, et al.
Published: (2025)
Similar Items
-
The AI Data Scientist
by: Akimov, Farkhad, et al.
Published: (2025) -
Sycophancy Hides Linearly in the Attention Heads
by: Genadi, Rifo, et al.
Published: (2026) -
RECALL: Library-Like Behavior In Language Models is Enhanced by Self-Referencing Causal Cycles
by: Nwadike, Munachiso, et al.
Published: (2025) -
LLM-BABYBENCH: Understanding and Evaluating Grounded Planning and Reasoning in LLMs
by: Choukrani, Omar, et al.
Published: (2025) -
Emergence of Primacy and Recency Effect in Mamba: A Mechanistic Point of View
by: Airlangga, Muhammad Cendekia, et al.
Published: (2025)