Understanding LLM Failures: A Multi-Tape Turing Machine Analysis of Systematic Errors in Language Model Reasoning
Fuente:
arXiv
Guardado en:
| Autor principal: | Boman, Magnus |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
On the Compatibility of Generative AI and Generative Linguistics
por: Portelance, Eva, et al.
Publicado: (2024)
por: Portelance, Eva, et al.
Publicado: (2024)
LlamBERT: Large-scale low-cost data annotation in NLP
por: Csanády, Bálint, et al.
Publicado: (2024)
por: Csanády, Bálint, et al.
Publicado: (2024)
GeoGalactica: A Scientific Large Language Model in Geoscience
por: Lin, Zhouhan, et al.
Publicado: (2023)
por: Lin, Zhouhan, et al.
Publicado: (2023)
Understanding Syllogistic Reasoning in LLMs from Formal and Natural Language Perspectives
por: Poddar, Aheli, et al.
Publicado: (2025)
por: Poddar, Aheli, et al.
Publicado: (2025)
LFC-DA: Logical Formula-Controlled Data Augmentation for Enhanced Logical Reasoning
por: Li, Shenghao
Publicado: (2025)
por: Li, Shenghao
Publicado: (2025)
Automated Theorem Provers Help Improve Large Language Model Reasoning
por: McGinness, Lachlan, et al.
Publicado: (2024)
por: McGinness, Lachlan, et al.
Publicado: (2024)
Beyond End-to-End Video Models: An LLM-Based Multi-Agent System for Educational Video Generation
por: Yan, Lingyong, et al.
Publicado: (2026)
por: Yan, Lingyong, et al.
Publicado: (2026)
Multi-Hierarchical Feature Detection for Large Language Model Generated Text
por: Zhang, Luyan, et al.
Publicado: (2025)
por: Zhang, Luyan, et al.
Publicado: (2025)
ReTreVal: Reasoning Tree with Validation -- A Hybrid Framework for Enhanced LLM Multi-Step Reasoning
por: HS, Abhishek, et al.
Publicado: (2026)
por: HS, Abhishek, et al.
Publicado: (2026)
Automating the Analysis of Parsing Algorithms (and other Dynamic Programs)
por: Vieira, Tim, et al.
Publicado: (2025)
por: Vieira, Tim, et al.
Publicado: (2025)
From Helpfulness to Toxic Proactivity: Diagnosing Behavioral Misalignment in LLM Agents
por: Wang, Xinyue, et al.
Publicado: (2026)
por: Wang, Xinyue, et al.
Publicado: (2026)
HalluScan: A Systematic Benchmark for Detecting and Mitigating Hallucinations in Instruction-Following LLMs
por: Cherif, Ahmed
Publicado: (2026)
por: Cherif, Ahmed
Publicado: (2026)
An Industrial-Scale Insurance LLM Achieving Verifiable Domain Mastery and Hallucination Control without Competence Trade-offs
por: Zhu, Qian, et al.
Publicado: (2026)
por: Zhu, Qian, et al.
Publicado: (2026)
From Classification to Ranking: Enhancing LLM Reasoning Capabilities for MBTI Personality Detection
por: Cao, Yuan, et al.
Publicado: (2026)
por: Cao, Yuan, et al.
Publicado: (2026)
LLM-Rubric: A Multidimensional, Calibrated Approach to Automated Evaluation of Natural Language Texts
por: Hashemi, Helia, et al.
Publicado: (2024)
por: Hashemi, Helia, et al.
Publicado: (2024)
Large Language Models as Software Components: A Taxonomy for LLM-Integrated Applications
por: Weber, Irene
Publicado: (2024)
por: Weber, Irene
Publicado: (2024)
Evaluating the efficacy of LLM Safety Solutions : The Palit Benchmark Dataset
por: Palit, Sayon, et al.
Publicado: (2025)
por: Palit, Sayon, et al.
Publicado: (2025)
STATe-of-Thoughts: Structured Action Templates for Tree-of-Thoughts
por: Bamberger, Zachary, et al.
Publicado: (2026)
por: Bamberger, Zachary, et al.
Publicado: (2026)
Vision-Language Reasoning for Geolocalization: A Reinforcement Learning Approach
por: Wu, Biao, et al.
Publicado: (2026)
por: Wu, Biao, et al.
Publicado: (2026)
Cost-Aware Model Selection for Text Classification: Multi-Objective Trade-offs Between Fine-Tuned Encoders and LLM Prompting in Production
por: Gonzalez, Alberto Andres Valdes
Publicado: (2026)
por: Gonzalez, Alberto Andres Valdes
Publicado: (2026)
Exploring and Mitigating Gender Bias in Encoder-Based Transformer Models
por: Hossain, Ariyan, et al.
Publicado: (2025)
por: Hossain, Ariyan, et al.
Publicado: (2025)
Graphemic Normalization of the Perso-Arabic Script
por: Doctor, Raiomond, et al.
Publicado: (2022)
por: Doctor, Raiomond, et al.
Publicado: (2022)
Beyond Arabic: Software for Perso-Arabic Script Manipulation
por: Gutkin, Alexander, et al.
Publicado: (2023)
por: Gutkin, Alexander, et al.
Publicado: (2023)
Text Generation: A Systematic Literature Review of Tasks, Evaluation, and Challenges
por: Becker, Jonas, et al.
Publicado: (2024)
por: Becker, Jonas, et al.
Publicado: (2024)
Automated MCQA Benchmarking at Scale: Evaluating Reasoning Traces as Retrieval Sources for Domain Adaptation of Small Language Models
por: Gokdemir, Ozan, et al.
Publicado: (2025)
por: Gokdemir, Ozan, et al.
Publicado: (2025)
Evaluating Large Language Models on Historical Health Crisis Knowledge in Resource-Limited Settings: A Hybrid Multi-Metric Study
por: Hasan, Mohammed Rakibul
Publicado: (2026)
por: Hasan, Mohammed Rakibul
Publicado: (2026)
Exploring Design of Multi-Agent LLM Dialogues for Research Ideation
por: Ueda, Keisuke, et al.
Publicado: (2025)
por: Ueda, Keisuke, et al.
Publicado: (2025)
Project Riley: Multimodal Multi-Agent LLM Collaboration with Emotional Reasoning and Voting
por: Ortigoso, Ana Rita, et al.
Publicado: (2025)
por: Ortigoso, Ana Rita, et al.
Publicado: (2025)
SERC: LDPC-Inspired Semantic Error Correction for Retrieval-Augmented Generation
por: Kim, Gyumin, et al.
Publicado: (2026)
por: Kim, Gyumin, et al.
Publicado: (2026)
Towards Conditioning Clinical Text Generation for User Control
por: Koraş, Osman Alperen, et al.
Publicado: (2025)
por: Koraş, Osman Alperen, et al.
Publicado: (2025)
Meta-Evaluation of Translation Evaluation Methods: a systematic up-to-date overview
por: Han, Lifeng, et al.
Publicado: (2016)
por: Han, Lifeng, et al.
Publicado: (2016)
Efficient Strategy for Improving Large Language Model (LLM) Capabilities
por: Gutiérrez, Julián Camilo Velandia
Publicado: (2025)
por: Gutiérrez, Julián Camilo Velandia
Publicado: (2025)
Stay Focused: Problem Drift in Multi-Agent Debate
por: Becker, Jonas, et al.
Publicado: (2025)
por: Becker, Jonas, et al.
Publicado: (2025)
Robust Reward Modeling for Large Language Models via Causal Decomposition
por: Lu, Yunsheng, et al.
Publicado: (2026)
por: Lu, Yunsheng, et al.
Publicado: (2026)
Do LLMs Game Formalization? Evaluating Faithfulness in Logical Reasoning
por: Kim, Kyuhee, et al.
Publicado: (2026)
por: Kim, Kyuhee, et al.
Publicado: (2026)
BLT: Can Large Language Models Handle Basic Legal Text?
por: Blair-Stanek, Andrew, et al.
Publicado: (2023)
por: Blair-Stanek, Andrew, et al.
Publicado: (2023)
Computational Economics in Large Language Models: Exploring Model Behavior and Incentive Design under Resource Constraints
por: Reddy, Sandeep, et al.
Publicado: (2025)
por: Reddy, Sandeep, et al.
Publicado: (2025)
Blocks Architecture (BloArk): Efficient, Cost-Effective, and Incremental Dataset Architecture for Wikipedia Revision History
por: Li, Lingxi, et al.
Publicado: (2024)
por: Li, Lingxi, et al.
Publicado: (2024)
Does Editing Provide Evidence for Localization?
por: Wang, Zihao, et al.
Publicado: (2025)
por: Wang, Zihao, et al.
Publicado: (2025)
Advances in Pre-trained Language Models for Domain-Specific Text Classification: A Systematic Review
por: Rostam, Zhyar Rzgar K., et al.
Publicado: (2025)
por: Rostam, Zhyar Rzgar K., et al.
Publicado: (2025)
Ejemplares similares
-
On the Compatibility of Generative AI and Generative Linguistics
por: Portelance, Eva, et al.
Publicado: (2024) -
LlamBERT: Large-scale low-cost data annotation in NLP
por: Csanády, Bálint, et al.
Publicado: (2024) -
GeoGalactica: A Scientific Large Language Model in Geoscience
por: Lin, Zhouhan, et al.
Publicado: (2023) -
Understanding Syllogistic Reasoning in LLMs from Formal and Natural Language Perspectives
por: Poddar, Aheli, et al.
Publicado: (2025) -
LFC-DA: Logical Formula-Controlled Data Augmentation for Enhanced Logical Reasoning
por: Li, Shenghao
Publicado: (2025)