Illusions of reflection: open-ended task reveals systematic failures in Large Language Models' reflective reasoning
Fuente:
arXiv
Saved in:
| Main Authors: | Weatherhead, Sion, Salim, Flora, Belbasis, Aaron |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Impute With Confidence: A Framework for Uncertainty Aware Multivariate Time Series Imputation
by: Weatherhead, Addison, et al.
Published: (2025)
by: Weatherhead, Addison, et al.
Published: (2025)
SOCIA-Nabla: Textual Gradient Meets Multi-Agent Orchestration for Automated Simulator Generation
by: Hua, Yuncheng, et al.
Published: (2025)
by: Hua, Yuncheng, et al.
Published: (2025)
SOCIA-$\nabla$: Textual Gradient Meets Multi-Agent Orchestration for Automated Simulator Generation
by: Hua, Yuncheng, et al.
Published: (2025)
by: Hua, Yuncheng, et al.
Published: (2025)
SOCIA-EVO: Automated Simulator Construction via Dual-Anchored Bi-Level Optimization
by: Hua, Yuncheng, et al.
Published: (2026)
by: Hua, Yuncheng, et al.
Published: (2026)
Large Language Models for Next Point-of-Interest Recommendation
by: Li, Peibo, et al.
Published: (2024)
by: Li, Peibo, et al.
Published: (2024)
Language models show human-like content effects on reasoning tasks
by: Dasgupta, Ishita, et al.
Published: (2022)
by: Dasgupta, Ishita, et al.
Published: (2022)
Cross-model Fairness: Empirical Study of Fairness and Ethics Under Model Multiplicity
by: Sokol, Kacper, et al.
Published: (2022)
by: Sokol, Kacper, et al.
Published: (2022)
Are UFOs Driving Innovation? The Illusion of Causality in Large Language Models
by: Carro, María Victoria, et al.
Published: (2024)
by: Carro, María Victoria, et al.
Published: (2024)
ASTER: Adaptive Spatio-Temporal Early Decision Model for Dynamic Resource Allocation
by: Chen, Shulun, et al.
Published: (2025)
by: Chen, Shulun, et al.
Published: (2025)
Consolidating TinyML Lifecycle with Large Language Models: Reality, Illusion, or Opportunity?
by: Wu, Guanghan, et al.
Published: (2025)
by: Wu, Guanghan, et al.
Published: (2025)
T-JEPA: A Joint-Embedding Predictive Architecture for Trajectory Similarity Computation
by: Li, Lihuan, et al.
Published: (2024)
by: Li, Lihuan, et al.
Published: (2024)
Causes in neuron diagrams, and testing causal reasoning in Large Language Models. A glimpse of the future of philosophy?
by: Vervoort, Louis, et al.
Published: (2025)
by: Vervoort, Louis, et al.
Published: (2025)
Refine-POI: Reinforcement Fine-Tuned Large Language Models for Next Point-of-Interest Recommendation
by: Li, Peibo, et al.
Published: (2025)
by: Li, Peibo, et al.
Published: (2025)
Critique of Impure Reason: Unveiling the reasoning behaviour of medical Large Language Models
by: Sim, Shamus, et al.
Published: (2024)
by: Sim, Shamus, et al.
Published: (2024)
Are Flat Minima an Illusion?
by: Bennett, Michael Timothy
Published: (2026)
by: Bennett, Michael Timothy
Published: (2026)
LLMCarbon: Modeling the end-to-end Carbon Footprint of Large Language Models
by: Faiz, Ahmad, et al.
Published: (2023)
by: Faiz, Ahmad, et al.
Published: (2023)
On-device System of Compositional Multi-tasking in Large Language Models
by: Bohdal, Ondrej, et al.
Published: (2025)
by: Bohdal, Ondrej, et al.
Published: (2025)
Efficient Compositional Multi-tasking for On-device Large Language Models
by: Bohdal, Ondrej, et al.
Published: (2025)
by: Bohdal, Ondrej, et al.
Published: (2025)
CGM-JEPA: Learning Consistent Continuous Glucose Monitor Representations via Predictive Self-Supervised Pretraining
by: Muhammad, Hada Melino, et al.
Published: (2026)
by: Muhammad, Hada Melino, et al.
Published: (2026)
XXLTraffic: Expanding and Extremely Long Traffic forecasting beyond test adaptation
by: Yin, Du, et al.
Published: (2024)
by: Yin, Du, et al.
Published: (2024)
ScheduleFree+: Scaling Learning-Rate-Free & Schedule-Free Learning to Large Language Models
by: Defazio, Aaron
Published: (2026)
by: Defazio, Aaron
Published: (2026)
A-UTE: Advection Informed, Uncertainty Aware Temperature Emulator
by: Saleem, Hira, et al.
Published: (2024)
by: Saleem, Hira, et al.
Published: (2024)
QuestBench: Can LLMs ask the right question to acquire information in reasoning tasks?
by: Li, Belinda Z., et al.
Published: (2025)
by: Li, Belinda Z., et al.
Published: (2025)
Performance of AI agents based on reasoning language models on ALD process optimization tasks
by: Yanguas-Gil, Angel
Published: (2026)
by: Yanguas-Gil, Angel
Published: (2026)
Long-term Fairness in Ride-Hailing Platform
by: Kang, Yufan, et al.
Published: (2024)
by: Kang, Yufan, et al.
Published: (2024)
T-REX: Mixture-of-Rank-One-Experts with Semantic-aware Intuition for Multi-task Large Language Model Finetuning
by: Zhang, Rongyu, et al.
Published: (2024)
by: Zhang, Rongyu, et al.
Published: (2024)
Towards Generalizable Human Activity Recognition: A Survey
by: Cai, Yize, et al.
Published: (2025)
by: Cai, Yize, et al.
Published: (2025)
A$^2$-LLM: An End-to-end Conversational Audio Avatar Large Language Model
by: Hu, Xiaolin, et al.
Published: (2026)
by: Hu, Xiaolin, et al.
Published: (2026)
BoostStep: Boosting mathematical capability of Large Language Models via improved single-step reasoning
by: Zhang, Beichen, et al.
Published: (2025)
by: Zhang, Beichen, et al.
Published: (2025)
AKD : Adversarial Knowledge Distillation For Large Language Models Alignment on Coding tasks
by: Oulkadda, Ilyas, et al.
Published: (2025)
by: Oulkadda, Ilyas, et al.
Published: (2025)
EIDOS: Latent-Space Predictive Learning for Time Series Foundation Models
by: Zhou, Xinxing, et al.
Published: (2026)
by: Zhou, Xinxing, et al.
Published: (2026)
AdaNODEs: Test Time Adaptation for Time Series Forecasting Using Neural ODEs
by: Dang, Ting, et al.
Published: (2026)
by: Dang, Ting, et al.
Published: (2026)
Nested Learning: The Illusion of Deep Learning Architectures
by: Behrouz, Ali, et al.
Published: (2025)
by: Behrouz, Ali, et al.
Published: (2025)
Inference Time Context Sparsity: Illusion or Opportunity?
by: Joshi, Sahil, et al.
Published: (2026)
by: Joshi, Sahil, et al.
Published: (2026)
Zero-shot Load Forecasting for Integrated Energy Systems: A Large Language Model-based Framework with Multi-task Learning
by: Li, Jiaheng, et al.
Published: (2025)
by: Li, Jiaheng, et al.
Published: (2025)
The Illusion of Specialization: Unveiling the Domain-Invariant "Standing Committee" in Mixture-of-Experts Models
by: Wang, Yan, et al.
Published: (2026)
by: Wang, Yan, et al.
Published: (2026)
UrbanVerse: Learning Urban Region Representation Across Cities and Tasks
by: Sun, Fengze, et al.
Published: (2026)
by: Sun, Fengze, et al.
Published: (2026)
Enhancing Chemical Reaction and Retrosynthesis Prediction with Large Language Model and Dual-task Learning
by: Lin, Xuan, et al.
Published: (2025)
by: Lin, Xuan, et al.
Published: (2025)
MetaTool: Facilitating Large Language Models to Master Tools with Meta-task Augmentation
by: Wang, Xiaohan, et al.
Published: (2024)
by: Wang, Xiaohan, et al.
Published: (2024)
Long-horizon Visual Instruction Generation with Logic and Attribute Self-reflection
by: Suo, Yucheng, et al.
Published: (2025)
by: Suo, Yucheng, et al.
Published: (2025)
Similar Items
-
Impute With Confidence: A Framework for Uncertainty Aware Multivariate Time Series Imputation
by: Weatherhead, Addison, et al.
Published: (2025) -
SOCIA-Nabla: Textual Gradient Meets Multi-Agent Orchestration for Automated Simulator Generation
by: Hua, Yuncheng, et al.
Published: (2025) -
SOCIA-$\nabla$: Textual Gradient Meets Multi-Agent Orchestration for Automated Simulator Generation
by: Hua, Yuncheng, et al.
Published: (2025) -
SOCIA-EVO: Automated Simulator Construction via Dual-Anchored Bi-Level Optimization
by: Hua, Yuncheng, et al.
Published: (2026) -
Large Language Models for Next Point-of-Interest Recommendation
by: Li, Peibo, et al.
Published: (2024)