Illusions of reflection: open-ended task reveals systematic failures in Large Language Models' reflective reasoning
Fuente:
arXiv
Salvato in:
| Autori principali: | Weatherhead, Sion, Salim, Flora, Belbasis, Aaron |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Impute With Confidence: A Framework for Uncertainty Aware Multivariate Time Series Imputation
di: Weatherhead, Addison, et al.
Pubblicazione: (2025)
di: Weatherhead, Addison, et al.
Pubblicazione: (2025)
SOCIA-Nabla: Textual Gradient Meets Multi-Agent Orchestration for Automated Simulator Generation
di: Hua, Yuncheng, et al.
Pubblicazione: (2025)
di: Hua, Yuncheng, et al.
Pubblicazione: (2025)
SOCIA-$\nabla$: Textual Gradient Meets Multi-Agent Orchestration for Automated Simulator Generation
di: Hua, Yuncheng, et al.
Pubblicazione: (2025)
di: Hua, Yuncheng, et al.
Pubblicazione: (2025)
SOCIA-EVO: Automated Simulator Construction via Dual-Anchored Bi-Level Optimization
di: Hua, Yuncheng, et al.
Pubblicazione: (2026)
di: Hua, Yuncheng, et al.
Pubblicazione: (2026)
Large Language Models for Next Point-of-Interest Recommendation
di: Li, Peibo, et al.
Pubblicazione: (2024)
di: Li, Peibo, et al.
Pubblicazione: (2024)
Language models show human-like content effects on reasoning tasks
di: Dasgupta, Ishita, et al.
Pubblicazione: (2022)
di: Dasgupta, Ishita, et al.
Pubblicazione: (2022)
Cross-model Fairness: Empirical Study of Fairness and Ethics Under Model Multiplicity
di: Sokol, Kacper, et al.
Pubblicazione: (2022)
di: Sokol, Kacper, et al.
Pubblicazione: (2022)
Are UFOs Driving Innovation? The Illusion of Causality in Large Language Models
di: Carro, María Victoria, et al.
Pubblicazione: (2024)
di: Carro, María Victoria, et al.
Pubblicazione: (2024)
ASTER: Adaptive Spatio-Temporal Early Decision Model for Dynamic Resource Allocation
di: Chen, Shulun, et al.
Pubblicazione: (2025)
di: Chen, Shulun, et al.
Pubblicazione: (2025)
Consolidating TinyML Lifecycle with Large Language Models: Reality, Illusion, or Opportunity?
di: Wu, Guanghan, et al.
Pubblicazione: (2025)
di: Wu, Guanghan, et al.
Pubblicazione: (2025)
T-JEPA: A Joint-Embedding Predictive Architecture for Trajectory Similarity Computation
di: Li, Lihuan, et al.
Pubblicazione: (2024)
di: Li, Lihuan, et al.
Pubblicazione: (2024)
Causes in neuron diagrams, and testing causal reasoning in Large Language Models. A glimpse of the future of philosophy?
di: Vervoort, Louis, et al.
Pubblicazione: (2025)
di: Vervoort, Louis, et al.
Pubblicazione: (2025)
Refine-POI: Reinforcement Fine-Tuned Large Language Models for Next Point-of-Interest Recommendation
di: Li, Peibo, et al.
Pubblicazione: (2025)
di: Li, Peibo, et al.
Pubblicazione: (2025)
Critique of Impure Reason: Unveiling the reasoning behaviour of medical Large Language Models
di: Sim, Shamus, et al.
Pubblicazione: (2024)
di: Sim, Shamus, et al.
Pubblicazione: (2024)
Are Flat Minima an Illusion?
di: Bennett, Michael Timothy
Pubblicazione: (2026)
di: Bennett, Michael Timothy
Pubblicazione: (2026)
LLMCarbon: Modeling the end-to-end Carbon Footprint of Large Language Models
di: Faiz, Ahmad, et al.
Pubblicazione: (2023)
di: Faiz, Ahmad, et al.
Pubblicazione: (2023)
On-device System of Compositional Multi-tasking in Large Language Models
di: Bohdal, Ondrej, et al.
Pubblicazione: (2025)
di: Bohdal, Ondrej, et al.
Pubblicazione: (2025)
Efficient Compositional Multi-tasking for On-device Large Language Models
di: Bohdal, Ondrej, et al.
Pubblicazione: (2025)
di: Bohdal, Ondrej, et al.
Pubblicazione: (2025)
CGM-JEPA: Learning Consistent Continuous Glucose Monitor Representations via Predictive Self-Supervised Pretraining
di: Muhammad, Hada Melino, et al.
Pubblicazione: (2026)
di: Muhammad, Hada Melino, et al.
Pubblicazione: (2026)
XXLTraffic: Expanding and Extremely Long Traffic forecasting beyond test adaptation
di: Yin, Du, et al.
Pubblicazione: (2024)
di: Yin, Du, et al.
Pubblicazione: (2024)
ScheduleFree+: Scaling Learning-Rate-Free & Schedule-Free Learning to Large Language Models
di: Defazio, Aaron
Pubblicazione: (2026)
di: Defazio, Aaron
Pubblicazione: (2026)
A-UTE: Advection Informed, Uncertainty Aware Temperature Emulator
di: Saleem, Hira, et al.
Pubblicazione: (2024)
di: Saleem, Hira, et al.
Pubblicazione: (2024)
QuestBench: Can LLMs ask the right question to acquire information in reasoning tasks?
di: Li, Belinda Z., et al.
Pubblicazione: (2025)
di: Li, Belinda Z., et al.
Pubblicazione: (2025)
Performance of AI agents based on reasoning language models on ALD process optimization tasks
di: Yanguas-Gil, Angel
Pubblicazione: (2026)
di: Yanguas-Gil, Angel
Pubblicazione: (2026)
Long-term Fairness in Ride-Hailing Platform
di: Kang, Yufan, et al.
Pubblicazione: (2024)
di: Kang, Yufan, et al.
Pubblicazione: (2024)
T-REX: Mixture-of-Rank-One-Experts with Semantic-aware Intuition for Multi-task Large Language Model Finetuning
di: Zhang, Rongyu, et al.
Pubblicazione: (2024)
di: Zhang, Rongyu, et al.
Pubblicazione: (2024)
Towards Generalizable Human Activity Recognition: A Survey
di: Cai, Yize, et al.
Pubblicazione: (2025)
di: Cai, Yize, et al.
Pubblicazione: (2025)
A$^2$-LLM: An End-to-end Conversational Audio Avatar Large Language Model
di: Hu, Xiaolin, et al.
Pubblicazione: (2026)
di: Hu, Xiaolin, et al.
Pubblicazione: (2026)
BoostStep: Boosting mathematical capability of Large Language Models via improved single-step reasoning
di: Zhang, Beichen, et al.
Pubblicazione: (2025)
di: Zhang, Beichen, et al.
Pubblicazione: (2025)
AKD : Adversarial Knowledge Distillation For Large Language Models Alignment on Coding tasks
di: Oulkadda, Ilyas, et al.
Pubblicazione: (2025)
di: Oulkadda, Ilyas, et al.
Pubblicazione: (2025)
EIDOS: Latent-Space Predictive Learning for Time Series Foundation Models
di: Zhou, Xinxing, et al.
Pubblicazione: (2026)
di: Zhou, Xinxing, et al.
Pubblicazione: (2026)
AdaNODEs: Test Time Adaptation for Time Series Forecasting Using Neural ODEs
di: Dang, Ting, et al.
Pubblicazione: (2026)
di: Dang, Ting, et al.
Pubblicazione: (2026)
Nested Learning: The Illusion of Deep Learning Architectures
di: Behrouz, Ali, et al.
Pubblicazione: (2025)
di: Behrouz, Ali, et al.
Pubblicazione: (2025)
Inference Time Context Sparsity: Illusion or Opportunity?
di: Joshi, Sahil, et al.
Pubblicazione: (2026)
di: Joshi, Sahil, et al.
Pubblicazione: (2026)
Zero-shot Load Forecasting for Integrated Energy Systems: A Large Language Model-based Framework with Multi-task Learning
di: Li, Jiaheng, et al.
Pubblicazione: (2025)
di: Li, Jiaheng, et al.
Pubblicazione: (2025)
The Illusion of Specialization: Unveiling the Domain-Invariant "Standing Committee" in Mixture-of-Experts Models
di: Wang, Yan, et al.
Pubblicazione: (2026)
di: Wang, Yan, et al.
Pubblicazione: (2026)
UrbanVerse: Learning Urban Region Representation Across Cities and Tasks
di: Sun, Fengze, et al.
Pubblicazione: (2026)
di: Sun, Fengze, et al.
Pubblicazione: (2026)
Enhancing Chemical Reaction and Retrosynthesis Prediction with Large Language Model and Dual-task Learning
di: Lin, Xuan, et al.
Pubblicazione: (2025)
di: Lin, Xuan, et al.
Pubblicazione: (2025)
MetaTool: Facilitating Large Language Models to Master Tools with Meta-task Augmentation
di: Wang, Xiaohan, et al.
Pubblicazione: (2024)
di: Wang, Xiaohan, et al.
Pubblicazione: (2024)
Long-horizon Visual Instruction Generation with Logic and Attribute Self-reflection
di: Suo, Yucheng, et al.
Pubblicazione: (2025)
di: Suo, Yucheng, et al.
Pubblicazione: (2025)
Documenti analoghi
-
Impute With Confidence: A Framework for Uncertainty Aware Multivariate Time Series Imputation
di: Weatherhead, Addison, et al.
Pubblicazione: (2025) -
SOCIA-Nabla: Textual Gradient Meets Multi-Agent Orchestration for Automated Simulator Generation
di: Hua, Yuncheng, et al.
Pubblicazione: (2025) -
SOCIA-$\nabla$: Textual Gradient Meets Multi-Agent Orchestration for Automated Simulator Generation
di: Hua, Yuncheng, et al.
Pubblicazione: (2025) -
SOCIA-EVO: Automated Simulator Construction via Dual-Anchored Bi-Level Optimization
di: Hua, Yuncheng, et al.
Pubblicazione: (2026) -
Large Language Models for Next Point-of-Interest Recommendation
di: Li, Peibo, et al.
Pubblicazione: (2024)