The Mirror Loop: Recursive Non-Convergence in Generative Reasoning Systems
Fuente:
arXiv
Saved in:
| Main Author: | DeVilling, Bentley |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
SafeRL-Lite: A Lightweight, Explainable, and Constrained Reinforcement Learning Library
by: Mishra, Satyam, et al.
Published: (2025)
by: Mishra, Satyam, et al.
Published: (2025)
Foundation Models as World Models: A Foundational Study in Text-Based GridWorlds
by: Sasso, Remo, et al.
Published: (2025)
by: Sasso, Remo, et al.
Published: (2025)
Exploration with Foundation Models: Capabilities, Limitations, and Hybrid Approaches
by: Sasso, Remo, et al.
Published: (2025)
by: Sasso, Remo, et al.
Published: (2025)
GoldenStart: Q-Guided Priors and Entropy Control for Distilling Flow Policies
by: Zhang, He, et al.
Published: (2026)
by: Zhang, He, et al.
Published: (2026)
An Improved Adaptive PID Optimizer with Enhanced Convergence and Stability for Deep Learning
by: Saini, Saurabh, et al.
Published: (2026)
by: Saini, Saurabh, et al.
Published: (2026)
Evolving machine learning workflows through interactive AutoML
by: Barbudo, Rafael, et al.
Published: (2024)
by: Barbudo, Rafael, et al.
Published: (2024)
Think-at-Hard: Selective Latent Iterations to Improve Reasoning Language Models
by: Fu, Tianyu, et al.
Published: (2025)
by: Fu, Tianyu, et al.
Published: (2025)
SMOSE: Sparse Mixture of Shallow Experts for Interpretable Reinforcement Learning in Continuous Control Tasks
by: Vincze, Mátyás, et al.
Published: (2024)
by: Vincze, Mátyás, et al.
Published: (2024)
Active Causal Experimentalist (ACE): Learning Intervention Strategies via Direct Preference Optimization
by: Cooper, Patrick, et al.
Published: (2026)
by: Cooper, Patrick, et al.
Published: (2026)
Streaming Continual Learning for Unified Adaptive Intelligence in Dynamic Environments
by: Giannini, Federico, et al.
Published: (2026)
by: Giannini, Federico, et al.
Published: (2026)
AI Agents: Evolution, Architecture, and Real-World Applications
by: Krishnan, Naveen
Published: (2025)
by: Krishnan, Naveen
Published: (2025)
Adaptive Minds: Empowering Agents with LoRA-as-Tools
by: Shekar, Pavan C, et al.
Published: (2025)
by: Shekar, Pavan C, et al.
Published: (2025)
Efficient Contextual Preferential Bayesian Optimization with Historical Examples
by: Khan, Farha A., et al.
Published: (2022)
by: Khan, Farha A., et al.
Published: (2022)
SALE-Based Offline Reinforcement Learning with Ensemble Q-Networks
by: Chun, Zheng
Published: (2025)
by: Chun, Zheng
Published: (2025)
Can a Bayesian Oracle Prevent Harm from an Agent?
by: Bengio, Yoshua, et al.
Published: (2024)
by: Bengio, Yoshua, et al.
Published: (2024)
Diffusion-MPC in Discrete Domains: Feasibility Constraints, Horizon Effects, and Critic Alignment: Case study with Tetris
by: Wang, Haochuan Kevin
Published: (2026)
by: Wang, Haochuan Kevin
Published: (2026)
On Divergence Measures for Training GFlowNets
by: da Silva, Tiago, et al.
Published: (2024)
by: da Silva, Tiago, et al.
Published: (2024)
Modeling and Controlling Deployment Reliability under Temporal Distribution Shift
by: Rahman, Naimur, et al.
Published: (2026)
by: Rahman, Naimur, et al.
Published: (2026)
From Imitation to Interaction: Mastering Game of Schnapsen with Shallow Reinforcement Learning
by: Klačan, Ján, et al.
Published: (2026)
by: Klačan, Ján, et al.
Published: (2026)
Pioneer Agent: Continual Improvement of Small Language Models in Production
by: Atreja, Dhruv, et al.
Published: (2026)
by: Atreja, Dhruv, et al.
Published: (2026)
Prediction-Based Markov Violation Scores for Detecting Non-Markovian Observations in Reinforcement Learning
by: Mysore, Naveen
Published: (2026)
by: Mysore, Naveen
Published: (2026)
Task Memory Engine (TME): Enhancing State Awareness for Multi-Step LLM Agent Tasks
by: Ye, Ye
Published: (2025)
by: Ye, Ye
Published: (2025)
AI and Machine Learning Approaches for Predicting Nanoparticles Toxicity The Critical Role of Physiochemical Properties
by: Yousaf, Iqra
Published: (2024)
by: Yousaf, Iqra
Published: (2024)
Differentiable Symbolic Planning: A Neural Architecture for Constraint Reasoning with Learned Feasibility
by: Oruganti, Venkatakrishna Reddy
Published: (2026)
by: Oruganti, Venkatakrishna Reddy
Published: (2026)
SuS: Strategy-aware Surprise for Intrinsic Exploration
by: Kashirskiy, Mark, et al.
Published: (2026)
by: Kashirskiy, Mark, et al.
Published: (2026)
Emotion-Inspired Learning Signals (EILS): A Homeostatic Framework for Adaptive Autonomous Agents
by: Tiwari, Dhruv
Published: (2025)
by: Tiwari, Dhruv
Published: (2025)
A Comparison Between Decision Transformers and Traditional Offline Reinforcement Learning Algorithms
by: Caunhye, Ali Murtaza, et al.
Published: (2025)
by: Caunhye, Ali Murtaza, et al.
Published: (2025)
Vis-CoT: A Human-in-the-Loop Framework for Interactive Visualization and Intervention in LLM Chain-of-Thought Reasoning
by: Pather, Kaviraj, et al.
Published: (2025)
by: Pather, Kaviraj, et al.
Published: (2025)
XAutoLM: Efficient Fine-Tuning of Language Models via Meta-Learning and AutoML
by: Estevanell-Valladares, Ernesto L., et al.
Published: (2025)
by: Estevanell-Valladares, Ernesto L., et al.
Published: (2025)
KL-Regularised Q-Learning: A Token-level Action-Value perspective on Online RLHF
by: Brown, Jason R, et al.
Published: (2025)
by: Brown, Jason R, et al.
Published: (2025)
Statistical Guarantees for Lifelong Reinforcement Learning using PAC-Bayes Theory
by: Zhang, Zhi, et al.
Published: (2024)
by: Zhang, Zhi, et al.
Published: (2024)
On the Generalization Gap in LLM Planning: Tests and Verifier-Reward RL
by: Belcamino, Valerio, et al.
Published: (2026)
by: Belcamino, Valerio, et al.
Published: (2026)
Improving Hyperparameter Optimization with Checkpointed Model Weights
by: Mehta, Nikhil, et al.
Published: (2024)
by: Mehta, Nikhil, et al.
Published: (2024)
GIRL: Generative Imagination Reinforcement Learning via Information-Theoretic Hallucination Control
by: Hiremath, Prakul Sunil
Published: (2026)
by: Hiremath, Prakul Sunil
Published: (2026)
ASkDAgger: Active Skill-level Data Aggregation for Interactive Imitation Learning
by: Luijkx, Jelle, et al.
Published: (2025)
by: Luijkx, Jelle, et al.
Published: (2025)
CogniLoad: A Synthetic Natural Language Reasoning Benchmark With Tunable Length, Intrinsic Difficulty, and Distractor Density
by: Kaiser, Daniel, et al.
Published: (2025)
by: Kaiser, Daniel, et al.
Published: (2025)
A Structural Threshold in Decision Capacity Governs Collapse in Self-Play Reinforcement Learning
by: Kujur, Arahan
Published: (2026)
by: Kujur, Arahan
Published: (2026)
A New Modeling to Feature Selection Based on the Fuzzy Rough Set Theory in Normal and Optimistic States on Hybrid Information Systems
by: Safarpour, Mohammad Hossein, et al.
Published: (2026)
by: Safarpour, Mohammad Hossein, et al.
Published: (2026)
Self-Supervised On-Policy Reinforcement Learning via Contrastive Proximal Policy Optimisation
by: Osman, Asim, et al.
Published: (2026)
by: Osman, Asim, et al.
Published: (2026)
Maximizing Rollout Informativeness under a Fixed Budget: A Submodular View of Tree Search for Tool-Use Agentic Reinforcement Learning
by: Hu, Yuelin, et al.
Published: (2026)
by: Hu, Yuelin, et al.
Published: (2026)
Similar Items
-
SafeRL-Lite: A Lightweight, Explainable, and Constrained Reinforcement Learning Library
by: Mishra, Satyam, et al.
Published: (2025) -
Foundation Models as World Models: A Foundational Study in Text-Based GridWorlds
by: Sasso, Remo, et al.
Published: (2025) -
Exploration with Foundation Models: Capabilities, Limitations, and Hybrid Approaches
by: Sasso, Remo, et al.
Published: (2025) -
GoldenStart: Q-Guided Priors and Entropy Control for Distilling Flow Policies
by: Zhang, He, et al.
Published: (2026) -
An Improved Adaptive PID Optimizer with Enhanced Convergence and Stability for Deep Learning
by: Saini, Saurabh, et al.
Published: (2026)