Circular Reasoning: Understanding Self-Reinforcing Loops in Large Reasoning Models
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Duan, Zenghao, Pang, Liang, Wei, Zihao, Duan, Wenbin, Tian, Yuxin, Xu, Shicheng, Deng, Jingcheng, Yin, Zhiyi, Cheng, Xueqi |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
SkillAttack: Automated Red Teaming of Agent Skills through Attack Path Refinement
von: Duan, Zenghao, et al.
Veröffentlicht: (2026)
von: Duan, Zenghao, et al.
Veröffentlicht: (2026)
LLM Latent Reasoning as Chain of Superposition
von: Deng, Jingcheng, et al.
Veröffentlicht: (2025)
von: Deng, Jingcheng, et al.
Veröffentlicht: (2025)
Latent-GRPO: Group Relative Policy Optimization for Latent Reasoning
von: Deng, Jingcheng, et al.
Veröffentlicht: (2026)
von: Deng, Jingcheng, et al.
Veröffentlicht: (2026)
The Evolution of Thought: Tracking LLM Overthinking via Reasoning Dynamics Analysis
von: Wei, Zihao, et al.
Veröffentlicht: (2025)
von: Wei, Zihao, et al.
Veröffentlicht: (2025)
RLKD: Distilling LLMs' Reasoning via Reinforcement Learning
von: Xu, Shicheng, et al.
Veröffentlicht: (2025)
von: Xu, Shicheng, et al.
Veröffentlicht: (2025)
Large Language Model Sourcing: A Survey
von: Pang, Liang, et al.
Veröffentlicht: (2025)
von: Pang, Liang, et al.
Veröffentlicht: (2025)
Related Knowledge Perturbation Matters: Rethinking Multiple Pieces of Knowledge Editing in Same-Subject
von: Duan, Zenghao, et al.
Veröffentlicht: (2025)
von: Duan, Zenghao, et al.
Veröffentlicht: (2025)
Stable Knowledge Editing in Large Language Models
von: Wei, Zihao, et al.
Veröffentlicht: (2024)
von: Wei, Zihao, et al.
Veröffentlicht: (2024)
MLaKE: Multilingual Knowledge Editing Benchmark for Large Language Models
von: Wei, Zihao, et al.
Veröffentlicht: (2024)
von: Wei, Zihao, et al.
Veröffentlicht: (2024)
GloSS over Toxicity: Understanding and Mitigating Toxicity in LLMs via Global Toxic Subspace
von: Duan, Zenghao, et al.
Veröffentlicht: (2025)
von: Duan, Zenghao, et al.
Veröffentlicht: (2025)
Everything is Editable: Extend Knowledge Editing to Unstructured Data in Large Language Models
von: Deng, Jingcheng, et al.
Veröffentlicht: (2024)
von: Deng, Jingcheng, et al.
Veröffentlicht: (2024)
Projecting Out the Malice: A Global Subspace Approach to LLM Detoxification
von: Duan, Zenghao, et al.
Veröffentlicht: (2026)
von: Duan, Zenghao, et al.
Veröffentlicht: (2026)
Reverse Physician-AI Relationship: Full-process Clinical Diagnosis Driven by a Large Language Model
von: Xu, Shicheng, et al.
Veröffentlicht: (2025)
von: Xu, Shicheng, et al.
Veröffentlicht: (2025)
Invisible Relevance Bias: Text-Image Retrieval Models Prefer AI-Generated Images
von: Xu, Shicheng, et al.
Veröffentlicht: (2023)
von: Xu, Shicheng, et al.
Veröffentlicht: (2023)
Cross-Modal Safety Mechanism Transfer in Large Vision-Language Models
von: Xu, Shicheng, et al.
Veröffentlicht: (2024)
von: Xu, Shicheng, et al.
Veröffentlicht: (2024)
Search-in-the-Chain: Interactively Enhancing Large Language Models with Search for Knowledge-intensive Tasks
von: Xu, Shicheng, et al.
Veröffentlicht: (2023)
von: Xu, Shicheng, et al.
Veröffentlicht: (2023)
A Theory for Token-Level Harmonization in Retrieval-Augmented Generation
von: Xu, Shicheng, et al.
Veröffentlicht: (2024)
von: Xu, Shicheng, et al.
Veröffentlicht: (2024)
List-aware Reranking-Truncation Joint Model for Search and Retrieval-augmented Generation
von: Xu, Shicheng, et al.
Veröffentlicht: (2024)
von: Xu, Shicheng, et al.
Veröffentlicht: (2024)
Following the Autoregressive Nature of LLM Embeddings via Compression and Alignment
von: Deng, Jingcheng, et al.
Veröffentlicht: (2025)
von: Deng, Jingcheng, et al.
Veröffentlicht: (2025)
OThink-SRR1: Search, Refine and Reasoning with Reinforced Learning for Large Language Models
von: Liang, Haijian, et al.
Veröffentlicht: (2026)
von: Liang, Haijian, et al.
Veröffentlicht: (2026)
Unsupervised Information Refinement Training of Large Language Models for Retrieval-Augmented Generation
von: Xu, Shicheng, et al.
Veröffentlicht: (2024)
von: Xu, Shicheng, et al.
Veröffentlicht: (2024)
from Benign import Toxic: Jailbreaking the Language Model via Adversarial Metaphors
von: Yan, Yu, et al.
Veröffentlicht: (2025)
von: Yan, Yu, et al.
Veröffentlicht: (2025)
Rowen: Adaptive Retrieval-Augmented Generation for Hallucination Mitigation in LLMs
von: Ding, Hanxing, et al.
Veröffentlicht: (2024)
von: Ding, Hanxing, et al.
Veröffentlicht: (2024)
D-Models and E-Models: Diversity-Stability Trade-offs in the Sampling Behavior of Large Language Models
von: Gu, Jia, et al.
Veröffentlicht: (2026)
von: Gu, Jia, et al.
Veröffentlicht: (2026)
Enhancing Training Data Attribution for Large Language Models with Fitting Error Consideration
von: Wu, Kangxi, et al.
Veröffentlicht: (2024)
von: Wu, Kangxi, et al.
Veröffentlicht: (2024)
Scalable Uncertainty Reasoning in Knowledge Graphs
von: Wu, Jingcheng
Veröffentlicht: (2026)
von: Wu, Jingcheng
Veröffentlicht: (2026)
Multi-tool Integration Application for Math Reasoning Using Large Language Model
von: Duan, Zhihua, et al.
Veröffentlicht: (2024)
von: Duan, Zhihua, et al.
Veröffentlicht: (2024)
Do LLMs Play Dice? Exploring Probability Distribution Sampling in Large Language Models for Behavioral Simulation
von: Gu, Jia, et al.
Veröffentlicht: (2024)
von: Gu, Jia, et al.
Veröffentlicht: (2024)
Think Before You Speak: Cultivating Communication Skills of Large Language Models via Inner Monologue
von: Zhou, Junkai, et al.
Veröffentlicht: (2023)
von: Zhou, Junkai, et al.
Veröffentlicht: (2023)
ToolCoder: A Systematic Code-Empowered Tool Learning Framework for Large Language Models
von: Ding, Hanxing, et al.
Veröffentlicht: (2025)
von: Ding, Hanxing, et al.
Veröffentlicht: (2025)
Innate Reasoning is Not Enough: In-Context Learning Enhances Reasoning Large Language Models with Less Overthinking
von: Ge, Yuyao, et al.
Veröffentlicht: (2025)
von: Ge, Yuyao, et al.
Veröffentlicht: (2025)
Reinforcement Learning with Knowledge Representation and Reasoning: A Brief Survey
von: Yu, Chao, et al.
Veröffentlicht: (2023)
von: Yu, Chao, et al.
Veröffentlicht: (2023)
Flexible Circular Antenna Sensor With Multi Loops for Plantar Pressure Detection
von: Chunyu Ao, et al.
Veröffentlicht: (2026)
von: Chunyu Ao, et al.
Veröffentlicht: (2026)
DELTA: Deliberative Multi-Agent Reasoning with Reinforcement Learning for Multimodal Psychological Counseling
von: Yang, Jiangnan, et al.
Veröffentlicht: (2026)
von: Yang, Jiangnan, et al.
Veröffentlicht: (2026)
Look Globally and Reason: Two-stage Path Reasoning over Sparse Knowledge Graphs
von: Guan, Saiping, et al.
Veröffentlicht: (2024)
von: Guan, Saiping, et al.
Veröffentlicht: (2024)
Internalizing Safety Understanding in Large Reasoning Models via Verification
von: Zhang, Yi, et al.
Veröffentlicht: (2026)
von: Zhang, Yi, et al.
Veröffentlicht: (2026)
Thinking Forward and Backward: Multi-Objective Reinforcement Learning for Retrieval-Augmented Reasoning
von: Wei, Wenda, et al.
Veröffentlicht: (2025)
von: Wei, Wenda, et al.
Veröffentlicht: (2025)
From Large to Small: Transferring CUDA Optimization Expertise via Reasoning Graph
von: Gong, Junfeng, et al.
Veröffentlicht: (2025)
von: Gong, Junfeng, et al.
Veröffentlicht: (2025)
Generative Ghost: Investigating Ranking Bias Hidden in AI-Generated Videos
von: Gao, Haowen, et al.
Veröffentlicht: (2025)
von: Gao, Haowen, et al.
Veröffentlicht: (2025)
Unilaw-R1: A Large Language Model for Legal Reasoning with Reinforcement Learning and Iterative Inference
von: Cai, Hua, et al.
Veröffentlicht: (2025)
von: Cai, Hua, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
SkillAttack: Automated Red Teaming of Agent Skills through Attack Path Refinement
von: Duan, Zenghao, et al.
Veröffentlicht: (2026) -
LLM Latent Reasoning as Chain of Superposition
von: Deng, Jingcheng, et al.
Veröffentlicht: (2025) -
Latent-GRPO: Group Relative Policy Optimization for Latent Reasoning
von: Deng, Jingcheng, et al.
Veröffentlicht: (2026) -
The Evolution of Thought: Tracking LLM Overthinking via Reasoning Dynamics Analysis
von: Wei, Zihao, et al.
Veröffentlicht: (2025) -
RLKD: Distilling LLMs' Reasoning via Reinforcement Learning
von: Xu, Shicheng, et al.
Veröffentlicht: (2025)