InSPO: Unlocking Intrinsic Self-Reflection for LLM Preference Optimization
Fuente:
arXiv
Saved in:
| Main Authors: | Li, Yu, Lan, Tian, Qi, Zhengling |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
OPPO: Bayesian Value Recursion for Token-Level Credit Assignment in LLM Reasoning
by: Li, Yu, et al.
Published: (2026)
by: Li, Yu, et al.
Published: (2026)
ARISE: Agent Reasoning with Intrinsic Skill Evolution in Hierarchical Reinforcement Learning
by: Li, Yu, et al.
Published: (2026)
by: Li, Yu, et al.
Published: (2026)
SPO: Sequential Monte Carlo Policy Optimisation
by: Macfarlane, Matthew V, et al.
Published: (2024)
by: Macfarlane, Matthew V, et al.
Published: (2024)
Reason in Chains, Learn in Trees: Self-Rectification and Grafting for Multi-turn Agent Policy Optimization
by: Li, Yu, et al.
Published: (2026)
by: Li, Yu, et al.
Published: (2026)
Reflective Preference Optimization (RPO): Enhancing On-Policy Alignment via Hint-Guided Reflection
by: Zhao, Zihui, et al.
Published: (2025)
by: Zhao, Zihui, et al.
Published: (2025)
When Right Meets Wrong: Bilateral Context Conditioning with Reward-Confidence Correction for GRPO
by: Li, Yu, et al.
Published: (2026)
by: Li, Yu, et al.
Published: (2026)
CriSPO: Multi-Aspect Critique-Suggestion-guided Automatic Prompt Optimization for Text Generation
by: He, Han, et al.
Published: (2024)
by: He, Han, et al.
Published: (2024)
SIRI: Self-Internalizing Reinforcement Learning with Intrinsic Skills for LLM Agent Training
by: He, Zhongyu, et al.
Published: (2026)
by: He, Zhongyu, et al.
Published: (2026)
TSO: Self-Training with Scaled Preference Optimization
by: Chen, Kaihui, et al.
Published: (2024)
by: Chen, Kaihui, et al.
Published: (2024)
Self-Improving Robust Preference Optimization
by: Choi, Eugene, et al.
Published: (2024)
by: Choi, Eugene, et al.
Published: (2024)
Enhancing LLM Planning Capabilities through Intrinsic Self-Critique
by: Bohnet, Bernd, et al.
Published: (2025)
by: Bohnet, Bernd, et al.
Published: (2025)
Robust Offline Reinforcement learning with Heavy-Tailed Rewards
by: Zhu, Jin, et al.
Published: (2023)
by: Zhu, Jin, et al.
Published: (2023)
Self-Consistency Preference Optimization
by: Prasad, Archiki, et al.
Published: (2024)
by: Prasad, Archiki, et al.
Published: (2024)
Beyond Demand Estimation: Consumer Surplus Evaluation via Cumulative Propensity Weights
by: Bian, Zeyu, et al.
Published: (2026)
by: Bian, Zeyu, et al.
Published: (2026)
Experiential Reflective Learning for Self-Improving LLM Agents
by: Allard, Marc-Antoine, et al.
Published: (2026)
by: Allard, Marc-Antoine, et al.
Published: (2026)
RLHF Can Speak Many Languages: Unlocking Multilingual Preference Optimization for LLMs
by: Dang, John, et al.
Published: (2024)
by: Dang, John, et al.
Published: (2024)
Time-Prompt: Integrated Heterogeneous Prompts for Unlocking LLMs in Time Series Forecasting
by: Wang, Zesen, et al.
Published: (2025)
by: Wang, Zesen, et al.
Published: (2025)
A Theoretical Study on Bridging Internal Probability and Self-Consistency for LLM Reasoning
by: Zhou, Zhi, et al.
Published: (2025)
by: Zhou, Zhi, et al.
Published: (2025)
PROF: An LLM-based Reward Code Preference Optimization Framework for Offline Imitation Learning
by: Sun, Shengjie, et al.
Published: (2025)
by: Sun, Shengjie, et al.
Published: (2025)
Distribution Preference Optimization: A Fine-grained Perspective for LLM Unlearning
by: Qin, Kai, et al.
Published: (2025)
by: Qin, Kai, et al.
Published: (2025)
Robust LLM Alignment via Distributionally Robust Direct Preference Optimization
by: Xu, Zaiyan, et al.
Published: (2025)
by: Xu, Zaiyan, et al.
Published: (2025)
Boosting In-Context Learning in LLMs Through the Lens of Classical Supervised Learning
by: Gundem, Korel, et al.
Published: (2025)
by: Gundem, Korel, et al.
Published: (2025)
Online Preference-based Reinforcement Learning with Self-augmented Feedback from Large Language Model
by: Tu, Songjun, et al.
Published: (2024)
by: Tu, Songjun, et al.
Published: (2024)
AMPO: Active Multi-Preference Optimization for Self-play Preference Selection
by: Gupta, Taneesh, et al.
Published: (2025)
by: Gupta, Taneesh, et al.
Published: (2025)
DPIC: Decoupling Prompt and Intrinsic Characteristics for LLM Generated Text Detection
by: Yu, Xiao, et al.
Published: (2023)
by: Yu, Xiao, et al.
Published: (2023)
Risk-aware Direct Preference Optimization under Nested Risk Measure
by: Zhang, Lijun, et al.
Published: (2025)
by: Zhang, Lijun, et al.
Published: (2025)
Quantifying and Mitigating Self-Preference Bias of LLM Judges
by: Yang, Jinming, et al.
Published: (2026)
by: Yang, Jinming, et al.
Published: (2026)
Intrinsic Muon: Spectral Optimization on Riemannian Matrix Manifolds
by: Li, Yibang, et al.
Published: (2026)
by: Li, Yibang, et al.
Published: (2026)
Thinking Preference Optimization
by: Yang, Wang, et al.
Published: (2025)
by: Yang, Wang, et al.
Published: (2025)
Beyond One-Preference-Fits-All Alignment: Multi-Objective Direct Preference Optimization
by: Zhou, Zhanhui, et al.
Published: (2023)
by: Zhou, Zhanhui, et al.
Published: (2023)
Interactive Critique-Revision Training for Reliable Structured LLM Generation
by: Yu, Fei Xu, et al.
Published: (2026)
by: Yu, Fei Xu, et al.
Published: (2026)
SSPO: Self-traced Step-wise Preference Optimization for Process Supervision and Reasoning Compression
by: Xu, Yuyang, et al.
Published: (2025)
by: Xu, Yuyang, et al.
Published: (2025)
Hindsight Preference Optimization for Financial Time Series Advisory
by: Cui, Yanwei, et al.
Published: (2026)
by: Cui, Yanwei, et al.
Published: (2026)
Preference as Reward, Maximum Preference Optimization with Importance Sampling
by: Jiang, Zaifan, et al.
Published: (2023)
by: Jiang, Zaifan, et al.
Published: (2023)
Reflective Policy Optimization
by: Gan, Yaozhong, et al.
Published: (2024)
by: Gan, Yaozhong, et al.
Published: (2024)
Light Alignment Improves LLM Safety via Model Self-Reflection with a Single Neuron
by: Shen, Sicheng, et al.
Published: (2026)
by: Shen, Sicheng, et al.
Published: (2026)
To Call or Not to Call: Diagnosing Intrinsic Over-Calling Bias in LLM Agents
by: Shi, Wei, et al.
Published: (2026)
by: Shi, Wei, et al.
Published: (2026)
On Symmetric Losses for Robust Policy Optimization with Noisy Preferences
by: Nishimori, Soichiro, et al.
Published: (2025)
by: Nishimori, Soichiro, et al.
Published: (2025)
Refining Alignment Framework for Diffusion Models with Intermediate-Step Preference Ranking
by: Ren, Jie, et al.
Published: (2025)
by: Ren, Jie, et al.
Published: (2025)
Self-Play Preference Optimization for Language Model Alignment
by: Wu, Yue, et al.
Published: (2024)
by: Wu, Yue, et al.
Published: (2024)
Similar Items
-
OPPO: Bayesian Value Recursion for Token-Level Credit Assignment in LLM Reasoning
by: Li, Yu, et al.
Published: (2026) -
ARISE: Agent Reasoning with Intrinsic Skill Evolution in Hierarchical Reinforcement Learning
by: Li, Yu, et al.
Published: (2026) -
SPO: Sequential Monte Carlo Policy Optimisation
by: Macfarlane, Matthew V, et al.
Published: (2024) -
Reason in Chains, Learn in Trees: Self-Rectification and Grafting for Multi-turn Agent Policy Optimization
by: Li, Yu, et al.
Published: (2026) -
Reflective Preference Optimization (RPO): Enhancing On-Policy Alignment via Hint-Guided Reflection
by: Zhao, Zihui, et al.
Published: (2025)