In-Token Rationality Optimization: Towards Accurate and Concise LLM Reasoning via Self-Feedback
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Zhu, Mingye, Liu, Yi, Fu, Zheren, Wang, Quan, Zhang, Yongdong |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Leveraging Robust Optimization for LLM Alignment under Distribution Shifts
von: Zhu, Mingye, et al.
Veröffentlicht: (2025)
von: Zhu, Mingye, et al.
Veröffentlicht: (2025)
DACL-RAG: Data Augmentation Strategy with Curriculum Learning for Retrieval-Augmented Generation
von: Wang, Shaohan, et al.
Veröffentlicht: (2025)
von: Wang, Shaohan, et al.
Veröffentlicht: (2025)
FlipGuard: Defending Preference Alignment against Update Regression with Constrained Optimization
von: Zhu, Mingye, et al.
Veröffentlicht: (2024)
von: Zhu, Mingye, et al.
Veröffentlicht: (2024)
Leveraging Importance Sampling to Detach Alignment Modules from Large Language Models
von: Liu, Yi, et al.
Veröffentlicht: (2025)
von: Liu, Yi, et al.
Veröffentlicht: (2025)
SparseRM: A Lightweight Preference Modeling with Sparse Autoencoder
von: Liu, Dengcan, et al.
Veröffentlicht: (2025)
von: Liu, Dengcan, et al.
Veröffentlicht: (2025)
Concise Reasoning via Reinforcement Learning
von: Fatemi, Mehdi, et al.
Veröffentlicht: (2025)
von: Fatemi, Mehdi, et al.
Veröffentlicht: (2025)
Towards Concise and Adaptive Thinking in Large Reasoning Models: A Survey
von: Zhu, Jason, et al.
Veröffentlicht: (2025)
von: Zhu, Jason, et al.
Veröffentlicht: (2025)
Walk Before You Run! Concise LLM Reasoning via Reinforcement Learning
von: Song, Mingyang, et al.
Veröffentlicht: (2025)
von: Song, Mingyang, et al.
Veröffentlicht: (2025)
Not All Tokens Matter: Towards Efficient LLM Reasoning via Token Significance in Reinforcement Learning
von: Liu, Hanbing, et al.
Veröffentlicht: (2025)
von: Liu, Hanbing, et al.
Veröffentlicht: (2025)
ConciseHint: Boosting Efficient Reasoning via Continuous Concise Hints during Generation
von: Tang, Siao, et al.
Veröffentlicht: (2025)
von: Tang, Siao, et al.
Veröffentlicht: (2025)
On-the-fly Preference Alignment via Principle-Guided Decoding
von: Zhu, Mingye, et al.
Veröffentlicht: (2025)
von: Zhu, Mingye, et al.
Veröffentlicht: (2025)
Training LLM-Based Agents with Synthetic Self-Reflected Trajectories and Partial Masking
von: Chen, Yihan, et al.
Veröffentlicht: (2025)
von: Chen, Yihan, et al.
Veröffentlicht: (2025)
Self-signals Driven Multi-LLM Debate for Efficient and Accurate Reasoning
von: Chen, Xuhang, et al.
Veröffentlicht: (2025)
von: Chen, Xuhang, et al.
Veröffentlicht: (2025)
Concise Thoughts: Impact of Output Length on LLM Reasoning and Cost
von: Nayab, Sania, et al.
Veröffentlicht: (2024)
von: Nayab, Sania, et al.
Veröffentlicht: (2024)
HAPO: Training Language Models to Reason Concisely via History-Aware Policy Optimization
von: Huang, Chengyu, et al.
Veröffentlicht: (2025)
von: Huang, Chengyu, et al.
Veröffentlicht: (2025)
Robust Reasoning via Dynamic Token Selection for Distribution-Aligned Self-Distillation
von: Zhang, Ruiqi, et al.
Veröffentlicht: (2026)
von: Zhang, Ruiqi, et al.
Veröffentlicht: (2026)
Learning to Reason via Self-Iterative Process Feedback for Small Language Models
von: Chen, Kaiyuan, et al.
Veröffentlicht: (2024)
von: Chen, Kaiyuan, et al.
Veröffentlicht: (2024)
SelfBudgeter: Adaptive Token Allocation for Efficient LLM Reasoning
von: Li, Zheng, et al.
Veröffentlicht: (2025)
von: Li, Zheng, et al.
Veröffentlicht: (2025)
Adaptive Group Policy Optimization: Towards Stable Training and Token-Efficient Reasoning
von: Li, Chen, et al.
Veröffentlicht: (2025)
von: Li, Chen, et al.
Veröffentlicht: (2025)
Token-Guard: Towards Token-Level Hallucination Control via Self-Checking Decoding
von: Zhu, Yifan, et al.
Veröffentlicht: (2026)
von: Zhu, Yifan, et al.
Veröffentlicht: (2026)
Zero-Shot Detection of LLM-Generated Text using Token Cohesiveness
von: Ma, Shixuan, et al.
Veröffentlicht: (2024)
von: Ma, Shixuan, et al.
Veröffentlicht: (2024)
Self-Training Elicits Concise Reasoning in Large Language Models
von: Munkhbat, Tergel, et al.
Veröffentlicht: (2025)
von: Munkhbat, Tergel, et al.
Veröffentlicht: (2025)
Reasoning in Token Economies: Budget-Aware Evaluation of LLM Reasoning Strategies
von: Wang, Junlin, et al.
Veröffentlicht: (2024)
von: Wang, Junlin, et al.
Veröffentlicht: (2024)
Enhancing Persona Following at Decoding Time via Dynamic Importance Estimation for Role-Playing Agents
von: Liu, Yuxin, et al.
Veröffentlicht: (2026)
von: Liu, Yuxin, et al.
Veröffentlicht: (2026)
Steering Large Reasoning Models towards Concise Reasoning via Flow Matching
von: Li, Yawei, et al.
Veröffentlicht: (2026)
von: Li, Yawei, et al.
Veröffentlicht: (2026)
LIRE: listwise reward enhancement for preference alignment
von: Zhu, Mingye, et al.
Veröffentlicht: (2024)
von: Zhu, Mingye, et al.
Veröffentlicht: (2024)
Accurate KV Cache Quantization with Outlier Tokens Tracing
von: Su, Yi, et al.
Veröffentlicht: (2025)
von: Su, Yi, et al.
Veröffentlicht: (2025)
Concise and Organized Perception Facilitates Reasoning in Large Language Models
von: Liu, Junjie, et al.
Veröffentlicht: (2023)
von: Liu, Junjie, et al.
Veröffentlicht: (2023)
Sample More to Think Less: Group Filtered Policy Optimization for Concise Reasoning
von: Shrivastava, Vaishnavi, et al.
Veröffentlicht: (2025)
von: Shrivastava, Vaishnavi, et al.
Veröffentlicht: (2025)
Preference Optimization for Reasoning with Pseudo Feedback
von: Jiao, Fangkai, et al.
Veröffentlicht: (2024)
von: Jiao, Fangkai, et al.
Veröffentlicht: (2024)
HAMburger: Accelerating LLM Inference via Token Smashing
von: Liu, Jingyu, et al.
Veröffentlicht: (2025)
von: Liu, Jingyu, et al.
Veröffentlicht: (2025)
AERO: Autonomous Evolutionary Reasoning Optimization via Endogenous Dual-Loop Feedback
von: Gao, Zhitao, et al.
Veröffentlicht: (2026)
von: Gao, Zhitao, et al.
Veröffentlicht: (2026)
Align Documents to Questions: Question-Oriented Document Rewriting for Retrieval-Augmented Generation
von: Li, Jiaang, et al.
Veröffentlicht: (2026)
von: Li, Jiaang, et al.
Veröffentlicht: (2026)
How Do Answer Tokens Read Reasoning Traces? Self-Reading Patterns in Thinking LLMs for Quantitative Reasoning
von: Chen, Haoyang, et al.
Veröffentlicht: (2026)
von: Chen, Haoyang, et al.
Veröffentlicht: (2026)
Self-Reflective Planning with Knowledge Graphs: Enhancing LLM Reasoning Reliability for Question Answering
von: Zhu, Jiajun, et al.
Veröffentlicht: (2025)
von: Zhu, Jiajun, et al.
Veröffentlicht: (2025)
SPA: Achieving Consensus in LLM Alignment via Self-Priority Optimization
von: Huang, Yue, et al.
Veröffentlicht: (2025)
von: Huang, Yue, et al.
Veröffentlicht: (2025)
EASYTOOL: Enhancing LLM-based Agents with Concise Tool Instruction
von: Yuan, Siyu, et al.
Veröffentlicht: (2024)
von: Yuan, Siyu, et al.
Veröffentlicht: (2024)
LLMs are Superior Feedback Providers: Bootstrapping Reasoning for Lie Detection with Self-Generated Feedback
von: Banerjee, Tanushree, et al.
Veröffentlicht: (2024)
von: Banerjee, Tanushree, et al.
Veröffentlicht: (2024)
Alignment-Enhanced Decoding:Defending via Token-Level Adaptive Refining of Probability Distributions
von: Liu, Quan, et al.
Veröffentlicht: (2024)
von: Liu, Quan, et al.
Veröffentlicht: (2024)
A Study on Leveraging Search and Self-Feedback for Agent Reasoning
von: K, Karthikeyan, et al.
Veröffentlicht: (2025)
von: K, Karthikeyan, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Leveraging Robust Optimization for LLM Alignment under Distribution Shifts
von: Zhu, Mingye, et al.
Veröffentlicht: (2025) -
DACL-RAG: Data Augmentation Strategy with Curriculum Learning for Retrieval-Augmented Generation
von: Wang, Shaohan, et al.
Veröffentlicht: (2025) -
FlipGuard: Defending Preference Alignment against Update Regression with Constrained Optimization
von: Zhu, Mingye, et al.
Veröffentlicht: (2024) -
Leveraging Importance Sampling to Detach Alignment Modules from Large Language Models
von: Liu, Yi, et al.
Veröffentlicht: (2025) -
SparseRM: A Lightweight Preference Modeling with Sparse Autoencoder
von: Liu, Dengcan, et al.
Veröffentlicht: (2025)