RPO:Reinforcement Fine-Tuning with Partial Reasoning Optimization
Fuente:
arXiv
Saved in:
| Main Authors: | Yi, Hongzhu, Wang, Xinming, zhang, Zhenghao, Zong, Tianyu, Wang, Yuanxiang, Xie, Jun, Yu, Tao, Jin, Haopeng, Xu, Kaixin, Chen, Feng, Chen, Jiahuan, Yang, Yujia, Guan, Zhenyu, Shi, Bingkang, Xu, Jungang |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
JTCSE: Joint Tensor-Modulus Constraints and Cross-Attention for Unsupervised Contrastive Learning of Sentence Embeddings
by: Zong, Tianyu, et al.
Published: (2025)
by: Zong, Tianyu, et al.
Published: (2025)
TNCSE: Tensor's Norm Constraints for Unsupervised Contrastive Learning of Sentence Embeddings
by: Zong, Tianyu, et al.
Published: (2025)
by: Zong, Tianyu, et al.
Published: (2025)
Dynamic Deep Graph Learning for Incomplete Multi-View Clustering with Masked Graph Reconstruction Loss
by: Zhang, Zhenghao, et al.
Published: (2025)
by: Zhang, Zhenghao, et al.
Published: (2025)
VDE Bench: Evaluating The Capability of Image Editing Models to Modify Visual Documents
by: Yi, Hongzhu, et al.
Published: (2026)
by: Yi, Hongzhu, et al.
Published: (2026)
FAIRGAMER: Evaluating Social Biases in LLM-Based Video Game NPCs
by: Shi, Bingkang, et al.
Published: (2025)
by: Shi, Bingkang, et al.
Published: (2025)
Omni IIE Bench: Benchmarking the Practical Capabilities of Image Editing Models
by: Yang, Yujia, et al.
Published: (2026)
by: Yang, Yujia, et al.
Published: (2026)
RPO: Fine-Tuning Visual Generative Models via Rich Vision-Language Preferences
by: Zhao, Hanyang, et al.
Published: (2025)
by: Zhao, Hanyang, et al.
Published: (2025)
Towards Principled Dataset Distillation: A Spectral Distribution Perspective
by: Wu, Ruixi, et al.
Published: (2026)
by: Wu, Ruixi, et al.
Published: (2026)
Multimodal Video Emotion Recognition with Reliable Reasoning Priors
by: Wang, Zhepeng, et al.
Published: (2025)
by: Wang, Zhepeng, et al.
Published: (2025)
Diffusion-RPO: Aligning Diffusion Models through Relative Preference Optimization
by: Gu, Yi, et al.
Published: (2024)
by: Gu, Yi, et al.
Published: (2024)
StaRPO: Stability-Augmented Reinforcement Policy Optimization
by: Zhang, Jinghan, et al.
Published: (2026)
by: Zhang, Jinghan, et al.
Published: (2026)
Reason-RFT: Reinforcement Fine-Tuning for Visual Reasoning of Vision Language Models
by: Tan, Huajie, et al.
Published: (2025)
by: Tan, Huajie, et al.
Published: (2025)
More Is Better: A MoE-Based Emotion Recognition Framework with Human Preference Alignment
by: Xie, Jun, et al.
Published: (2025)
by: Xie, Jun, et al.
Published: (2025)
MR-Align: Meta-Reasoning Informed Factuality Alignment for Large Reasoning Models
by: Wang, Xinming, et al.
Published: (2025)
by: Wang, Xinming, et al.
Published: (2025)
QFFT, Question-Free Fine-Tuning for Adaptive Reasoning
by: Liu, Wanlong, et al.
Published: (2025)
by: Liu, Wanlong, et al.
Published: (2025)
Four Eyes Are Better Than Two: Harnessing the Collaborative Potential of Large Models via Differentiated Thinking and Complementary Ensembles
by: Xie, Jun, et al.
Published: (2025)
by: Xie, Jun, et al.
Published: (2025)
Team of One: Cracking Complex Video QA with Model Synergy
by: Xie, Jun, et al.
Published: (2025)
by: Xie, Jun, et al.
Published: (2025)
HY-Himmel Technical Report: Hierarchical Interleaved Multi-stream Motion Encoding for Long Video Understanding
by: Jin, Haopeng, et al.
Published: (2026)
by: Jin, Haopeng, et al.
Published: (2026)
TreeRPO: Tree Relative Policy Optimization
by: Yang, Zhicheng, et al.
Published: (2025)
by: Yang, Zhicheng, et al.
Published: (2025)
Clue Matters: Leveraging Latent Visual Clues to Empower Video Reasoning
by: zhang, Kaixin, et al.
Published: (2026)
by: zhang, Kaixin, et al.
Published: (2026)
Mitigating Catastrophic Forgetting with Adaptive Transformer Block Expansion in Federated Fine-Tuning
by: Huo, Yujia, et al.
Published: (2025)
by: Huo, Yujia, et al.
Published: (2025)
HFT: Half Fine-Tuning for Large Language Models
by: Hui, Tingfeng, et al.
Published: (2024)
by: Hui, Tingfeng, et al.
Published: (2024)
SRFT: A Single-Stage Method with Supervised and Reinforcement Fine-Tuning for Reasoning
by: Fu, Yuqian, et al.
Published: (2025)
by: Fu, Yuqian, et al.
Published: (2025)
FRoD: Full-Rank Efficient Fine-Tuning with Rotational Degrees for Fast Convergence
by: Wan, Guoan, et al.
Published: (2025)
by: Wan, Guoan, et al.
Published: (2025)
ReFT: Reasoning with Reinforced Fine-Tuning
by: Luong, Trung Quoc, et al.
Published: (2024)
by: Luong, Trung Quoc, et al.
Published: (2024)
Fine‐Tuning Electron Transfer for Nanozyme Design
by: Xia Zong, et al.
Published: (2024)
by: Xia Zong, et al.
Published: (2024)
Spatio-Temporal Partial Sensing Forecast for Long-term Traffic
by: Liu, Zibo, et al.
Published: (2024)
by: Liu, Zibo, et al.
Published: (2024)
Design Topological Materials by Reinforcement Fine-Tuned Generative Model
by: Xu, Haosheng, et al.
Published: (2025)
by: Xu, Haosheng, et al.
Published: (2025)
TS-SAM: Fine-Tuning Segment-Anything Model for Downstream Tasks
by: Yu, Yang, et al.
Published: (2024)
by: Yu, Yang, et al.
Published: (2024)
ScRPO: From Errors to Insights
by: Li, Lianrui, et al.
Published: (2025)
by: Li, Lianrui, et al.
Published: (2025)
PhySe-RPO: Physics and Semantics Guided Relative Policy Optimization for Diffusion-Based Surgical Smoke Removal
by: Fang, Zining, et al.
Published: (2026)
by: Fang, Zining, et al.
Published: (2026)
PaperX: A Unified Framework for Multimodal Academic Presentation Generation with Scholar DAG
by: Yu, Tao, et al.
Published: (2026)
by: Yu, Tao, et al.
Published: (2026)
Blending Supervised and Reinforcement Fine-Tuning with Prefix Sampling
by: Huang, Zeyu, et al.
Published: (2025)
by: Huang, Zeyu, et al.
Published: (2025)
ShotFinder: Imagination-Driven Open-Domain Video Shot Retrieval via Web Search
by: Yu, Tao, et al.
Published: (2026)
by: Yu, Tao, et al.
Published: (2026)
Reinforcement Fine-Tuning for Materials Design
by: Cao, Zhendong, et al.
Published: (2025)
by: Cao, Zhendong, et al.
Published: (2025)
ReAlign: Optimizing the Visual Document Retriever with Reasoning-Guided Fine-Grained Alignment
by: Yang, Hao, et al.
Published: (2026)
by: Yang, Hao, et al.
Published: (2026)
VLA-GSE: Boosting Parameter-Efficient Fine-Tuning in VLA with Generalized and Specialized Experts
by: Jiang, Yuhua, et al.
Published: (2026)
by: Jiang, Yuhua, et al.
Published: (2026)
Lang2Act: Fine-Grained Visual Reasoning through Self-Emergent Linguistic Toolchains
by: Xiong, Yuqi, et al.
Published: (2026)
by: Xiong, Yuqi, et al.
Published: (2026)
CReFT-CAD: Boosting Orthographic Projection Reasoning for CAD via Reinforcement Fine-Tuning
by: Niu, Ke, et al.
Published: (2025)
by: Niu, Ke, et al.
Published: (2025)
EgoMind: Activating Spatial Cognition through Linguistic Reasoning in MLLMs
by: Chen, Zhenghao, et al.
Published: (2026)
by: Chen, Zhenghao, et al.
Published: (2026)
Similar Items
-
JTCSE: Joint Tensor-Modulus Constraints and Cross-Attention for Unsupervised Contrastive Learning of Sentence Embeddings
by: Zong, Tianyu, et al.
Published: (2025) -
TNCSE: Tensor's Norm Constraints for Unsupervised Contrastive Learning of Sentence Embeddings
by: Zong, Tianyu, et al.
Published: (2025) -
Dynamic Deep Graph Learning for Incomplete Multi-View Clustering with Masked Graph Reconstruction Loss
by: Zhang, Zhenghao, et al.
Published: (2025) -
VDE Bench: Evaluating The Capability of Image Editing Models to Modify Visual Documents
by: Yi, Hongzhu, et al.
Published: (2026) -
FAIRGAMER: Evaluating Social Biases in LLM-Based Video Game NPCs
by: Shi, Bingkang, et al.
Published: (2025)