SRPO: A Cross-Domain Implementation of Large-Scale Reinforcement Learning on LLM
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Zhang, Xiaojiang, Wang, Jinghui, Cheng, Zifei, Zhuang, Wenhao, Lin, Zheng, Zhang, Minglei, Wang, Shaojie, Cui, Yinghan, Wang, Chao, Peng, Junyi, Jiang, Shimiao, Kuang, Shiqi, Yin, Shouyu, Wen, Chaohang, Zhang, Haotian, Chen, Bin, Yu, Bing |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Tree Training: Accelerating Agentic LLMs Training via Shared Prefix Reuse
von: Wang, Jinghui, et al.
Veröffentlicht: (2025)
von: Wang, Jinghui, et al.
Veröffentlicht: (2025)
SeamlessFlow: A Trainer Agent Isolation RL Framework Achieving Bubble-Free Pipelines via Tag Scheduling
von: Wang, Jinghui, et al.
Veröffentlicht: (2025)
von: Wang, Jinghui, et al.
Veröffentlicht: (2025)
SRPO: Enhancing Multimodal LLM Reasoning via Reflection-Aware Reinforcement Learning
von: Wan, Zhongwei, et al.
Veröffentlicht: (2025)
von: Wan, Zhongwei, et al.
Veröffentlicht: (2025)
SRPO: Self-Referential Policy Optimization for Vision-Language-Action Models
von: Fei, Senyu, et al.
Veröffentlicht: (2025)
von: Fei, Senyu, et al.
Veröffentlicht: (2025)
A Practical Introduction to Deep Reinforcement Learning
von: Sun, Yinghan, et al.
Veröffentlicht: (2025)
von: Sun, Yinghan, et al.
Veröffentlicht: (2025)
Unravelling the Oxygen Dynamics at Gas‐Solid Interface of CaGd 2 (MoO 4 ) 4 :Yb,Er Scheelite Ceramics via Upconversion Luminescence
von: Yinghan Wang, et al.
Veröffentlicht: (2025)
von: Yinghan Wang, et al.
Veröffentlicht: (2025)
KAT-Coder Technical Report
von: Zhan, Zizheng, et al.
Veröffentlicht: (2025)
von: Zhan, Zizheng, et al.
Veröffentlicht: (2025)
Effect of the cycle number of multistep sintering on the properties of FeSe 0.4 Te 0.6 superconductors
von: Chaohang Miao, et al.
Veröffentlicht: (2024)
von: Chaohang Miao, et al.
Veröffentlicht: (2024)
Amorphous Sulfo‐Halide Solid Electrolytes With Enhanced Anion Dynamics for Highly Stable All‐Solid‐State Sodium Batteries
von: Zifei Shi, et al.
Veröffentlicht: (2026)
von: Zifei Shi, et al.
Veröffentlicht: (2026)
$Δl =1$ coupling of single-particle orbitals in octupole deformed nuclei
von: Wang, XuDong, et al.
Veröffentlicht: (2026)
von: Wang, XuDong, et al.
Veröffentlicht: (2026)
Large and Small Yellow Croakers Feeding and Living Together Make Large Yellow Croaker Population Recovery Difficult: A Guild Perspective.
von: Cai, Pengyu, et al.
Veröffentlicht: (2024)
von: Cai, Pengyu, et al.
Veröffentlicht: (2024)
Exploiting sparse structures and synergy designs to advance situational awareness of electrical power grid
von: Li, Shimiao
Veröffentlicht: (2024)
von: Li, Shimiao
Veröffentlicht: (2024)
Two Birds with One Stone: Enhancing Uncertainty Quantification and Interpretability with Graph Functional Neural Process
von: Kong, Lingkai, et al.
Veröffentlicht: (2025)
von: Kong, Lingkai, et al.
Veröffentlicht: (2025)
From Implicit to Explicit: Token-Efficient Logical Supervision for Mathematical Reasoning in LLMs
von: Wang, Shaojie, et al.
Veröffentlicht: (2026)
von: Wang, Shaojie, et al.
Veröffentlicht: (2026)
From Meta-Thought to Execution: Cognitively Aligned Post-Training for Generalizable and Reliable LLM Reasoning
von: Wang, Shaojie, et al.
Veröffentlicht: (2026)
von: Wang, Shaojie, et al.
Veröffentlicht: (2026)
Mechanical Experiments and ABAQUS Simulations of Fiber–Coal Gangue Stabilized Expansive Soil Under Freeze–Thaw Cycles
von: Yan Zhang, et al.
Veröffentlicht: (2024)
von: Yan Zhang, et al.
Veröffentlicht: (2024)
Reliable Intelligent Diagnostics With Uncertainty Quantification for Mechanical System Condition Assessment
von: Guangjun Jiang, et al.
Veröffentlicht: (2026)
von: Guangjun Jiang, et al.
Veröffentlicht: (2026)
Modified Arthroscopic Technique for Repair of Medial Meniscus Posterior Root Tear and Centralization of the Extruded Medial Meniscus Through a Double Transtibial Tunnel (Without Use of Accessory Ports)
von: Parvind Grish Dev Sookun, et al.
Veröffentlicht: (2025)
von: Parvind Grish Dev Sookun, et al.
Veröffentlicht: (2025)
A good neighbor, a found treasure: Do local neighbors affect corporate innovation?
von: Shaopeng Cao, et al.
Veröffentlicht: (2024)
von: Shaopeng Cao, et al.
Veröffentlicht: (2024)
Credit Where It is Due: Cross-Modality Connectivity Drives Precise Reinforcement Learning for MLLM Reasoning
von: Jiao, Zhengbo, et al.
Veröffentlicht: (2026)
von: Jiao, Zhengbo, et al.
Veröffentlicht: (2026)
Divide-and-Conquer Decoupled Network for Cross-Domain Few-Shot Segmentation
von: Cong, Runmin, et al.
Veröffentlicht: (2025)
von: Cong, Runmin, et al.
Veröffentlicht: (2025)
A Camera‐Like Dual‐Defocus Curvature Wavefront Sensor With GPU Acceleration for Real‐Time Quantitative Phase Imaging
von: Wei Wang, et al.
Veröffentlicht: (2025)
von: Wei Wang, et al.
Veröffentlicht: (2025)
Balancing the Reasoning Load: Difficulty-Differentiated Policy Optimization with Length Redistribution for Efficient and Robust Reinforcement Learning
von: Xia, Yinan, et al.
Veröffentlicht: (2026)
von: Xia, Yinan, et al.
Veröffentlicht: (2026)
On the Dirichlet Problem at Infinity and Poisson Boundary for Certain Manifolds without Conjugate Points
von: Liu, Fei, et al.
Veröffentlicht: (2025)
von: Liu, Fei, et al.
Veröffentlicht: (2025)
Dispersion-Domain Detection for Mobile Molecular Communication Under Multiplicative Geometry Uncertainty
von: Zhang, Shaojie, et al.
Veröffentlicht: (2026)
von: Zhang, Shaojie, et al.
Veröffentlicht: (2026)
The Impact of Quantization on Large Reasoning Model Reinforcement Learning
von: Kumar, Medha, et al.
Veröffentlicht: (2025)
von: Kumar, Medha, et al.
Veröffentlicht: (2025)
Do Institutional Cross‐Owners Obstruct Corporate Environmental Information Disclosure? Evidence From China
von: Xin Cui, et al.
Veröffentlicht: (2025)
von: Xin Cui, et al.
Veröffentlicht: (2025)
Exploring Task-Solving Paradigm for Generalized Cross-Domain Face Anti-Spoofing via Reinforcement Fine-Tuning
von: Jiang, Fangling, et al.
Veröffentlicht: (2025)
von: Jiang, Fangling, et al.
Veröffentlicht: (2025)
Microquasar Remnants as Pevatrons Illuminating the Galactic Cosmic Ray Knee
von: Zhang, Bing Theodore, et al.
Veröffentlicht: (2026)
von: Zhang, Bing Theodore, et al.
Veröffentlicht: (2026)
A Unified Framework for 10 TeV to EeV Diffuse Neutrino Sky and KM3-230213A
von: Yu, Shiqi, et al.
Veröffentlicht: (2026)
von: Yu, Shiqi, et al.
Veröffentlicht: (2026)
Multi-Messenger Modeling of Low-Luminosity Gamma-Ray Bursts
von: Yu, Shiqi, et al.
Veröffentlicht: (2026)
von: Yu, Shiqi, et al.
Veröffentlicht: (2026)
BTL-UI: Blink-Think-Link Reasoning Model for GUI Agent
von: Zhang, Shaojie, et al.
Veröffentlicht: (2025)
von: Zhang, Shaojie, et al.
Veröffentlicht: (2025)
Discrepancies in Pedestrian Crossing of Static vs. Dynamic Crowds: An Experimental Study
von: Wang, Jinghui, et al.
Veröffentlicht: (2024)
von: Wang, Jinghui, et al.
Veröffentlicht: (2024)
OASIS: Conditional Distribution Shaping for Offline Safe Reinforcement Learning
von: Yao, Yihang, et al.
Veröffentlicht: (2024)
von: Yao, Yihang, et al.
Veröffentlicht: (2024)
Reinforcement Learning for Tool-Integrated Interleaved Thinking towards Cross-Domain Generalization
von: Chen, Zhengyu, et al.
Veröffentlicht: (2025)
von: Chen, Zhengyu, et al.
Veröffentlicht: (2025)
Contrastive Representation for Data Filtering in Cross-Domain Offline Reinforcement Learning
von: Wen, Xiaoyu, et al.
Veröffentlicht: (2024)
von: Wen, Xiaoyu, et al.
Veröffentlicht: (2024)
Towards Monotonic Improvement in In-Context Reinforcement Learning
von: Zhang, Wenhao, et al.
Veröffentlicht: (2025)
von: Zhang, Wenhao, et al.
Veröffentlicht: (2025)
Preferences for Processing Facial Information at Different Orientations in Individuals With Developmental Prosopagnosia
von: Jialin Ma, et al.
Veröffentlicht: (2026)
von: Jialin Ma, et al.
Veröffentlicht: (2026)
ClassEval-Pro: A Cross-Domain Benchmark for Class-Level Code Generation
von: Chen, Yeheng, et al.
Veröffentlicht: (2026)
von: Chen, Yeheng, et al.
Veröffentlicht: (2026)
Prototype-Regularized Federated Learning for Cross-Domain Aspect Sentiment Triplet Extraction
von: Cai, Zongming, et al.
Veröffentlicht: (2026)
von: Cai, Zongming, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
Tree Training: Accelerating Agentic LLMs Training via Shared Prefix Reuse
von: Wang, Jinghui, et al.
Veröffentlicht: (2025) -
SeamlessFlow: A Trainer Agent Isolation RL Framework Achieving Bubble-Free Pipelines via Tag Scheduling
von: Wang, Jinghui, et al.
Veröffentlicht: (2025) -
SRPO: Enhancing Multimodal LLM Reasoning via Reflection-Aware Reinforcement Learning
von: Wan, Zhongwei, et al.
Veröffentlicht: (2025) -
SRPO: Self-Referential Policy Optimization for Vision-Language-Action Models
von: Fei, Senyu, et al.
Veröffentlicht: (2025) -
A Practical Introduction to Deep Reinforcement Learning
von: Sun, Yinghan, et al.
Veröffentlicht: (2025)