SERL: Self-Examining Reinforcement Learning on Open-Domain
Fuente:
arXiv
Saved in:
| Main Authors: | Ou, Weixuan, Zheng, Yanzhao, Sun, Shuoshuo, Zhang, Wei, Dong, Baohua, Zhu, Hangcheng, Huang, Ruohui, Yu, Gang, Yan, Pengwei, Qiao, Yifan |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Reason from Future: Reverse Thought Chain Enhances LLM Reasoning
by: Xu, Yinlong, et al.
Published: (2025)
by: Xu, Yinlong, et al.
Published: (2025)
Rubrics to Tokens: Bridging Response-level Rubrics and Token-level Rewards in Instruction Following Tasks
by: Xu, Tianze, et al.
Published: (2026)
by: Xu, Tianze, et al.
Published: (2026)
SkillRouter: Skill Routing for LLM Agents at Scale
by: Zheng, YanZhao, et al.
Published: (2026)
by: Zheng, YanZhao, et al.
Published: (2026)
SERL: A Software Suite for Sample-Efficient Robotic Reinforcement Learning
by: Luo, Jianlan, et al.
Published: (2024)
by: Luo, Jianlan, et al.
Published: (2024)
Conifer: Improving Complex Constrained Instruction-Following Ability of Large Language Models
by: Sun, Haoran, et al.
Published: (2024)
by: Sun, Haoran, et al.
Published: (2024)
Structural Optimization of Lightweight Bipedal Robot via SERL
by: Cheng, Yi, et al.
Published: (2024)
by: Cheng, Yi, et al.
Published: (2024)
Comment on: “Evaluation of Postoperative Anastomotic Patency in Lymphaticovenular Anastomosis Using Photoacoustic Imaging”
by: Ruohui Huang, et al.
Published: (2025)
by: Ruohui Huang, et al.
Published: (2025)
Evaluating the predictors of poor corporal integrity in penile implant recipients
by: Ruohui Huang, et al.
Published: (2025)
by: Ruohui Huang, et al.
Published: (2025)
Insider and stealth trading with dynamic legal risk
by: Qiao, Bixing, et al.
Published: (2026)
by: Qiao, Bixing, et al.
Published: (2026)
CAGS: Open-Vocabulary 3D Scene Understanding with Context-Aware Gaussian Splatting
by: Sun, Wei, et al.
Published: (2025)
by: Sun, Wei, et al.
Published: (2025)
High Strength, Strain, and Resilience of Gold Nanoparticle Reinforced Eutectogels for Multifunctional Sensors
by: Yingxiang Huang, et al.
Published: (2025)
by: Yingxiang Huang, et al.
Published: (2025)
Domain-Conditioned Transformer for Fully Test-time Adaptation
by: Tang, Yushun, et al.
Published: (2024)
by: Tang, Yushun, et al.
Published: (2024)
The Role and Mechanism of Deep Statistical Machine Learning In Biological Target Screening and Immune Microenvironment Regulation of Asthma
by: Zhu, Pengwei
Published: (2025)
by: Zhu, Pengwei
Published: (2025)
Curriculum Sampling: A Two-Phase Curriculum for Efficient Training of Flow Matching
by: Sun, Pengwei
Published: (2026)
by: Sun, Pengwei
Published: (2026)
Investigation of rainfall‐runoff and sediment yield dynamics under varying slope land use patterns in the Three Gorges Reservoir Area of China
by: Xianmeng Meng, et al.
Published: (2024)
by: Xianmeng Meng, et al.
Published: (2024)
SHE: Stepwise Hybrid Examination Reinforcement Learning Framework for E-commerce Search Relevance
by: Jiao, Pengkun, et al.
Published: (2025)
by: Jiao, Pengkun, et al.
Published: (2025)
Early Warning of Low‐Frequency Oscillations in Power System Using Rough Set and Cloud Model
by: Miao Yu, et al.
Published: (2025)
by: Miao Yu, et al.
Published: (2025)
AMHR2 mutation in persistent Müllerian duct syndrome: A case of transverse testicular ectopia
by: Hangcheng Fu, et al.
Published: (2025)
by: Hangcheng Fu, et al.
Published: (2025)
Low-frequency fiber-optic vibration sensing with a Floquet-engineered optical lattice clock
by: Yin, Mojuan, et al.
Published: (2026)
by: Yin, Mojuan, et al.
Published: (2026)
Differentially Private Reinforcement Learning with Self-Play
by: Qiao, Dan, et al.
Published: (2024)
by: Qiao, Dan, et al.
Published: (2024)
Dual-Path Adversarial Lifting for Domain Shift Correction in Online Test-time Adaptation
by: Tang, Yushun, et al.
Published: (2024)
by: Tang, Yushun, et al.
Published: (2024)
Classical symmetry enriched topological orders and distinct monopole charges for dipole-octupole spin ices
by: Zhao, Pengwei, et al.
Published: (2025)
by: Zhao, Pengwei, et al.
Published: (2025)
Z$_2$ topological orders in kagomé dipolar systems: Feedback from Rydberg quantum simulator
by: Zhao, Pengwei, et al.
Published: (2024)
by: Zhao, Pengwei, et al.
Published: (2024)
Inelastic neutron scattering on Ce pyrochlores: Signatures of electric monopoles
by: Zhao, Pengwei, et al.
Published: (2024)
by: Zhao, Pengwei, et al.
Published: (2024)
Domain Expansion and Boundary Growth for Open-Set Single-Source Domain Generalization
by: Jiao, Pengkun, et al.
Published: (2024)
by: Jiao, Pengkun, et al.
Published: (2024)
Strategic Response of News Publishers to Generative AI
by: Zhao, Hangcheng, et al.
Published: (2025)
by: Zhao, Hangcheng, et al.
Published: (2025)
Algorithmic Collusion of Pricing and Advertising on E-commerce Platforms
by: Zhao, Hangcheng, et al.
Published: (2025)
by: Zhao, Hangcheng, et al.
Published: (2025)
SALAD-Bench: A Hierarchical and Comprehensive Safety Benchmark for Large Language Models
by: Li, Lijun, et al.
Published: (2024)
by: Li, Lijun, et al.
Published: (2024)
CO2: Efficient Distributed Training with Full Communication-Computation Overlap
by: Sun, Weigao, et al.
Published: (2024)
by: Sun, Weigao, et al.
Published: (2024)
The Emperor's New Clothes in Benchmarking? A Rigorous Examination of Mitigation Strategies for LLM Benchmark Data Contamination
by: Sun, Yifan, et al.
Published: (2025)
by: Sun, Yifan, et al.
Published: (2025)
From Friction to Fulfillment: Examining When and How Spousal Active‐Destructive Responsiveness to Employees' Sharing of Positive Events Benefits the Work Domain
by: Xin Liu, et al.
Published: (2025)
by: Xin Liu, et al.
Published: (2025)
A leucine‐rich‐repeat receptor‐like kinase SERL1 phosphorylates and stabilizes OsALDH2B1 to promote alkaline tolerance and grain size in rice
by: Zemin Ma, et al.
Published: (2026)
by: Zemin Ma, et al.
Published: (2026)
Enhancing Large Language Models with Domain-Specific Knowledge: The Case in Topological Materials
by: Xu, HuangChao, et al.
Published: (2024)
by: Xu, HuangChao, et al.
Published: (2024)
The High Cost of Inexpensive Medicines
by: George Selali Blewusi, et al.
Published: (2026)
by: George Selali Blewusi, et al.
Published: (2026)
DECEIVE-AFC: Adversarial Claim Attacks against Search-Enabled LLM-based Fact-Checking Systems
by: Ou, Haoran, et al.
Published: (2026)
by: Ou, Haoran, et al.
Published: (2026)
Engineering topological states and quantum-inspired information processing using classical circuits
by: Chen, Tian, et al.
Published: (2024)
by: Chen, Tian, et al.
Published: (2024)
Engineering Topological States and Quantum‐Inspired Information Processing Using Classical Circuits
by: Tian Chen, et al.
Published: (2025)
by: Tian Chen, et al.
Published: (2025)
Physical Prompt Injection Attacks on Large Vision-Language Models
by: Ling, Chen, et al.
Published: (2026)
by: Ling, Chen, et al.
Published: (2026)
Near-Optimal Reinforcement Learning with Self-Play under Adaptivity Constraints
by: Qiao, Dan, et al.
Published: (2024)
by: Qiao, Dan, et al.
Published: (2024)
Towards Multimodal Open-Set Domain Generalization and Adaptation through Self-supervision
by: Dong, Hao, et al.
Published: (2024)
by: Dong, Hao, et al.
Published: (2024)
Similar Items
-
Reason from Future: Reverse Thought Chain Enhances LLM Reasoning
by: Xu, Yinlong, et al.
Published: (2025) -
Rubrics to Tokens: Bridging Response-level Rubrics and Token-level Rewards in Instruction Following Tasks
by: Xu, Tianze, et al.
Published: (2026) -
SkillRouter: Skill Routing for LLM Agents at Scale
by: Zheng, YanZhao, et al.
Published: (2026) -
SERL: A Software Suite for Sample-Efficient Robotic Reinforcement Learning
by: Luo, Jianlan, et al.
Published: (2024) -
Conifer: Improving Complex Constrained Instruction-Following Ability of Large Language Models
by: Sun, Haoran, et al.
Published: (2024)