Saved in:
| Main Authors: | Zhang, Xiaojiang, Wang, Jinghui, Cheng, Zifei, Zhuang, Wenhao, Lin, Zheng, Zhang, Minglei, Wang, Shaojie, Cui, Yinghan, Wang, Chao, Peng, Junyi, Jiang, Shimiao, Kuang, Shiqi, Yin, Shouyu, Wen, Chaohang, Zhang, Haotian, Chen, Bin, Yu, Bing |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2504.14286 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Tree Training: Accelerating Agentic LLMs Training via Shared Prefix Reuse
by: Wang, Jinghui, et al.
Published: (2025)
by: Wang, Jinghui, et al.
Published: (2025)
SeamlessFlow: A Trainer Agent Isolation RL Framework Achieving Bubble-Free Pipelines via Tag Scheduling
by: Wang, Jinghui, et al.
Published: (2025)
by: Wang, Jinghui, et al.
Published: (2025)
SRPO: Enhancing Multimodal LLM Reasoning via Reflection-Aware Reinforcement Learning
by: Wan, Zhongwei, et al.
Published: (2025)
by: Wan, Zhongwei, et al.
Published: (2025)
SRPO: Self-Referential Policy Optimization for Vision-Language-Action Models
by: Fei, Senyu, et al.
Published: (2025)
by: Fei, Senyu, et al.
Published: (2025)
KAT-Coder Technical Report
by: Zhan, Zizheng, et al.
Published: (2025)
by: Zhan, Zizheng, et al.
Published: (2025)
A Practical Introduction to Deep Reinforcement Learning
by: Sun, Yinghan, et al.
Published: (2025)
by: Sun, Yinghan, et al.
Published: (2025)
Amorphous Sulfo‐Halide Solid Electrolytes With Enhanced Anion Dynamics for Highly Stable All‐Solid‐State Sodium Batteries
by: Zifei Shi, et al.
Published: (2026)
by: Zifei Shi, et al.
Published: (2026)
Unravelling the Oxygen Dynamics at Gas‐Solid Interface of CaGd 2 (MoO 4 ) 4 :Yb,Er Scheelite Ceramics via Upconversion Luminescence
by: Yinghan Wang, et al.
Published: (2025)
by: Yinghan Wang, et al.
Published: (2025)
Effect of the cycle number of multistep sintering on the properties of FeSe 0.4 Te 0.6 superconductors
by: Chaohang Miao, et al.
Published: (2024)
by: Chaohang Miao, et al.
Published: (2024)
$Δl =1$ coupling of single-particle orbitals in octupole deformed nuclei
by: Wang, XuDong, et al.
Published: (2026)
by: Wang, XuDong, et al.
Published: (2026)
Large and Small Yellow Croakers Feeding and Living Together Make Large Yellow Croaker Population Recovery Difficult: A Guild Perspective.
by: Cai, Pengyu, et al.
Published: (2024)
by: Cai, Pengyu, et al.
Published: (2024)
Two Birds with One Stone: Enhancing Uncertainty Quantification and Interpretability with Graph Functional Neural Process
by: Kong, Lingkai, et al.
Published: (2025)
by: Kong, Lingkai, et al.
Published: (2025)
Exploiting sparse structures and synergy designs to advance situational awareness of electrical power grid
by: Li, Shimiao
Published: (2024)
by: Li, Shimiao
Published: (2024)
Reliable Intelligent Diagnostics With Uncertainty Quantification for Mechanical System Condition Assessment
by: Guangjun Jiang, et al.
Published: (2026)
by: Guangjun Jiang, et al.
Published: (2026)
Mechanical Experiments and ABAQUS Simulations of Fiber–Coal Gangue Stabilized Expansive Soil Under Freeze–Thaw Cycles
by: Yan Zhang, et al.
Published: (2024)
by: Yan Zhang, et al.
Published: (2024)
A good neighbor, a found treasure: Do local neighbors affect corporate innovation?
by: Shaopeng Cao, et al.
Published: (2024)
by: Shaopeng Cao, et al.
Published: (2024)
Modified Arthroscopic Technique for Repair of Medial Meniscus Posterior Root Tear and Centralization of the Extruded Medial Meniscus Through a Double Transtibial Tunnel (Without Use of Accessory Ports)
by: Parvind Grish Dev Sookun, et al.
Published: (2025)
by: Parvind Grish Dev Sookun, et al.
Published: (2025)
From Implicit to Explicit: Token-Efficient Logical Supervision for Mathematical Reasoning in LLMs
by: Wang, Shaojie, et al.
Published: (2026)
by: Wang, Shaojie, et al.
Published: (2026)
From Meta-Thought to Execution: Cognitively Aligned Post-Training for Generalizable and Reliable LLM Reasoning
by: Wang, Shaojie, et al.
Published: (2026)
by: Wang, Shaojie, et al.
Published: (2026)
A Camera‐Like Dual‐Defocus Curvature Wavefront Sensor With GPU Acceleration for Real‐Time Quantitative Phase Imaging
by: Wei Wang, et al.
Published: (2025)
by: Wei Wang, et al.
Published: (2025)
Do Institutional Cross‐Owners Obstruct Corporate Environmental Information Disclosure? Evidence From China
by: Xin Cui, et al.
Published: (2025)
by: Xin Cui, et al.
Published: (2025)
BTL-UI: Blink-Think-Link Reasoning Model for GUI Agent
by: Zhang, Shaojie, et al.
Published: (2025)
by: Zhang, Shaojie, et al.
Published: (2025)
The Impact of Quantization on Large Reasoning Model Reinforcement Learning
by: Kumar, Medha, et al.
Published: (2025)
by: Kumar, Medha, et al.
Published: (2025)
On the Dirichlet Problem at Infinity and Poisson Boundary for Certain Manifolds without Conjugate Points
by: Liu, Fei, et al.
Published: (2025)
by: Liu, Fei, et al.
Published: (2025)
Dispersion-Domain Detection for Mobile Molecular Communication Under Multiplicative Geometry Uncertainty
by: Zhang, Shaojie, et al.
Published: (2026)
by: Zhang, Shaojie, et al.
Published: (2026)
OASIS: Conditional Distribution Shaping for Offline Safe Reinforcement Learning
by: Yao, Yihang, et al.
Published: (2024)
by: Yao, Yihang, et al.
Published: (2024)
Credit Where It is Due: Cross-Modality Connectivity Drives Precise Reinforcement Learning for MLLM Reasoning
by: Jiao, Zhengbo, et al.
Published: (2026)
by: Jiao, Zhengbo, et al.
Published: (2026)
Balancing the Reasoning Load: Difficulty-Differentiated Policy Optimization with Length Redistribution for Efficient and Robust Reinforcement Learning
by: Xia, Yinan, et al.
Published: (2026)
by: Xia, Yinan, et al.
Published: (2026)
Microquasar Remnants as Pevatrons Illuminating the Galactic Cosmic Ray Knee
by: Zhang, Bing Theodore, et al.
Published: (2026)
by: Zhang, Bing Theodore, et al.
Published: (2026)
A Unified Framework for 10 TeV to EeV Diffuse Neutrino Sky and KM3-230213A
by: Yu, Shiqi, et al.
Published: (2026)
by: Yu, Shiqi, et al.
Published: (2026)
Multi-Messenger Modeling of Low-Luminosity Gamma-Ray Bursts
by: Yu, Shiqi, et al.
Published: (2026)
by: Yu, Shiqi, et al.
Published: (2026)
Divide-and-Conquer Decoupled Network for Cross-Domain Few-Shot Segmentation
by: Cong, Runmin, et al.
Published: (2025)
by: Cong, Runmin, et al.
Published: (2025)
Preferences for Processing Facial Information at Different Orientations in Individuals With Developmental Prosopagnosia
by: Jialin Ma, et al.
Published: (2026)
by: Jialin Ma, et al.
Published: (2026)
Discrepancies in Pedestrian Crossing of Static vs. Dynamic Crowds: An Experimental Study
by: Wang, Jinghui, et al.
Published: (2024)
by: Wang, Jinghui, et al.
Published: (2024)
Improving LLM Code Generation via Requirement-Aware Curriculum Reinforcement Learning
by: Yin, Shouyu, et al.
Published: (2026)
by: Yin, Shouyu, et al.
Published: (2026)
Enhancing Video Music Recommendation with Transformer-Driven Audio-Visual Embeddings
by: Liu, Shimiao, et al.
Published: (2025)
by: Liu, Shimiao, et al.
Published: (2025)
Exploring Task-Solving Paradigm for Generalized Cross-Domain Face Anti-Spoofing via Reinforcement Fine-Tuning
by: Jiang, Fangling, et al.
Published: (2025)
by: Jiang, Fangling, et al.
Published: (2025)
Prototype-Regularized Federated Learning for Cross-Domain Aspect Sentiment Triplet Extraction
by: Cai, Zongming, et al.
Published: (2026)
by: Cai, Zongming, et al.
Published: (2026)
Towards Monotonic Improvement in In-Context Reinforcement Learning
by: Zhang, Wenhao, et al.
Published: (2025)
by: Zhang, Wenhao, et al.
Published: (2025)
Impact of Radio Frequency Power on Columnar and Filamentary Modes in Atmospheric Pressure Very Low Frequency Plasma within Pores
by: Wang, Haozhe, et al.
Published: (2025)
by: Wang, Haozhe, et al.
Published: (2025)
Similar Items
-
Tree Training: Accelerating Agentic LLMs Training via Shared Prefix Reuse
by: Wang, Jinghui, et al.
Published: (2025) -
SeamlessFlow: A Trainer Agent Isolation RL Framework Achieving Bubble-Free Pipelines via Tag Scheduling
by: Wang, Jinghui, et al.
Published: (2025) -
SRPO: Enhancing Multimodal LLM Reasoning via Reflection-Aware Reinforcement Learning
by: Wan, Zhongwei, et al.
Published: (2025) -
SRPO: Self-Referential Policy Optimization for Vision-Language-Action Models
by: Fei, Senyu, et al.
Published: (2025) -
KAT-Coder Technical Report
by: Zhan, Zizheng, et al.
Published: (2025)