Rethinking the Comparison Unit in Sequence-Level Reinforcement Learning: An Equal-Length Paired Training Framework from Loss Correction to Sample Construction
Fuente:
arXiv
Saved in:
| Main Authors: | Ding, Fei, Zhang, Yongkang, Liu, Runhao, Liao, Yuhao, Zeng, Zijian, Yang, Huiming, wang, Sibo, Liao, Linglin |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Internalizing Outcome Supervision into Process Supervision: A New Paradigm for Reinforcement Learning for Reasoning
by: Ding, Fei, et al.
Published: (2026)
by: Ding, Fei, et al.
Published: (2026)
PoseStreamer: A Multi-modal Framework for 3D Tracking of Unseen Moving Objects
by: Yang, Huiming, et al.
Published: (2025)
by: Yang, Huiming, et al.
Published: (2025)
Difference Feedback: Generating Multimodal Process-Level Supervision for VLM Reinforcement Learning
by: Feiding, et al.
Published: (2026)
by: Feiding, et al.
Published: (2026)
Design Conditions for Intra-Group Learning of Sequence-Level Rewards: Token Gradient Cancellation
by: Ding, Fei, et al.
Published: (2026)
by: Ding, Fei, et al.
Published: (2026)
DexSim2Real: Foundation Model-Guided Sim-to-Real Transfer for Generalizable Dexterous Manipulation
by: Zeng, Zijian, et al.
Published: (2026)
by: Zeng, Zijian, et al.
Published: (2026)
Spatial-aware Symmetric Alignment for Text-guided Medical Image Segmentation
by: Liao, Linglin, et al.
Published: (2025)
by: Liao, Linglin, et al.
Published: (2025)
HELM: Harness-Enhanced Long-horizon Memory for Vision-Language-Action Manipulation
by: Zeng, Zijian, et al.
Published: (2026)
by: Zeng, Zijian, et al.
Published: (2026)
Reducing Credit Assignment Variance via Counterfactual Reasoning Paths
by: Ding, Fei, et al.
Published: (2026)
by: Ding, Fei, et al.
Published: (2026)
Not All LLM-Generated Data Are Equal: Rethinking Data Weighting in Text Classification
by: Kuo, Hsun-Yu, et al.
Published: (2024)
by: Kuo, Hsun-Yu, et al.
Published: (2024)
Hyperon Pair Production at BESIII
by: Shang, Zijie, et al.
Published: (2025)
by: Shang, Zijie, et al.
Published: (2025)
Balancing the Reasoning Load: Difficulty-Differentiated Policy Optimization with Length Redistribution for Efficient and Robust Reinforcement Learning
by: Xia, Yinan, et al.
Published: (2026)
by: Xia, Yinan, et al.
Published: (2026)
Sequence-Level Unsupervised Training in Speech Recognition: A Theoretical Study
by: Yang, Zijian, et al.
Published: (2026)
by: Yang, Zijian, et al.
Published: (2026)
Rethinking External Slow-Thinking: From Snowball Errors to Probability of Correct Reasoning
by: Gan, Zeyu, et al.
Published: (2025)
by: Gan, Zeyu, et al.
Published: (2025)
ShareDP: Finding k Disjoint Paths for Multiple Vertex Pairs
by: Yuan, Zhiqiu, et al.
Published: (2025)
by: Yuan, Zhiqiu, et al.
Published: (2025)
Memory Sequence Length of Data Sampling Impacts the Adaptation of Meta-Reinforcement Learning Agents
by: Zhang, Menglong, et al.
Published: (2024)
by: Zhang, Menglong, et al.
Published: (2024)
The equivalent condition for GRL codes to be MDS, AMDS or self-dual
by: Liang, Zhonghao, et al.
Published: (2025)
by: Liang, Zhonghao, et al.
Published: (2025)
The asymptotic estimation for two classes of generalized Fibonacci sub-sequences
by: Wan, Yongkang, et al.
Published: (2025)
by: Wan, Yongkang, et al.
Published: (2025)
The inverse of the (alternating) infinite sum of the reciprocal of the weighted sum for generalized Fibonacci sub-sequences
by: Wan, Yongkang, et al.
Published: (2025)
by: Wan, Yongkang, et al.
Published: (2025)
Not All Negative Samples Are Equal: LLMs Learn Better from Plausible Reasoning
by: Di, Zixiang, et al.
Published: (2026)
by: Di, Zixiang, et al.
Published: (2026)
On Duplication-Free Codes for Disjoint or Equal-Length Errors
by: Yu, Wenjun, et al.
Published: (2024)
by: Yu, Wenjun, et al.
Published: (2024)
Clip Your Sequences Fairly: Enforcing Length Fairness for Sequence-Level RL
by: Mao, Hanyi, et al.
Published: (2025)
by: Mao, Hanyi, et al.
Published: (2025)
Multi-Layer GRPO: Enhancing Reasoning and Self-Correction in Large Language Models
by: Ding, Fei, et al.
Published: (2025)
by: Ding, Fei, et al.
Published: (2025)
Constructing $k$-ary Orientable Sequences with Asymptotically Optimal Length
by: Gabrić, Daniel, et al.
Published: (2024)
by: Gabrić, Daniel, et al.
Published: (2024)
Broad Band Equal-Length And Equal-Width Substrate Integrated Waveguide Four Channel Power Divider
by: Masoud Khoubroo Eslamloo
Published: (2015)
by: Masoud Khoubroo Eslamloo
Published: (2015)
Correction to ‘Rethinking Dopamine‐Guided Action Sequence Learning’
Published: (2025)
Published: (2025)
sprofing the DLSCA
by: wang
Published: (2025)
by: wang
Published: (2025)
Touch to Pair: Secure and Usable IoT Pairing without Information Loss
by: Wu, Chuxiong, et al.
Published: (2024)
by: Wu, Chuxiong, et al.
Published: (2024)
Adversarial Samples Are Not Created Equal
by: Crawford, Jennifer, et al.
Published: (2026)
by: Crawford, Jennifer, et al.
Published: (2026)
AI for Equitable Tennis Training: Leveraging AI for Equitable and Accurate Classification of Tennis Skill Levels and Training Phases
by: Gao, Gyanna, et al.
Published: (2024)
by: Gao, Gyanna, et al.
Published: (2024)
Conversation Disentanglement with Bi-Level Contrastive Learning
by: Huang, Chengyu, et al.
Published: (2022)
by: Huang, Chengyu, et al.
Published: (2022)
Box-Level Class-Balanced Sampling for Active Object Detection
by: Liao, Jingyi, et al.
Published: (2025)
by: Liao, Jingyi, et al.
Published: (2025)
From Production Envelopes to Executable Schedules: Sound Constructive Refinement for High-Mix Manufacturing
by: Liu, Runhao, et al.
Published: (2025)
by: Liu, Runhao, et al.
Published: (2025)
GDEPO: Group Dual-dynamic and Equal-right Advantage Policy Optimization with Enhanced Training Data Utilization for Sample-Constrained Reinforcement Learning
by: Yan, Zhengqing, et al.
Published: (2026)
by: Yan, Zhengqing, et al.
Published: (2026)
Artificial Intelligence in Power System Security and Stability Analysis: A Comprehensive Review
by: Zhang, Runhao
Published: (2024)
by: Zhang, Runhao
Published: (2024)
Selective Matching Losses -- Not All Scores Are Created Equal
by: Shamir, Gil I., et al.
Published: (2025)
by: Shamir, Gil I., et al.
Published: (2025)
A Fast Algorithm for Scheduling Equal-Length Jobs on Identical Machines
by: Nodari Vakhania
Published: (1998)
by: Nodari Vakhania
Published: (1998)
Length-Unbiased Sequence Policy Optimization: Revealing and Controlling Response Length Variation in RLVR
by: Liu, Fanfan, et al.
Published: (2026)
by: Liu, Fanfan, et al.
Published: (2026)
Rethinking Model Evaluation as Narrowing the Socio-Technical Gap
by: Liao, Q. Vera, et al.
Published: (2023)
by: Liao, Q. Vera, et al.
Published: (2023)
Power-SMC: Low-Latency Sequence-Level Power Sampling for Training-Free LLM Reasoning
by: Azizi, Seyedarmin, et al.
Published: (2026)
by: Azizi, Seyedarmin, et al.
Published: (2026)
Dataset Decomposition: Faster LLM Training with Variable Sequence Length Curriculum
by: Pouransari, Hadi, et al.
Published: (2024)
by: Pouransari, Hadi, et al.
Published: (2024)
Similar Items
-
Internalizing Outcome Supervision into Process Supervision: A New Paradigm for Reinforcement Learning for Reasoning
by: Ding, Fei, et al.
Published: (2026) -
PoseStreamer: A Multi-modal Framework for 3D Tracking of Unseen Moving Objects
by: Yang, Huiming, et al.
Published: (2025) -
Difference Feedback: Generating Multimodal Process-Level Supervision for VLM Reinforcement Learning
by: Feiding, et al.
Published: (2026) -
Design Conditions for Intra-Group Learning of Sequence-Level Rewards: Token Gradient Cancellation
by: Ding, Fei, et al.
Published: (2026) -
DexSim2Real: Foundation Model-Guided Sim-to-Real Transfer for Generalizable Dexterous Manipulation
by: Zeng, Zijian, et al.
Published: (2026)