Length-Unbiased Sequence Policy Optimization: Revealing and Controlling Response Length Variation in RLVR
Fuente:
arXiv
Saved in:
| Main Authors: | Liu, Fanfan, Yin, Youyang, Shi, Peng, Yang, Siqi, Zeng, Zhixiong, Qiu, Haibo |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Flexible Entropy Control in RLVR with a Gradient-Preserving Perspective
by: Chen, Kun, et al.
Published: (2026)
by: Chen, Kun, et al.
Published: (2026)
Context Cascade Compression: Exploring the Upper Limits of Text Compression
by: Liu, Fanfan, et al.
Published: (2025)
by: Liu, Fanfan, et al.
Published: (2025)
Metis-SPECS: Decoupling Multimodal Learning via Self-distilled Preference-based Cold Start
by: Chen, Kun, et al.
Published: (2025)
by: Chen, Kun, et al.
Published: (2025)
Learning When to Look: A Disentangled Curriculum for Strategic Perception in Multimodal Reasoning
by: Yang, Siqi, et al.
Published: (2025)
by: Yang, Siqi, et al.
Published: (2025)
What is the Best Sequence Length for BABYLM?
by: Salhan, Suchir, et al.
Published: (2025)
by: Salhan, Suchir, et al.
Published: (2025)
LAPO: Internalizing Reasoning Efficiency via Length-Adaptive Policy Optimization
by: Wu, Xingyu, et al.
Published: (2025)
by: Wu, Xingyu, et al.
Published: (2025)
Length Desensitization in Direct Preference Optimization
by: Liu, Wei, et al.
Published: (2024)
by: Liu, Wei, et al.
Published: (2024)
ReCUT: Balancing Reasoning Length and Accuracy in LLMs via Stepwise Trails and Preference Optimization
by: Jin, Zhensheng, et al.
Published: (2025)
by: Jin, Zhensheng, et al.
Published: (2025)
Length Value Model: Scalable Value Pretraining for Token-Level Length Modeling
by: Zhang, Zhen, et al.
Published: (2026)
by: Zhang, Zhen, et al.
Published: (2026)
Parallelizing Linear Transformers with the Delta Rule over Sequence Length
by: Yang, Songlin, et al.
Published: (2024)
by: Yang, Songlin, et al.
Published: (2024)
LSPO: Length-aware Dynamic Sampling for Policy Optimization in LLM Reasoning
by: Chen, Weizhe, et al.
Published: (2025)
by: Chen, Weizhe, et al.
Published: (2025)
Length-Controlled Margin-Based Preference Optimization without Reference Model
by: Li, Gengxu, et al.
Published: (2025)
by: Li, Gengxu, et al.
Published: (2025)
Improving Variable-Length Generation in Diffusion Language Models via Length Regularization
by: Cheng, Zicong, et al.
Published: (2026)
by: Cheng, Zicong, et al.
Published: (2026)
Efficient Pretraining Length Scaling
by: Wu, Bohong, et al.
Published: (2025)
by: Wu, Bohong, et al.
Published: (2025)
Beyond Fixed Length: Bucket Pre-training is All You Need
by: Yang, Qing, et al.
Published: (2024)
by: Yang, Qing, et al.
Published: (2024)
Clip Your Sequences Fairly: Enforcing Length Fairness for Sequence-Level RL
by: Mao, Hanyi, et al.
Published: (2025)
by: Mao, Hanyi, et al.
Published: (2025)
Reading or Reasoning? Format Decoupled Reinforcement Learning for Document OCR
by: Zhong, Yufeng, et al.
Published: (2025)
by: Zhong, Yufeng, et al.
Published: (2025)
Metis-HOME: Hybrid Optimized Mixture-of-Experts for Multimodal Reasoning
by: Lan, Xiaohan, et al.
Published: (2025)
by: Lan, Xiaohan, et al.
Published: (2025)
Uniform-Correct Policy Optimization: Breaking RLVR's Indifference to Diversity
by: Lochab, Anamika, et al.
Published: (2026)
by: Lochab, Anamika, et al.
Published: (2026)
Length Controlled Generation for Black-box LLMs
by: Gu, Yuxuan, et al.
Published: (2024)
by: Gu, Yuxuan, et al.
Published: (2024)
Zero-Shot Strategies for Length-Controllable Summarization
by: Retkowski, Fabian, et al.
Published: (2024)
by: Retkowski, Fabian, et al.
Published: (2024)
An Empirical Study on the Characteristics of Bias upon Context Length Variation for Bangla
by: Sadhu, Jayanta, et al.
Published: (2024)
by: Sadhu, Jayanta, et al.
Published: (2024)
Efficient Interactive LLM Serving with Proxy Model-based Sequence Length Prediction
by: Qiu, Haoran, et al.
Published: (2024)
by: Qiu, Haoran, et al.
Published: (2024)
CLIPO: Contrastive Learning in Policy Optimization Generalizes RLVR
by: Cui, Sijia, et al.
Published: (2026)
by: Cui, Sijia, et al.
Published: (2026)
Gecko: An Efficient Neural Architecture Inherently Processing Sequences with Arbitrary Lengths
by: Ma, Xuezhe, et al.
Published: (2026)
by: Ma, Xuezhe, et al.
Published: (2026)
Prompt-Based Length Controlled Generation with Multiple Control Types
by: Jie, Renlong, et al.
Published: (2024)
by: Jie, Renlong, et al.
Published: (2024)
Prompting LLMs: Length Control for Isometric Machine Translation
by: Javorský, Dávid, et al.
Published: (2025)
by: Javorský, Dávid, et al.
Published: (2025)
Rethinking Perplexity: Revealing the Impact of Input Length on Perplexity Evaluation in LLMs
by: Cheng, Letian, et al.
Published: (2026)
by: Cheng, Letian, et al.
Published: (2026)
Investigating Length Issues in Document-level Machine Translation
by: Peng, Ziqian, et al.
Published: (2024)
by: Peng, Ziqian, et al.
Published: (2024)
Dataset Decomposition: Faster LLM Training with Variable Sequence Length Curriculum
by: Pouransari, Hadi, et al.
Published: (2024)
by: Pouransari, Hadi, et al.
Published: (2024)
Diverse, not Short: A Length-Controlled Data Selection Strategy for Improving Response Diversity of Language Models
by: Deshpande, Vijeta, et al.
Published: (2025)
by: Deshpande, Vijeta, et al.
Published: (2025)
Optimizing Length Compression in Large Reasoning Models
by: Cheng, Zhengxiang, et al.
Published: (2025)
by: Cheng, Zhengxiang, et al.
Published: (2025)
Eliminating Biased Length Reliance of Direct Preference Optimization via Down-Sampled KL Divergence
by: Lu, Junru, et al.
Published: (2024)
by: Lu, Junru, et al.
Published: (2024)
Recycling Failures: Salvaging Exploration in RLVR via Fine-Grained Off-Policy Guidance
by: Ren, Yanwei, et al.
Published: (2026)
by: Ren, Yanwei, et al.
Published: (2026)
Ruler: A Model-Agnostic Method to Control Generated Length for Large Language Models
by: Li, Jiaming, et al.
Published: (2024)
by: Li, Jiaming, et al.
Published: (2024)
Provable Length Generalization in Sequence Prediction via Spectral Filtering
by: Marsden, Annie, et al.
Published: (2024)
by: Marsden, Annie, et al.
Published: (2024)
Intrinsic Entropy of Context Length Scaling in LLMs
by: Shi, Jingzhe, et al.
Published: (2025)
by: Shi, Jingzhe, et al.
Published: (2025)
CEPO: RLVR Self-Distillation using Contrastive Evidence Policy Optimization
by: Heakl, Ahmed, et al.
Published: (2026)
by: Heakl, Ahmed, et al.
Published: (2026)
Optimizing Speech-Input Length for Speaker-Independent Depression Classification
by: Rutowski, Tomasz, et al.
Published: (2024)
by: Rutowski, Tomasz, et al.
Published: (2024)
LongStory: Coherent, Complete and Length Controlled Long story Generation
by: Park, Kyeongman, et al.
Published: (2023)
by: Park, Kyeongman, et al.
Published: (2023)
Similar Items
-
Flexible Entropy Control in RLVR with a Gradient-Preserving Perspective
by: Chen, Kun, et al.
Published: (2026) -
Context Cascade Compression: Exploring the Upper Limits of Text Compression
by: Liu, Fanfan, et al.
Published: (2025) -
Metis-SPECS: Decoupling Multimodal Learning via Self-distilled Preference-based Cold Start
by: Chen, Kun, et al.
Published: (2025) -
Learning When to Look: A Disentangled Curriculum for Strategic Perception in Multimodal Reasoning
by: Yang, Siqi, et al.
Published: (2025) -
What is the Best Sequence Length for BABYLM?
by: Salhan, Suchir, et al.
Published: (2025)