Beyond Token-level Supervision: Unlocking the Potential of Decoding-based Regression via Reinforcement Learning
Fuente:
arXiv
Saved in:
| Main Authors: | Chen, Ming, Tang, Sheng, Tan, Rong-Xi, Li, Ziniu, Chen, Jiacheng, Xue, Ke, Qian, Chao |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Towards Universal Offline Black-Box Optimization via Learning Language Model Embeddings
by: Tan, Rong-Xi, et al.
Published: (2025)
by: Tan, Rong-Xi, et al.
Published: (2025)
Supervised Fine-Tuning Needs to Unlock the Potential of Token Priority
by: Shen, Zhanming, et al.
Published: (2026)
by: Shen, Zhanming, et al.
Published: (2026)
Offline Multi-Objective Optimization
by: Xue, Ke, et al.
Published: (2024)
by: Xue, Ke, et al.
Published: (2024)
Offline Model-Based Optimization by Learning to Rank
by: Tan, Rong-Xi, et al.
Published: (2024)
by: Tan, Rong-Xi, et al.
Published: (2024)
Zero Token-Driven Deep Thinking in LLMs: Unlocking the Full Potential of Existing Parameters via Cyclic Refinement
by: Li, Guanghao, et al.
Published: (2025)
by: Li, Guanghao, et al.
Published: (2025)
Reinforcement Learning Policy as Macro Regulator Rather than Macro Placer
by: Xue, Ke, et al.
Published: (2024)
by: Xue, Ke, et al.
Published: (2024)
Revisiting End-to-End Learning with Slide-level Supervision in Computational Pathology
by: Tang, Wenhao, et al.
Published: (2025)
by: Tang, Wenhao, et al.
Published: (2025)
Beyond Generation: Unlocking Universal Editing via Self-Supervised Fine-Tuning
by: Chen, Harold Haodong, et al.
Published: (2024)
by: Chen, Harold Haodong, et al.
Published: (2024)
Token Constraint Decoding Improves Robustness on Question Answering for Large Language Models
by: Yao, Jui-Ming, et al.
Published: (2025)
by: Yao, Jui-Ming, et al.
Published: (2025)
Quality-Diversity Red-Teaming: Automated Generation of High-Quality and Diverse Attackers for Large Language Models
by: Wang, Ren-Jian, et al.
Published: (2025)
by: Wang, Ren-Jian, et al.
Published: (2025)
BBOPlace-Bench: Benchmarking Black-Box Optimization for Chip Placement
by: Xue, Ke, et al.
Published: (2025)
by: Xue, Ke, et al.
Published: (2025)
Optimized Multi-Token Joint Decoding with Auxiliary Model for LLM Inference
by: Qin, Zongyue, et al.
Published: (2024)
by: Qin, Zongyue, et al.
Published: (2024)
Unlocking Zero-shot Potential of Semi-dense Image Matching via Gaussian Splatting
by: Chen, Juncheng, et al.
Published: (2025)
by: Chen, Juncheng, et al.
Published: (2025)
Token-Guard: Towards Token-Level Hallucination Control via Self-Checking Decoding
by: Zhu, Yifan, et al.
Published: (2026)
by: Zhu, Yifan, et al.
Published: (2026)
Decomposing the Entropy-Performance Exchange: The Missing Keys to Unlocking Effective Reinforcement Learning
by: Deng, Jia, et al.
Published: (2025)
by: Deng, Jia, et al.
Published: (2025)
SEER: Facilitating Structured Reasoning and Explanation via Reinforcement Learning
by: Chen, Guoxin, et al.
Published: (2024)
by: Chen, Guoxin, et al.
Published: (2024)
Dataset Distillation for Offline Reinforcement Learning
by: Light, Jonathan, et al.
Published: (2024)
by: Light, Jonathan, et al.
Published: (2024)
Revisiting Entropy Regularization: Adaptive Coefficient Unlocks Its Potential for LLM Reinforcement Learning
by: Zhang, Xiaoyun, et al.
Published: (2025)
by: Zhang, Xiaoyun, et al.
Published: (2025)
Task Assignment and Exploration Optimization for Low Altitude UAV Rescue via Generative AI Enhanced Multi-agent Reinforcement Learning
by: Tang, Xin, et al.
Published: (2025)
by: Tang, Xin, et al.
Published: (2025)
Physically Interpretable Interatomic Potentials via Symbolic Regression and Reinforcement Learning
by: Varughese, Bilvin, et al.
Published: (2025)
by: Varughese, Bilvin, et al.
Published: (2025)
Adaptive Ability Decomposing for Unlocking Large Reasoning Model Effective Reinforcement Learning
by: Chen, Zhipeng, et al.
Published: (2026)
by: Chen, Zhipeng, et al.
Published: (2026)
Unlock the Correlation between Supervised Fine-Tuning and Reinforcement Learning in Training Code Large Language Models
by: Chen, Jie, et al.
Published: (2024)
by: Chen, Jie, et al.
Published: (2024)
Unlocking Potential Catalysts: A Machine Learning Approach with Bayesian and Regression Models
by: Chandra Chowdhury
Published: (2024)
by: Chandra Chowdhury
Published: (2024)
D2C: Unlocking the Potential of Continuous Autoregressive Image Generation with Discrete Tokens
by: Wang, Panpan, et al.
Published: (2025)
by: Wang, Panpan, et al.
Published: (2025)
Beyond Reasoning: Reinforcement Learning Unlocks Parametric Knowledge in LLMs
by: Yang, Wanli, et al.
Published: (2026)
by: Yang, Wanli, et al.
Published: (2026)
Knapsack RL: Unlocking Exploration of LLMs via Optimizing Budget Allocation
by: Li, Ziniu, et al.
Published: (2025)
by: Li, Ziniu, et al.
Published: (2025)
Network Sparsity Unlocks the Scaling Potential of Deep Reinforcement Learning
by: Ma, Guozheng, et al.
Published: (2025)
by: Ma, Guozheng, et al.
Published: (2025)
Dual-Decoder Consistency via Pseudo-Labels Guided Data Augmentation for Semi-Supervised Medical Image Segmentation
by: Chen, Yuanbin, et al.
Published: (2023)
by: Chen, Yuanbin, et al.
Published: (2023)
Decoding-based Regression
by: Song, Xingyou, et al.
Published: (2025)
by: Song, Xingyou, et al.
Published: (2025)
AtomicVLA: Unlocking the Potential of Atomic Skill Learning in Robots
by: Zhang, Likui, et al.
Published: (2026)
by: Zhang, Likui, et al.
Published: (2026)
Reinforcement Learning with Token-level Feedback for Controllable Text Generation
by: Li, Wendi, et al.
Published: (2024)
by: Li, Wendi, et al.
Published: (2024)
Unlocking Telemetry Potential: Self-Supervised Learning for Continuous Clinical Electrocardiogram Monitoring
by: Kite, Thomas, et al.
Published: (2024)
by: Kite, Thomas, et al.
Published: (2024)
DiffuSpec: Unlocking Diffusion Language Models for Speculative Decoding
by: Li, Guanghao, et al.
Published: (2025)
by: Li, Guanghao, et al.
Published: (2025)
On the Learnability of Offline Model-Based Optimization: A Ranking Perspective
by: Lyu, Shen-Huan, et al.
Published: (2026)
by: Lyu, Shen-Huan, et al.
Published: (2026)
A Note on Hybrid Online Reinforcement and Imitation Learning for LLMs: Formulations and Algorithms
by: Li, Yingru, et al.
Published: (2025)
by: Li, Yingru, et al.
Published: (2025)
Rec-R1: Bridging Generative Large Language Models and User-Centric Recommendation Systems via Reinforcement Learning
by: Lin, Jiacheng, et al.
Published: (2025)
by: Lin, Jiacheng, et al.
Published: (2025)
Generalized Category Discovery via Token Manifold Capacity Learning
by: Tang, Luyao, et al.
Published: (2025)
by: Tang, Luyao, et al.
Published: (2025)
Multifunctional Separator Engineering: Unlocking the Potential of Aqueous Zinc‐Ion Batteries
by: Chen Qian, et al.
Published: (2025)
by: Chen Qian, et al.
Published: (2025)
Think Natively: Unlocking Multilingual Reasoning with Consistency-Enhanced Reinforcement Learning
by: Zhang, Xue, et al.
Published: (2025)
by: Zhang, Xue, et al.
Published: (2025)
FoldToken: Learning Protein Language via Vector Quantization and Beyond
by: Gao, Zhangyang, et al.
Published: (2024)
by: Gao, Zhangyang, et al.
Published: (2024)
Similar Items
-
Towards Universal Offline Black-Box Optimization via Learning Language Model Embeddings
by: Tan, Rong-Xi, et al.
Published: (2025) -
Supervised Fine-Tuning Needs to Unlock the Potential of Token Priority
by: Shen, Zhanming, et al.
Published: (2026) -
Offline Multi-Objective Optimization
by: Xue, Ke, et al.
Published: (2024) -
Offline Model-Based Optimization by Learning to Rank
by: Tan, Rong-Xi, et al.
Published: (2024) -
Zero Token-Driven Deep Thinking in LLMs: Unlocking the Full Potential of Existing Parameters via Cyclic Refinement
by: Li, Guanghao, et al.
Published: (2025)