SPS: Steering Probability Squeezing for Better Exploration in Reinforcement Learning for Large Language Models
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Huo, Yifu, Wang, Chenglong, Zhu, Ziming, Xing, Shunjie, Feng, Peinan, Liu, Tongran, He, Qiaozhi, Zhou, Tianhua, Chang, Xiaojia, Zhu, Jingbo, Yu, Zhengtao, Xiao, Tong |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
HEAL: A Hypothesis-Based Preference-Aware Analysis Framework
von: Huo, Yifu, et al.
Veröffentlicht: (2025)
von: Huo, Yifu, et al.
Veröffentlicht: (2025)
SERM: Self-Evolving Relevance Model with Agent-Driven Learning from Massive Query Streams
von: Wang, Chenglong, et al.
Veröffentlicht: (2026)
von: Wang, Chenglong, et al.
Veröffentlicht: (2026)
APR: Penalizing Structural Redundancy in Large Reasoning Models via Anchor-based Process Rewards
von: Chang, Kaiyan, et al.
Veröffentlicht: (2026)
von: Chang, Kaiyan, et al.
Veröffentlicht: (2026)
LRHP: Learning Representations for Human Preferences via Preference Pairs
von: Wang, Chenglong, et al.
Veröffentlicht: (2024)
von: Wang, Chenglong, et al.
Veröffentlicht: (2024)
MSRL: Scaling Generative Multimodal Reward Modeling via Multi-Stage Reinforcement Learning
von: Wang, Chenglong, et al.
Veröffentlicht: (2026)
von: Wang, Chenglong, et al.
Veröffentlicht: (2026)
Probing Preference Representations: A Multi-Dimensional Evaluation and Analysis Method for Reward Models
von: Wang, Chenglong, et al.
Veröffentlicht: (2025)
von: Wang, Chenglong, et al.
Veröffentlicht: (2025)
GRAM: A Generative Foundation Reward Model for Reward Generalization
von: Wang, Chenglong, et al.
Veröffentlicht: (2025)
von: Wang, Chenglong, et al.
Veröffentlicht: (2025)
LaTeXTrans: Structured LaTeX Translation with Multi-Agent Coordination
von: Zhu, Ziming, et al.
Veröffentlicht: (2025)
von: Zhu, Ziming, et al.
Veröffentlicht: (2025)
RoVRM: A Robust Visual Reward Model Optimized via Auxiliary Textual Preference Data
von: Wang, Chenglong, et al.
Veröffentlicht: (2024)
von: Wang, Chenglong, et al.
Veröffentlicht: (2024)
Revealing the Parallel Multilingual Learning within Large Language Models
von: Mu, Yongyu, et al.
Veröffentlicht: (2024)
von: Mu, Yongyu, et al.
Veröffentlicht: (2024)
Hybrid Alignment Training for Large Language Models
von: Wang, Chenglong, et al.
Veröffentlicht: (2024)
von: Wang, Chenglong, et al.
Veröffentlicht: (2024)
When Scaling Fails: Mitigating Audio Perception Decay of LALMs via Multi-Step Perception-Aware Reasoning
von: Mao, Ruixiang, et al.
Veröffentlicht: (2026)
von: Mao, Ruixiang, et al.
Veröffentlicht: (2026)
Learning Evaluation Models from Large Language Models for Sequence Generation
von: Wang, Chenglong, et al.
Veröffentlicht: (2023)
von: Wang, Chenglong, et al.
Veröffentlicht: (2023)
DaPT: A Dual-Path Framework for Multilingual Multi-hop Question Answering
von: Wang, Yilin, et al.
Veröffentlicht: (2026)
von: Wang, Yilin, et al.
Veröffentlicht: (2026)
MRO: Enhancing Reasoning in Diffusion Language Models via Multi-Reward Optimization
von: Wang, Chenglong, et al.
Veröffentlicht: (2025)
von: Wang, Chenglong, et al.
Veröffentlicht: (2025)
GRAM-R$^2$: Self-Training Generative Foundation Reward Models for Reward Reasoning
von: Wang, Chenglong, et al.
Veröffentlicht: (2025)
von: Wang, Chenglong, et al.
Veröffentlicht: (2025)
EfficientGraph-RAG: Structured Retrieval-State Management for Cross-Task Retrieval-Augmented Generation
von: Niu, Miaohe, et al.
Veröffentlicht: (2026)
von: Niu, Miaohe, et al.
Veröffentlicht: (2026)
Boosting Text-To-Image Generation via Multilingual Prompting in Large Multimodal Models
von: Mu, Yongyu, et al.
Veröffentlicht: (2025)
von: Mu, Yongyu, et al.
Veröffentlicht: (2025)
Foundations of Large Language Models
von: Xiao, Tong, et al.
Veröffentlicht: (2025)
von: Xiao, Tong, et al.
Veröffentlicht: (2025)
M-CIF: Multi-Scale Alignment For CIF-Based Non-Autoregressive ASR
von: Mao, Ruixiang, et al.
Veröffentlicht: (2025)
von: Mao, Ruixiang, et al.
Veröffentlicht: (2025)
Efficient Prompting Methods for Large Language Models: A Survey
von: Chang, Kaiyan, et al.
Veröffentlicht: (2024)
von: Chang, Kaiyan, et al.
Veröffentlicht: (2024)
Prior Constraints-based Reward Model Training for Aligning Large Language Models
von: Zhou, Hang, et al.
Veröffentlicht: (2024)
von: Zhou, Hang, et al.
Veröffentlicht: (2024)
NiuTrans.LMT: Toward Inclusive and Scalable Multilingual Machine Translation with LLMs
von: Luo, Yingfeng, et al.
Veröffentlicht: (2025)
von: Luo, Yingfeng, et al.
Veröffentlicht: (2025)
Cross-layer Attention Sharing for Pre-trained Large Language Models
von: Mu, Yongyu, et al.
Veröffentlicht: (2024)
von: Mu, Yongyu, et al.
Veröffentlicht: (2024)
Beyond Decoder-only: Large Language Models Can be Good Encoders for Machine Translation
von: Luo, Yingfeng, et al.
Veröffentlicht: (2025)
von: Luo, Yingfeng, et al.
Veröffentlicht: (2025)
FLEXI: Benchmarking Full-duplex Human-LLM Speech Interaction
von: Ge, Yuan, et al.
Veröffentlicht: (2025)
von: Ge, Yuan, et al.
Veröffentlicht: (2025)
Attention2Probability: Attention-Driven Terminology Probability Estimation for Robust Speech-to-Text System
von: Du, Yanfan, et al.
Veröffentlicht: (2025)
von: Du, Yanfan, et al.
Veröffentlicht: (2025)
One Size Does Not Fit All: A Distribution-Aware Sparsification for More Precise Model Merging
von: Luo, Yingfeng, et al.
Veröffentlicht: (2025)
von: Luo, Yingfeng, et al.
Veröffentlicht: (2025)
RankPrompt: Step-by-Step Comparisons Make Language Models Better Reasoners
von: Hu, Chi, et al.
Veröffentlicht: (2024)
von: Hu, Chi, et al.
Veröffentlicht: (2024)
RouteLMT: Learned Sample Routing for Hybrid LLM Translation Deployment
von: Luo, Yingfeng, et al.
Veröffentlicht: (2026)
von: Luo, Yingfeng, et al.
Veröffentlicht: (2026)
Striking Gold in Advertising: Standardization and Exploration of Ad Text Generation
von: Mita, Masato, et al.
Veröffentlicht: (2023)
von: Mita, Masato, et al.
Veröffentlicht: (2023)
Autoencoding-Free Context Compression for LLMs via Contextual Semantic Anchors
von: Liu, Xin, et al.
Veröffentlicht: (2025)
von: Liu, Xin, et al.
Veröffentlicht: (2025)
MTP-S2UT: Enhancing Speech-to-Speech Translation Quality with Multi-token Prediction
von: Wang, Jianjin, et al.
Veröffentlicht: (2025)
von: Wang, Jianjin, et al.
Veröffentlicht: (2025)
On the Emotion Understanding of Synthesized Speech
von: Ge, Yuan, et al.
Veröffentlicht: (2026)
von: Ge, Yuan, et al.
Veröffentlicht: (2026)
Training Tales: Steer Squeeze Training at the Nashville Zoo
von: Sproul, Kate
Veröffentlicht: (2016)
von: Sproul, Kate
Veröffentlicht: (2016)
EIT: Enhanced Interactive Transformer
von: Zheng, Tong, et al.
Veröffentlicht: (2022)
von: Zheng, Tong, et al.
Veröffentlicht: (2022)
A Review of Electronic Early Warning Systems for Acute Kidney Injury
von: Xiangxiang Wang, et al.
Veröffentlicht: (2024)
von: Xiangxiang Wang, et al.
Veröffentlicht: (2024)
Prognostic Value of Serum Cystatin C for Predicting Contrast‐Induced Nephropathy in Elderly Individuals Aged Over 80 Years
von: Xiangxiang Wang, et al.
Veröffentlicht: (2026)
von: Xiangxiang Wang, et al.
Veröffentlicht: (2026)
Step-level Verifier-guided Hybrid Test-Time Scaling for Large Language Models
von: Chang, Kaiyan, et al.
Veröffentlicht: (2025)
von: Chang, Kaiyan, et al.
Veröffentlicht: (2025)
Leveraging Unit Language Guidance to Advance Speech Modeling in Textless Speech-to-Speech Translation
von: Zhang, Yuhao, et al.
Veröffentlicht: (2025)
von: Zhang, Yuhao, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
HEAL: A Hypothesis-Based Preference-Aware Analysis Framework
von: Huo, Yifu, et al.
Veröffentlicht: (2025) -
SERM: Self-Evolving Relevance Model with Agent-Driven Learning from Massive Query Streams
von: Wang, Chenglong, et al.
Veröffentlicht: (2026) -
APR: Penalizing Structural Redundancy in Large Reasoning Models via Anchor-based Process Rewards
von: Chang, Kaiyan, et al.
Veröffentlicht: (2026) -
LRHP: Learning Representations for Human Preferences via Preference Pairs
von: Wang, Chenglong, et al.
Veröffentlicht: (2024) -
MSRL: Scaling Generative Multimodal Reward Modeling via Multi-Stage Reinforcement Learning
von: Wang, Chenglong, et al.
Veröffentlicht: (2026)