Reinforcement Learning with Semantic Rewards Enables Low-Resource Language Expansion without Alignment Tax
Fuente:
arXiv
Guardado en:
| Autores principales: | Su, Zeli, Zhang, Ziyin, Liu, Zhou, Song, Xuexian, Xu, Zhankai, Zheng, Longfei, Zhang, Xiaolu, Fu, Rong, Xu, Guixian, Zhang, Wentao |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Source-Grounded Semantic Reinforcement Learning for Low-Resource Target-Language Generation
por: Su, Zeli, et al.
Publicado: (2026)
por: Su, Zeli, et al.
Publicado: (2026)
FTibSuite: A Comprehensive Resource Suite for Tibetan Vision-Language Modeling
por: Xu, Guixian, et al.
Publicado: (2026)
por: Xu, Guixian, et al.
Publicado: (2026)
The Curse of Helpfulness: Inverse Scaling Law in Robustness to Distractor Instructions via DistractionIF
por: Su, Zeli, et al.
Publicado: (2026)
por: Su, Zeli, et al.
Publicado: (2026)
SHRP: Specialized Head Routing and Pruning for Efficient Encoder Compression
por: Su, Zeli, et al.
Publicado: (2025)
por: Su, Zeli, et al.
Publicado: (2025)
Multilingual Encoder Knows more than You Realize: Shared Weights Pretraining for Extremely Low-Resource Languages
por: Su, Zeli, et al.
Publicado: (2025)
por: Su, Zeli, et al.
Publicado: (2025)
CMHG: A Dataset and Benchmark for Headline Generation of Minority Languages in China
por: Xu, Guixian, et al.
Publicado: (2025)
por: Xu, Guixian, et al.
Publicado: (2025)
Group Symmetry Enables Faster Optimization in Inverse Problems
por: Tang, Junqi, et al.
Publicado: (2025)
por: Tang, Junqi, et al.
Publicado: (2025)
Multilingual LLMs Inherently Reward In-Language Time-Sensitive Semantic Alignment for Low-Resource Languages
por: Bajpai, Ashutosh, et al.
Publicado: (2024)
por: Bajpai, Ashutosh, et al.
Publicado: (2024)
Pareto-Guided Optimal Transport for Multi-Reward Alignment
por: Ba, Ying, et al.
Publicado: (2026)
por: Ba, Ying, et al.
Publicado: (2026)
Online Merging Optimizers for Boosting Rewards and Mitigating Tax in Alignment
por: Lu, Keming, et al.
Publicado: (2024)
por: Lu, Keming, et al.
Publicado: (2024)
Mitigating the Alignment Tax of RLHF
por: Lin, Yong, et al.
Publicado: (2023)
por: Lin, Yong, et al.
Publicado: (2023)
Safety Alignment as Continual Learning: Mitigating the Alignment Tax via Orthogonal Gradient Projection
por: Sun, Guanglong, et al.
Publicado: (2026)
por: Sun, Guanglong, et al.
Publicado: (2026)
FuRL: Visual-Language Models as Fuzzy Rewards for Reinforcement Learning
por: Fu, Yuwei, et al.
Publicado: (2024)
por: Fu, Yuwei, et al.
Publicado: (2024)
Policy Expansion for Bridging Offline-to-Online Reinforcement Learning
por: Zhang, Haichao, et al.
Publicado: (2023)
por: Zhang, Haichao, et al.
Publicado: (2023)
Anatomy-Aware Low-Dose CT Denoising via Pretrained Vision Models and Semantic-Guided Contrastive Learning
por: Wang, Runze, et al.
Publicado: (2025)
por: Wang, Runze, et al.
Publicado: (2025)
Focal Reward: Balanced Reinforcement Learning under Rubric-Based Rewards
por: Huang, Yu, et al.
Publicado: (2026)
por: Huang, Yu, et al.
Publicado: (2026)
The Hallucination Tax of Reinforcement Finetuning
por: Song, Linxin, et al.
Publicado: (2025)
por: Song, Linxin, et al.
Publicado: (2025)
Taxes without Taxpayers: The Invisibility of Taxes in Chile
por: Andrés Biehl
Publicado: (2019)
por: Andrés Biehl
Publicado: (2019)
Zeroth-Order Policy Gradient for Reinforcement Learning from Human Feedback without Reward Inference
por: Zhang, Qining, et al.
Publicado: (2024)
por: Zhang, Qining, et al.
Publicado: (2024)
Does Climate Finance Accelerate Progress Toward SDG 7? Evidence From Low‐and Middle‐Income Countries
por: Junyi Tian, et al.
Publicado: (2026)
por: Junyi Tian, et al.
Publicado: (2026)
Enhancing Reward Models for High-quality Image Generation: Beyond Text-Image Alignment
por: Ba, Ying, et al.
Publicado: (2025)
por: Ba, Ying, et al.
Publicado: (2025)
Prototypical Progressive Alignment and Reweighting for Generalizable Semantic Segmentation
por: Zhang, Yuhang, et al.
Publicado: (2025)
por: Zhang, Yuhang, et al.
Publicado: (2025)
MERIT: Multilingual Expert-Reward Informed Tuning for Chinese-Centric Low-Resource Machine Translation
por: Lu, Zhixiang, et al.
Publicado: (2026)
por: Lu, Zhixiang, et al.
Publicado: (2026)
Constraints as Rewards: Reinforcement Learning for Robots without Reward Functions
por: Ishihara, Yu, et al.
Publicado: (2025)
por: Ishihara, Yu, et al.
Publicado: (2025)
Inverse Reinforcement Learning without an Optimal Demonstrator: A Feasible Reward Set Approach
por: Kim, Kihyun, et al.
Publicado: (2026)
por: Kim, Kihyun, et al.
Publicado: (2026)
RMB: Comprehensively Benchmarking Reward Models in LLM Alignment
por: Zhou, Enyu, et al.
Publicado: (2024)
por: Zhou, Enyu, et al.
Publicado: (2024)
What Is the Alignment Tax?
por: Young, Robin
Publicado: (2026)
por: Young, Robin
Publicado: (2026)
Locating a shortest vector in certain $2$-dimensional lattices
por: Zou, Guixian
Publicado: (2026)
por: Zou, Guixian
Publicado: (2026)
Fast Gradient Methods for Data-Consistent Local Super-Resolution of Medical Images
por: Tang, Junqi, et al.
Publicado: (2022)
por: Tang, Junqi, et al.
Publicado: (2022)
A Comparative Study of Variational Autoencoders, Normalizing Flows, and Score-based Diffusion Models for Electrical Impedance Tomography
por: Wang, Huihui, et al.
Publicado: (2023)
por: Wang, Huihui, et al.
Publicado: (2023)
Fast Equivariant Imaging: Acceleration for Unsupervised Learning via Augmented Lagrangian and Auxiliary PnP Denoisers
por: Xu, Guixian, et al.
Publicado: (2025)
por: Xu, Guixian, et al.
Publicado: (2025)
A New Convergence Analysis of Plug-and-Play Proximal Gradient Descent Under Prior Mismatch
por: Xu, Guixian, et al.
Publicado: (2026)
por: Xu, Guixian, et al.
Publicado: (2026)
Equivariant Test-Time Training with Operator Sketching for Imaging Inverse Problems
por: Xu, Guixian, et al.
Publicado: (2024)
por: Xu, Guixian, et al.
Publicado: (2024)
Enhancing Electrical Impedance Tomography reconstruction using Learned Half-Quadratic Splitting Networks with Anderson Acceleration
por: Xu, Guixian, et al.
Publicado: (2023)
por: Xu, Guixian, et al.
Publicado: (2023)
Conscious Gaze: Adaptive Attention Mechanisms for Hallucination Mitigation in Vision-Language Models
por: Bu, Weijue, et al.
Publicado: (2025)
por: Bu, Weijue, et al.
Publicado: (2025)
NALA: an Effective and Interpretable Entity Alignment Method
por: Xu, Chuanhao, et al.
Publicado: (2024)
por: Xu, Chuanhao, et al.
Publicado: (2024)
Joint Semantic Transmission and Resource Allocation for Intelligent Computation Task Offloading in MEC Systems
por: Zheng, Yuanpeng, et al.
Publicado: (2025)
por: Zheng, Yuanpeng, et al.
Publicado: (2025)
Pay Attention Later: Isolating the Semantic Alignment Tax via Iterative Semantic Map Refinement
por: YILDIRIM, Alper, et al.
Publicado: (2025)
por: YILDIRIM, Alper, et al.
Publicado: (2025)
Pay Attention Later: Isolating the Semantic Alignment Tax via Iterative Semantic Map Refinement
por: YILDIRIM, Alper, et al.
Publicado: (2025)
por: YILDIRIM, Alper, et al.
Publicado: (2025)
Three Mechanisms of Feature Learning in a Linear Network
por: Xu, Yizhou, et al.
Publicado: (2024)
por: Xu, Yizhou, et al.
Publicado: (2024)
Ejemplares similares
-
Source-Grounded Semantic Reinforcement Learning for Low-Resource Target-Language Generation
por: Su, Zeli, et al.
Publicado: (2026) -
FTibSuite: A Comprehensive Resource Suite for Tibetan Vision-Language Modeling
por: Xu, Guixian, et al.
Publicado: (2026) -
The Curse of Helpfulness: Inverse Scaling Law in Robustness to Distractor Instructions via DistractionIF
por: Su, Zeli, et al.
Publicado: (2026) -
SHRP: Specialized Head Routing and Pruning for Efficient Encoder Compression
por: Su, Zeli, et al.
Publicado: (2025) -
Multilingual Encoder Knows more than You Realize: Shared Weights Pretraining for Extremely Low-Resource Languages
por: Su, Zeli, et al.
Publicado: (2025)