Saved in:
| Main Authors: | , , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2602.04884 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866908831105155072 |
|---|---|
| author | Li, Bangzheng Ni, Jianmo Qu, Chen Miao, Ian Yang, Liu Fu, Xingyu Chen, Muhao Cheng, Derek Zhiyuan |
| author_facet | Li, Bangzheng Ni, Jianmo Qu, Chen Miao, Ian Yang, Liu Fu, Xingyu Chen, Muhao Cheng, Derek Zhiyuan |
| contents | Post-training with Reinforcement Learning (RL) has substantially improved reasoning in Large Language Models (LLMs) via test-time scaling. However, extending this paradigm to Multimodal LLMs (MLLMs) through verbose rationales yields limited gains for perception and can even degrade performance.
We propose Reinforced Attention Learning (RAL), a policy-gradient framework that directly optimizes internal attention distributions rather than output token sequences. By shifting optimization from what to generate to where to attend, RAL promotes effective information allocation and improved grounding in complex multimodal inputs. Experiments across diverse image and video benchmarks show consistent gains over GRPO and other baselines. We further introduce On-Policy Attention Distillation, demonstrating that transferring latent attention behaviors yields stronger cross-modal alignment than standard knowledge distillation. Our results position attention policies as a principled and general alternative for multimodal post-training. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2602_04884 |
| institution | arXiv |
| publishDate | 2026 |
| record_format | arxiv |
| spellingShingle | Reinforced Attention Learning Li, Bangzheng Ni, Jianmo Qu, Chen Miao, Ian Yang, Liu Fu, Xingyu Chen, Muhao Cheng, Derek Zhiyuan Computation and Language Computer Vision and Pattern Recognition Machine Learning Post-training with Reinforcement Learning (RL) has substantially improved reasoning in Large Language Models (LLMs) via test-time scaling. However, extending this paradigm to Multimodal LLMs (MLLMs) through verbose rationales yields limited gains for perception and can even degrade performance. We propose Reinforced Attention Learning (RAL), a policy-gradient framework that directly optimizes internal attention distributions rather than output token sequences. By shifting optimization from what to generate to where to attend, RAL promotes effective information allocation and improved grounding in complex multimodal inputs. Experiments across diverse image and video benchmarks show consistent gains over GRPO and other baselines. We further introduce On-Policy Attention Distillation, demonstrating that transferring latent attention behaviors yields stronger cross-modal alignment than standard knowledge distillation. Our results position attention policies as a principled and general alternative for multimodal post-training. |
| title | Reinforced Attention Learning |
| topic | Computation and Language Computer Vision and Pattern Recognition Machine Learning |
| url | https://arxiv.org/abs/2602.04884 |