Saved in:
Bibliographic Details
Main Authors: Li, Bangzheng, Ni, Jianmo, Qu, Chen, Miao, Ian, Yang, Liu, Fu, Xingyu, Chen, Muhao, Cheng, Derek Zhiyuan
Format: Preprint
Published: 2026
Subjects:
Online Access:https://arxiv.org/abs/2602.04884
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908831105155072
author Li, Bangzheng
Ni, Jianmo
Qu, Chen
Miao, Ian
Yang, Liu
Fu, Xingyu
Chen, Muhao
Cheng, Derek Zhiyuan
author_facet Li, Bangzheng
Ni, Jianmo
Qu, Chen
Miao, Ian
Yang, Liu
Fu, Xingyu
Chen, Muhao
Cheng, Derek Zhiyuan
contents Post-training with Reinforcement Learning (RL) has substantially improved reasoning in Large Language Models (LLMs) via test-time scaling. However, extending this paradigm to Multimodal LLMs (MLLMs) through verbose rationales yields limited gains for perception and can even degrade performance. We propose Reinforced Attention Learning (RAL), a policy-gradient framework that directly optimizes internal attention distributions rather than output token sequences. By shifting optimization from what to generate to where to attend, RAL promotes effective information allocation and improved grounding in complex multimodal inputs. Experiments across diverse image and video benchmarks show consistent gains over GRPO and other baselines. We further introduce On-Policy Attention Distillation, demonstrating that transferring latent attention behaviors yields stronger cross-modal alignment than standard knowledge distillation. Our results position attention policies as a principled and general alternative for multimodal post-training.
format Preprint
id arxiv_https___arxiv_org_abs_2602_04884
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Reinforced Attention Learning
Li, Bangzheng
Ni, Jianmo
Qu, Chen
Miao, Ian
Yang, Liu
Fu, Xingyu
Chen, Muhao
Cheng, Derek Zhiyuan
Computation and Language
Computer Vision and Pattern Recognition
Machine Learning
Post-training with Reinforcement Learning (RL) has substantially improved reasoning in Large Language Models (LLMs) via test-time scaling. However, extending this paradigm to Multimodal LLMs (MLLMs) through verbose rationales yields limited gains for perception and can even degrade performance. We propose Reinforced Attention Learning (RAL), a policy-gradient framework that directly optimizes internal attention distributions rather than output token sequences. By shifting optimization from what to generate to where to attend, RAL promotes effective information allocation and improved grounding in complex multimodal inputs. Experiments across diverse image and video benchmarks show consistent gains over GRPO and other baselines. We further introduce On-Policy Attention Distillation, demonstrating that transferring latent attention behaviors yields stronger cross-modal alignment than standard knowledge distillation. Our results position attention policies as a principled and general alternative for multimodal post-training.
title Reinforced Attention Learning
topic Computation and Language
Computer Vision and Pattern Recognition
Machine Learning
url https://arxiv.org/abs/2602.04884