TQL: Scaling Q-Functions with Transformers by Preventing Attention Collapse

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Dong, Perry, Hung, Kuo-Han, Swerdlow, Alexander, Sadigh, Dorsa, Finn, Chelsea
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914300359081984
author Dong, Perry
Hung, Kuo-Han
Swerdlow, Alexander
Sadigh, Dorsa
Finn, Chelsea
author_facet Dong, Perry
Hung, Kuo-Han
Swerdlow, Alexander
Sadigh, Dorsa
Finn, Chelsea
contents Despite scale driving substantial recent advancements in machine learning, reinforcement learning (RL) methods still primarily use small value functions. Naively scaling value functions -- including with a transformer architecture, which is known to be highly scalable -- often results in learning instability and worse performance. In this work, we ask what prevents transformers from scaling effectively for value functions? Through empirical analysis, we identify the critical failure mode in this scaling: attention scores collapse as capacity increases. Our key insight is that we can effectively prevent this collapse and stabilize training by controlling the entropy of the attention scores, thereby enabling the use of larger models. To this end, we propose Transformer Q-Learning (TQL), a method that unlocks the scaling potential of transformers in learning value functions in RL. Our approach yields up to a 43% improvement in performance when scaling from the smallest to the largest network sizes, while prior methods suffer from performance degradation.
format Preprint
id arxiv_https___arxiv_org_abs_2602_01439
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle TQL: Scaling Q-Functions with Transformers by Preventing Attention Collapse
Dong, Perry
Hung, Kuo-Han
Swerdlow, Alexander
Sadigh, Dorsa
Finn, Chelsea
Machine Learning
Artificial Intelligence
Despite scale driving substantial recent advancements in machine learning, reinforcement learning (RL) methods still primarily use small value functions. Naively scaling value functions -- including with a transformer architecture, which is known to be highly scalable -- often results in learning instability and worse performance. In this work, we ask what prevents transformers from scaling effectively for value functions? Through empirical analysis, we identify the critical failure mode in this scaling: attention scores collapse as capacity increases. Our key insight is that we can effectively prevent this collapse and stabilize training by controlling the entropy of the attention scores, thereby enabling the use of larger models. To this end, we propose Transformer Q-Learning (TQL), a method that unlocks the scaling potential of transformers in learning value functions in RL. Our approach yields up to a 43% improvement in performance when scaling from the smallest to the largest network sizes, while prior methods suffer from performance degradation.
title TQL: Scaling Q-Functions with Transformers by Preventing Attention Collapse
topic Machine Learning
Artificial Intelligence
url https://arxiv.org/abs/2602.01439