Softmax is not Enough (for Sharp Size Generalisation)

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Veličković, Petar, Perivolaropoulos, Christos, Barbero, Federico, Pascanu, Razvan
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915314512429056
author Veličković, Petar
Perivolaropoulos, Christos
Barbero, Federico
Pascanu, Razvan
author_facet Veličković, Petar
Perivolaropoulos, Christos
Barbero, Federico
Pascanu, Razvan
contents A key property of reasoning systems is the ability to make sharp decisions on their input data. For contemporary AI systems, a key carrier of sharp behaviour is the softmax function, with its capability to perform differentiable query-key lookups. It is a common belief that the predictive power of networks leveraging softmax arises from "circuits" which sharply perform certain kinds of computations consistently across many diverse inputs. However, for these circuits to be robust, they would need to generalise well to arbitrary valid inputs. In this paper, we dispel this myth: even for tasks as simple as finding the maximum key, any learned circuitry must disperse as the number of items grows at test time. We attribute this to a fundamental limitation of the softmax function to robustly approximate sharp functions with increasing problem size, prove this phenomenon theoretically, and propose adaptive temperature as an ad-hoc technique for improving the sharpness of softmax at inference time.
format Preprint
id arxiv_https___arxiv_org_abs_2410_01104
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Softmax is not Enough (for Sharp Size Generalisation)
Veličković, Petar
Perivolaropoulos, Christos
Barbero, Federico
Pascanu, Razvan
Machine Learning
Artificial Intelligence
Information Theory
A key property of reasoning systems is the ability to make sharp decisions on their input data. For contemporary AI systems, a key carrier of sharp behaviour is the softmax function, with its capability to perform differentiable query-key lookups. It is a common belief that the predictive power of networks leveraging softmax arises from "circuits" which sharply perform certain kinds of computations consistently across many diverse inputs. However, for these circuits to be robust, they would need to generalise well to arbitrary valid inputs. In this paper, we dispel this myth: even for tasks as simple as finding the maximum key, any learned circuitry must disperse as the number of items grows at test time. We attribute this to a fundamental limitation of the softmax function to robustly approximate sharp functions with increasing problem size, prove this phenomenon theoretically, and propose adaptive temperature as an ad-hoc technique for improving the sharpness of softmax at inference time.
title Softmax is not Enough (for Sharp Size Generalisation)
topic Machine Learning
Artificial Intelligence
Information Theory
url https://arxiv.org/abs/2410.01104