Deriving the Scaled-Dot-Function via Maximum Likelihood Estimation and Maximum Entropy Approach

Fuente: arXiv
Saved in:
Bibliographic Details
Main Author: Ma, Jiyong
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914039054991360
author Ma, Jiyong
author_facet Ma, Jiyong
contents In this paper, we present a maximum likelihood estimation approach to determine the value vector in transformer models. We model the sequence of value vectors, key vectors, and the query vector as a sequence of Gaussian distributions. The variance in each Gaussian distribution depends on the time step, the corresponding key vector, and the query vector. The mean value in each Gaussian distribution depends on the time step, and the corresponding value vector. This analysis may offer a new explanation of the scaled-dot-product function or softmax function used in transformer architectures [1]. Another explanation, inspired by [4], is based on the maximum entropy approach in natural language processing [5]. In this approach, a query vector and key vectors are used to derive the feature functions for the maximum entropy model.
format Preprint
id arxiv_https___arxiv_org_abs_2509_12285
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Deriving the Scaled-Dot-Function via Maximum Likelihood Estimation and Maximum Entropy Approach
Ma, Jiyong
Machine Learning
Artificial Intelligence
In this paper, we present a maximum likelihood estimation approach to determine the value vector in transformer models. We model the sequence of value vectors, key vectors, and the query vector as a sequence of Gaussian distributions. The variance in each Gaussian distribution depends on the time step, the corresponding key vector, and the query vector. The mean value in each Gaussian distribution depends on the time step, and the corresponding value vector. This analysis may offer a new explanation of the scaled-dot-product function or softmax function used in transformer architectures [1]. Another explanation, inspired by [4], is based on the maximum entropy approach in natural language processing [5]. In this approach, a query vector and key vectors are used to derive the feature functions for the maximum entropy model.
title Deriving the Scaled-Dot-Function via Maximum Likelihood Estimation and Maximum Entropy Approach
topic Machine Learning
Artificial Intelligence
url https://arxiv.org/abs/2509.12285