Generalized Neural Sorting Networks with Error-Free Differentiable Swap Functions

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Kim, Jungtaek, Yoon, Jeongbeen, Cho, Minsu
Format: Preprint
Published: 2023
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910366604197888
author Kim, Jungtaek
Yoon, Jeongbeen
Cho, Minsu
author_facet Kim, Jungtaek
Yoon, Jeongbeen
Cho, Minsu
contents Sorting is a fundamental operation of all computer systems, having been a long-standing significant research topic. Beyond the problem formulation of traditional sorting algorithms, we consider sorting problems for more abstract yet expressive inputs, e.g., multi-digit images and image fragments, through a neural sorting network. To learn a mapping from a high-dimensional input to an ordinal variable, the differentiability of sorting networks needs to be guaranteed. In this paper we define a softening error by a differentiable swap function, and develop an error-free swap function that holds a non-decreasing condition and differentiability. Furthermore, a permutation-equivariant Transformer network with multi-head attention is adopted to capture dependency between given inputs and also leverage its model capacity with self-attention. Experiments on diverse sorting benchmarks show that our methods perform better than or comparable to baseline methods.
format Preprint
id arxiv_https___arxiv_org_abs_2310_07174
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle Generalized Neural Sorting Networks with Error-Free Differentiable Swap Functions
Kim, Jungtaek
Yoon, Jeongbeen
Cho, Minsu
Machine Learning
Sorting is a fundamental operation of all computer systems, having been a long-standing significant research topic. Beyond the problem formulation of traditional sorting algorithms, we consider sorting problems for more abstract yet expressive inputs, e.g., multi-digit images and image fragments, through a neural sorting network. To learn a mapping from a high-dimensional input to an ordinal variable, the differentiability of sorting networks needs to be guaranteed. In this paper we define a softening error by a differentiable swap function, and develop an error-free swap function that holds a non-decreasing condition and differentiability. Furthermore, a permutation-equivariant Transformer network with multi-head attention is adopted to capture dependency between given inputs and also leverage its model capacity with self-attention. Experiments on diverse sorting benchmarks show that our methods perform better than or comparable to baseline methods.
title Generalized Neural Sorting Networks with Error-Free Differentiable Swap Functions
topic Machine Learning
url https://arxiv.org/abs/2310.07174