The Inverse Lyndon Array: Definition, Properties, and Linear-Time Construction

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Negri, Pietro, Sica, Manuel, Zaccagnino, Rocco, Zizza, Rosalba
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912972597624832
author Negri, Pietro
Sica, Manuel
Zaccagnino, Rocco
Zizza, Rosalba
author_facet Negri, Pietro
Sica, Manuel
Zaccagnino, Rocco
Zizza, Rosalba
contents The Lyndon array stores, at each position of a word, the length of the longest maximal Lyndon subword starting at that position, and plays an important role in combinatorics on words, for example in the construction of fundamental data structures such as the suffix array. In this paper, we introduce the Inverse Lyndon Array, the analogous structure for inverse Lyndon words, namely words that are lexicographically greater than all their proper suffixes. Unlike standard Lyndon words, inverse Lyndon words may have non-trivial borders, which introduces a genuine theoretical difficulty. We show that the inverse Lyndon array can be characterized in terms of the next greater suffix array together with a border-correction term, and prove that this correction coincides with a longest common extension (LCE) value. Building on this characterization, we adapt the nearest-suffix framework underlying Ellert's linear-time construction of the Lyndon array to the inverse setting, obtaining an O(n)-time algorithm for general ordered alphabets. Finally, we discuss implications for suffix comparison and report experiments on random, structured, and real datasets showing that the inverse construction exhibits the same practical linear-time behavior as the standard one.
format Preprint
id arxiv_https___arxiv_org_abs_2603_17537
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle The Inverse Lyndon Array: Definition, Properties, and Linear-Time Construction
Negri, Pietro
Sica, Manuel
Zaccagnino, Rocco
Zizza, Rosalba
Data Structures and Algorithms
Formal Languages and Automata Theory
The Lyndon array stores, at each position of a word, the length of the longest maximal Lyndon subword starting at that position, and plays an important role in combinatorics on words, for example in the construction of fundamental data structures such as the suffix array. In this paper, we introduce the Inverse Lyndon Array, the analogous structure for inverse Lyndon words, namely words that are lexicographically greater than all their proper suffixes. Unlike standard Lyndon words, inverse Lyndon words may have non-trivial borders, which introduces a genuine theoretical difficulty. We show that the inverse Lyndon array can be characterized in terms of the next greater suffix array together with a border-correction term, and prove that this correction coincides with a longest common extension (LCE) value. Building on this characterization, we adapt the nearest-suffix framework underlying Ellert's linear-time construction of the Lyndon array to the inverse setting, obtaining an O(n)-time algorithm for general ordered alphabets. Finally, we discuss implications for suffix comparison and report experiments on random, structured, and real datasets showing that the inverse construction exhibits the same practical linear-time behavior as the standard one.
title The Inverse Lyndon Array: Definition, Properties, and Linear-Time Construction
topic Data Structures and Algorithms
Formal Languages and Automata Theory
url https://arxiv.org/abs/2603.17537