SOMBRERO: Measuring and Steering Boundary Placement in End-to-End Hierarchical Sequence Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Neitemeier, Pit, Serra, Alessio, Li, Jiaze, Wirges, Sascha, Balles, Lukas, Metzen, Jan Hendrik
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917234658508800
author Neitemeier, Pit
Serra, Alessio
Li, Jiaze
Wirges, Sascha
Balles, Lukas
Metzen, Jan Hendrik
author_facet Neitemeier, Pit
Serra, Alessio
Li, Jiaze
Wirges, Sascha
Balles, Lukas
Metzen, Jan Hendrik
contents Hierarchical sequence models replace fixed tokenization with learned segmentations that compress long byte sequences for efficient autoregressive modeling. While recent end-to-end methods can learn meaningful boundaries from the language-modeling objective alone, it remains difficult to quantitatively assess and systematically steer where compute is spent. We introduce a router-agnostic metric of boundary quality, boundary enrichment B, which measures how strongly chunk starts concentrate on positions with high next-byte surprisal. Guided by this metric, we propose Sombrero, which steers boundary placement toward predictive difficulty via a confidence-alignment boundary loss and stabilizes boundary learning by applying confidence-weighted smoothing at the input level rather than on realized chunks. On 1B scale, across UTF-8 corpora covering English and German text as well as code and mathematical content, Sombrero improves the accuracy-efficiency trade-off and yields boundaries that more consistently align compute with hard-to-predict positions.
format Preprint
id arxiv_https___arxiv_org_abs_2601_22805
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle SOMBRERO: Measuring and Steering Boundary Placement in End-to-End Hierarchical Sequence Models
Neitemeier, Pit
Serra, Alessio
Li, Jiaze
Wirges, Sascha
Balles, Lukas
Metzen, Jan Hendrik
Machine Learning
Artificial Intelligence
Computation and Language
Hierarchical sequence models replace fixed tokenization with learned segmentations that compress long byte sequences for efficient autoregressive modeling. While recent end-to-end methods can learn meaningful boundaries from the language-modeling objective alone, it remains difficult to quantitatively assess and systematically steer where compute is spent. We introduce a router-agnostic metric of boundary quality, boundary enrichment B, which measures how strongly chunk starts concentrate on positions with high next-byte surprisal. Guided by this metric, we propose Sombrero, which steers boundary placement toward predictive difficulty via a confidence-alignment boundary loss and stabilizes boundary learning by applying confidence-weighted smoothing at the input level rather than on realized chunks. On 1B scale, across UTF-8 corpora covering English and German text as well as code and mathematical content, Sombrero improves the accuracy-efficiency trade-off and yields boundaries that more consistently align compute with hard-to-predict positions.
title SOMBRERO: Measuring and Steering Boundary Placement in End-to-End Hierarchical Sequence Models
topic Machine Learning
Artificial Intelligence
Computation and Language
url https://arxiv.org/abs/2601.22805