End-to-End Action Segmentation Transformer

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Wang, Tieqiao, Todorovic, Sinisa
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866909754179190784
author Wang, Tieqiao
Todorovic, Sinisa
author_facet Wang, Tieqiao
Todorovic, Sinisa
contents Most recent work on action segmentation relies on pre-computed frame features from models trained on other tasks and typically focuses on framewise encoding and labeling without explicitly modeling action segments. To overcome these limitations, we introduce the End-to-End Action Segmentation Transformer (EAST), which processes raw video frames directly -- eliminating the need for pre-extracted features and enabling true end-to-end training. Our contributions are as follows: (1) a lightweight adapter design for effective fine-tuning of large backbones; (2) an efficient segmentation-by-detection framework for leveraging action proposals predicted over a coarsely downsampled video; and (3) a novel action-proposal-based data augmentation strategy. EAST achieves SOTA performance on standard benchmarks, including GTEA, 50Salads, Breakfast, and Assembly-101.
format Preprint
id arxiv_https___arxiv_org_abs_2503_06316
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle End-to-End Action Segmentation Transformer
Wang, Tieqiao
Todorovic, Sinisa
Computer Vision and Pattern Recognition
Image and Video Processing
Most recent work on action segmentation relies on pre-computed frame features from models trained on other tasks and typically focuses on framewise encoding and labeling without explicitly modeling action segments. To overcome these limitations, we introduce the End-to-End Action Segmentation Transformer (EAST), which processes raw video frames directly -- eliminating the need for pre-extracted features and enabling true end-to-end training. Our contributions are as follows: (1) a lightweight adapter design for effective fine-tuning of large backbones; (2) an efficient segmentation-by-detection framework for leveraging action proposals predicted over a coarsely downsampled video; and (3) a novel action-proposal-based data augmentation strategy. EAST achieves SOTA performance on standard benchmarks, including GTEA, 50Salads, Breakfast, and Assembly-101.
title End-to-End Action Segmentation Transformer
topic Computer Vision and Pattern Recognition
Image and Video Processing
url https://arxiv.org/abs/2503.06316