End-to-End Action Segmentation Transformer
Fuente:
arXiv
Salvato in:
| Autori principali: | , |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
| _version_ | 1866909754179190784 |
|---|---|
| author | Wang, Tieqiao Todorovic, Sinisa |
| author_facet | Wang, Tieqiao Todorovic, Sinisa |
| contents | Most recent work on action segmentation relies on pre-computed frame features from models trained on other tasks and typically focuses on framewise encoding and labeling without explicitly modeling action segments. To overcome these limitations, we introduce the End-to-End Action Segmentation Transformer (EAST), which processes raw video frames directly -- eliminating the need for pre-extracted features and enabling true end-to-end training. Our contributions are as follows: (1) a lightweight adapter design for effective fine-tuning of large backbones; (2) an efficient segmentation-by-detection framework for leveraging action proposals predicted over a coarsely downsampled video; and (3) a novel action-proposal-based data augmentation strategy. EAST achieves SOTA performance on standard benchmarks, including GTEA, 50Salads, Breakfast, and Assembly-101. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2503_06316 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | End-to-End Action Segmentation Transformer Wang, Tieqiao Todorovic, Sinisa Computer Vision and Pattern Recognition Image and Video Processing Most recent work on action segmentation relies on pre-computed frame features from models trained on other tasks and typically focuses on framewise encoding and labeling without explicitly modeling action segments. To overcome these limitations, we introduce the End-to-End Action Segmentation Transformer (EAST), which processes raw video frames directly -- eliminating the need for pre-extracted features and enabling true end-to-end training. Our contributions are as follows: (1) a lightweight adapter design for effective fine-tuning of large backbones; (2) an efficient segmentation-by-detection framework for leveraging action proposals predicted over a coarsely downsampled video; and (3) a novel action-proposal-based data augmentation strategy. EAST achieves SOTA performance on standard benchmarks, including GTEA, 50Salads, Breakfast, and Assembly-101. |
| title | End-to-End Action Segmentation Transformer |
| topic | Computer Vision and Pattern Recognition Image and Video Processing |
| url | https://arxiv.org/abs/2503.06316 |