Multi-Stage Boundary-Aware Transformer Network for Action Segmentation in Untrimmed Surgical Videos
Fuente:
arXiv
Saved in:
| Main Authors: | Shuvo, Rezowan, Mekala, M S, Elyan, Eyad |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Adapting Short-Term Transformers for Action Detection in Untrimmed Videos
by: Yang, Min, et al.
Published: (2023)
by: Yang, Min, et al.
Published: (2023)
Multi-Level LVLM Guidance for Untrimmed Video Action Recognition
by: Peng, Liyang, et al.
Published: (2025)
by: Peng, Liyang, et al.
Published: (2025)
MultiCounter: Multiple Action Agnostic Repetition Counting in Untrimmed Videos
by: Tang, Yin, et al.
Published: (2024)
by: Tang, Yin, et al.
Published: (2024)
Explainable Forensics of Manipulated Segments in Untrimmed Long Videos
by: Feng, Yue, et al.
Published: (2026)
by: Feng, Yue, et al.
Published: (2026)
Localizing Moments of Actions in Untrimmed Videos of Infants with Autism Spectrum Disorder
by: Helvaci, Halil Ismail, et al.
Published: (2024)
by: Helvaci, Halil Ismail, et al.
Published: (2024)
What, when, and where? -- Self-Supervised Spatio-Temporal Grounding in Untrimmed Multi-Action Videos from Narrated Instructions
by: Chen, Brian, et al.
Published: (2023)
by: Chen, Brian, et al.
Published: (2023)
MS-Temba: Multi-Scale Temporal Mamba for Understanding Long Untrimmed Videos
by: Sinha, Arkaprava, et al.
Published: (2025)
by: Sinha, Arkaprava, et al.
Published: (2025)
MUVR: A Multi-Modal Untrimmed Video Retrieval Benchmark with Multi-Level Visual Correspondence
by: Feng, Yue, et al.
Published: (2025)
by: Feng, Yue, et al.
Published: (2025)
Boundary-Aware Vision Transformer for Angiography Vascular Network Segmentation
by: Hezil, Nabil, et al.
Published: (2025)
by: Hezil, Nabil, et al.
Published: (2025)
Hierarchical Multi-Stage Transformer Architecture for Context-Aware Temporal Action Localization
by: Ullah, Hayat, et al.
Published: (2025)
by: Ullah, Hayat, et al.
Published: (2025)
Friends Across Time: Multi-Scale Action Segmentation Transformer for Surgical Phase Recognition
by: Zhang, Bokai, et al.
Published: (2024)
by: Zhang, Bokai, et al.
Published: (2024)
TAM-VT: Transformation-Aware Multi-scale Video Transformer for Segmentation and Tracking
by: Goyal, Raghav, et al.
Published: (2023)
by: Goyal, Raghav, et al.
Published: (2023)
Empowering LLMs with Pseudo-Untrimmed Videos for Audio-Visual Temporal Understanding
by: Tang, Yolo Yunlong, et al.
Published: (2024)
by: Tang, Yolo Yunlong, et al.
Published: (2024)
Efficient and Effective Weakly-Supervised Action Segmentation via Action-Transition-Aware Boundary Alignment
by: Xu, Angchi, et al.
Published: (2024)
by: Xu, Angchi, et al.
Published: (2024)
Cutup and Detect: Human Fall Detection on Cutup Untrimmed Videos Using a Large Foundational Video Understanding Model
by: Grutschus, Till, et al.
Published: (2024)
by: Grutschus, Till, et al.
Published: (2024)
Mamba-OTR: a Mamba-based Solution for Online Take and Release Detection from Untrimmed Egocentric Video
by: Catinello, Alessandro Sebastiano, et al.
Published: (2025)
by: Catinello, Alessandro Sebastiano, et al.
Published: (2025)
Grounding Surgical Action Triplets with Instrument Instance Segmentation: A Dataset and Target-Aware Fusion Approach
by: Alabi, Oluwatosin, et al.
Published: (2025)
by: Alabi, Oluwatosin, et al.
Published: (2025)
Boundary-Centric Active Learning for Temporal Action Segmentation
by: Helvaci, Halil Ismail, et al.
Published: (2026)
by: Helvaci, Halil Ismail, et al.
Published: (2026)
ActionVOS: Actions as Prompts for Video Object Segmentation
by: Ouyang, Liangyang, et al.
Published: (2024)
by: Ouyang, Liangyang, et al.
Published: (2024)
Combining Boundary Supervision and Segment-Level Regularization for Fine-Grained Action Segmentation
by: Mitsuoka, Hinako, et al.
Published: (2026)
by: Mitsuoka, Hinako, et al.
Published: (2026)
Local-Global Context Aware Transformer for Language-Guided Video Segmentation
by: Liang, Chen, et al.
Published: (2022)
by: Liang, Chen, et al.
Published: (2022)
Boundary-Aware Instance Segmentation in Microscopy Imaging
by: Mendelson, Thomas, et al.
Published: (2026)
by: Mendelson, Thomas, et al.
Published: (2026)
ViTALS: Vision Transformer for Action Localization in Surgical Nephrectomy
by: Chandra, Soumyadeep, et al.
Published: (2024)
by: Chandra, Soumyadeep, et al.
Published: (2024)
MATIS: Masked-Attention Transformers for Surgical Instrument Segmentation
by: Ayobi, Nicolás, et al.
Published: (2023)
by: Ayobi, Nicolás, et al.
Published: (2023)
Surgical Scene Segmentation by Transformer With Asymmetric Feature Enhancement
by: Yuan, Cheng, et al.
Published: (2024)
by: Yuan, Cheng, et al.
Published: (2024)
Permutation-Aware Action Segmentation via Unsupervised Frame-to-Segment Alignment
by: Tran, Quoc-Huy, et al.
Published: (2023)
by: Tran, Quoc-Huy, et al.
Published: (2023)
A Semantic and Motion-Aware Spatiotemporal Transformer Network for Action Detection
by: Korban, Matthew, et al.
Published: (2024)
by: Korban, Matthew, et al.
Published: (2024)
Boundary-Recovering Network for Temporal Action Detection
by: Kim, Jihwan, et al.
Published: (2024)
by: Kim, Jihwan, et al.
Published: (2024)
Joint Self-Supervised Video Alignment and Action Segmentation
by: Ali, Ali Shah, et al.
Published: (2025)
by: Ali, Ali Shah, et al.
Published: (2025)
Pose-Aware Weakly-Supervised Action Segmentation
by: Zhao, Seth Z., et al.
Published: (2025)
by: Zhao, Seth Z., et al.
Published: (2025)
Efficient Temporal Action Segmentation via Boundary-aware Query Voting
by: Wang, Peiyao, et al.
Published: (2024)
by: Wang, Peiyao, et al.
Published: (2024)
Dual Semantic-Aware Network for Noise Suppressed Ultrasound Video Segmentation
by: Zhou, Ling, et al.
Published: (2025)
by: Zhou, Ling, et al.
Published: (2025)
Toward Real-Time Surgical Scene Segmentation via a Spike-Driven Video Transformer with Spike-Informed Pretraining
by: Zou, Shihao, et al.
Published: (2025)
by: Zou, Shihao, et al.
Published: (2025)
Multi-Stage Contrastive Regression for Action Quality Assessment
by: An, Qi, et al.
Published: (2024)
by: An, Qi, et al.
Published: (2024)
Context-Aware Semantic Segmentation via Stage-Wise Attention
by: Carreaud, Antoine, et al.
Published: (2026)
by: Carreaud, Antoine, et al.
Published: (2026)
A Deep Learning Framework for Boundary-Aware Semantic Segmentation
by: An, Tai, et al.
Published: (2025)
by: An, Tai, et al.
Published: (2025)
Memory-Augmented SAM2 for Training-Free Surgical Video Segmentation
by: Yin, Ming, et al.
Published: (2025)
by: Yin, Ming, et al.
Published: (2025)
Systematic Evaluation and Guidelines for Segment Anything Model in Surgical Video Analysis
by: Yuan, Cheng, et al.
Published: (2024)
by: Yuan, Cheng, et al.
Published: (2024)
FlowBlending: Stage-Aware Multi-Model Sampling for Fast and High-Fidelity Video Generation
by: Song, Jibin, et al.
Published: (2025)
by: Song, Jibin, et al.
Published: (2025)
Augmenting Efficient Real-time Surgical Instrument Segmentation in Video with Point Tracking and Segment Anything
by: Wu, Zijian, et al.
Published: (2024)
by: Wu, Zijian, et al.
Published: (2024)
Similar Items
-
Adapting Short-Term Transformers for Action Detection in Untrimmed Videos
by: Yang, Min, et al.
Published: (2023) -
Multi-Level LVLM Guidance for Untrimmed Video Action Recognition
by: Peng, Liyang, et al.
Published: (2025) -
MultiCounter: Multiple Action Agnostic Repetition Counting in Untrimmed Videos
by: Tang, Yin, et al.
Published: (2024) -
Explainable Forensics of Manipulated Segments in Untrimmed Long Videos
by: Feng, Yue, et al.
Published: (2026) -
Localizing Moments of Actions in Untrimmed Videos of Infants with Autism Spectrum Disorder
by: Helvaci, Halil Ismail, et al.
Published: (2024)