ViTALS: Vision Transformer for Action Localization in Surgical Nephrectomy
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Chandra, Soumyadeep, Chowdhury, Sayeed Shafayet, Yong, Courtney, Sundaram, Chandru P., Roy, Kaushik |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Towards Visual Syntactical Understanding
von: Chowdhury, Sayeed Shafayet, et al.
Veröffentlicht: (2024)
von: Chowdhury, Sayeed Shafayet, et al.
Veröffentlicht: (2024)
REMAP: Regularized Matching and Partial Alignment of Video Embeddings
von: Chandra, Soumyadeep, et al.
Veröffentlicht: (2025)
von: Chandra, Soumyadeep, et al.
Veröffentlicht: (2025)
2D-ThermAl: Physics-Informed Framework for Thermal Analysis of Circuits using Generative AI
von: Chandra, Soumyadeep, et al.
Veröffentlicht: (2025)
von: Chandra, Soumyadeep, et al.
Veröffentlicht: (2025)
Loss Knows Best: Detecting Annotation Errors in Videos via Loss Trajectories
von: Alwis, Praditha, et al.
Veröffentlicht: (2026)
von: Alwis, Praditha, et al.
Veröffentlicht: (2026)
LocalViT: Analyzing Locality in Vision Transformers
von: Li, Yawei, et al.
Veröffentlicht: (2021)
von: Li, Yawei, et al.
Veröffentlicht: (2021)
VAT: Vision Action Transformer by Unlocking Full Representation of ViT
von: Li, Wenhao, et al.
Veröffentlicht: (2025)
von: Li, Wenhao, et al.
Veröffentlicht: (2025)
IML-ViT: Benchmarking Image Manipulation Localization by Vision Transformer
von: Ma, Xiaochen, et al.
Veröffentlicht: (2023)
von: Ma, Xiaochen, et al.
Veröffentlicht: (2023)
MangoLeafViT: Leveraging Lightweight Vision Transformer with Runtime Augmentation for Efficient Mango Leaf Disease Classification
von: Chowdhury, Rafi Hassan, et al.
Veröffentlicht: (2025)
von: Chowdhury, Rafi Hassan, et al.
Veröffentlicht: (2025)
BornoViT: A Novel Efficient Vision Transformer for Bengali Handwritten Basic Characters Classification
von: Chowdhury, Rafi Hassan, et al.
Veröffentlicht: (2026)
von: Chowdhury, Rafi Hassan, et al.
Veröffentlicht: (2026)
ViT-5: Vision Transformers for The Mid-2020s
von: Wang, Feng, et al.
Veröffentlicht: (2026)
von: Wang, Feng, et al.
Veröffentlicht: (2026)
FTerViT: Fully Ternary Vision Transformer
von: Ruciński, Szymon, et al.
Veröffentlicht: (2026)
von: Ruciński, Szymon, et al.
Veröffentlicht: (2026)
ViTAR: Vision Transformer with Any Resolution
von: Fan, Qihang, et al.
Veröffentlicht: (2024)
von: Fan, Qihang, et al.
Veröffentlicht: (2024)
Retina Vision Transformer (RetinaViT): Introducing Scaled Patches into Vision Transformers
von: Shu, Yuyang, et al.
Veröffentlicht: (2024)
von: Shu, Yuyang, et al.
Veröffentlicht: (2024)
Fine-Grained Action Segmentation for Renorrhaphy in Robot-Assisted Partial Nephrectomy
von: Dai, Jiaheng, et al.
Veröffentlicht: (2026)
von: Dai, Jiaheng, et al.
Veröffentlicht: (2026)
ViTOC: Vision Transformer and Object-aware Captioner
von: Huang, Feiyang
Veröffentlicht: (2024)
von: Huang, Feiyang
Veröffentlicht: (2024)
ACC-ViT : Atrous Convolution's Comeback in Vision Transformers
von: Ibtehaz, Nabil, et al.
Veröffentlicht: (2024)
von: Ibtehaz, Nabil, et al.
Veröffentlicht: (2024)
EA-ViT: Efficient Adaptation for Elastic Vision Transformer
von: Zhu, Chen, et al.
Veröffentlicht: (2025)
von: Zhu, Chen, et al.
Veröffentlicht: (2025)
ViTCN: Vision Transformer Contrastive Network For Reasoning
von: Song, Bo, et al.
Veröffentlicht: (2024)
von: Song, Bo, et al.
Veröffentlicht: (2024)
Towards Scalable Modeling of Compressed Videos for Efficient Action Recognition
von: Biswas, Shristi Das, et al.
Veröffentlicht: (2025)
von: Biswas, Shristi Das, et al.
Veröffentlicht: (2025)
LoViT: Long Video Transformer for Surgical Phase Recognition
von: Liu, Yang, et al.
Veröffentlicht: (2023)
von: Liu, Yang, et al.
Veröffentlicht: (2023)
ViT-AdaLA: Adapting Vision Transformers with Linear Attention
von: Li, Yifan, et al.
Veröffentlicht: (2026)
von: Li, Yifan, et al.
Veröffentlicht: (2026)
UniViTAR: Unified Vision Transformer with Native Resolution
von: Qiao, Limeng, et al.
Veröffentlicht: (2025)
von: Qiao, Limeng, et al.
Veröffentlicht: (2025)
ChangeViT: Unleashing Plain Vision Transformers for Change Detection
von: Zhu, Duowang, et al.
Veröffentlicht: (2024)
von: Zhu, Duowang, et al.
Veröffentlicht: (2024)
ViTGaze: Gaze Following with Interaction Features in Vision Transformers
von: Song, Yuehao, et al.
Veröffentlicht: (2024)
von: Song, Yuehao, et al.
Veröffentlicht: (2024)
ThinkingViT: Matryoshka Thinking Vision Transformer for Elastic Inference
von: Hojjat, Ali, et al.
Veröffentlicht: (2025)
von: Hojjat, Ali, et al.
Veröffentlicht: (2025)
SurgLaVi: Large-Scale Hierarchical Dataset for Surgical Vision-Language Representation Learning
von: Perez, Alejandra, et al.
Veröffentlicht: (2025)
von: Perez, Alejandra, et al.
Veröffentlicht: (2025)
LoLA-SpecViT: Local Attention SwiGLU Vision Transformer with LoRA for Hyperspectral Imaging
von: Zidi, Fadi Abdeladhim, et al.
Veröffentlicht: (2025)
von: Zidi, Fadi Abdeladhim, et al.
Veröffentlicht: (2025)
ViT-FIQA: Assessing Face Image Quality using Vision Transformers
von: Atzori, Andrea, et al.
Veröffentlicht: (2025)
von: Atzori, Andrea, et al.
Veröffentlicht: (2025)
MPTQ-ViT: Mixed-Precision Post-Training Quantization for Vision Transformer
von: Tai, Yu-Shan, et al.
Veröffentlicht: (2024)
von: Tai, Yu-Shan, et al.
Veröffentlicht: (2024)
MAGIC: Multimodal Alignment & Grounding-aware Instruction Coreset for Vision-Language Models
von: Biswas, Shristi Das, et al.
Veröffentlicht: (2026)
von: Biswas, Shristi Das, et al.
Veröffentlicht: (2026)
FairViT: Fair Vision Transformer via Adaptive Masking
von: Tian, Bowei, et al.
Veröffentlicht: (2024)
von: Tian, Bowei, et al.
Veröffentlicht: (2024)
Multiscale Vision Transformers meet Bipartite Matching for efficient single-stage Action Localization
von: Ntinou, Ioanna, et al.
Veröffentlicht: (2023)
von: Ntinou, Ioanna, et al.
Veröffentlicht: (2023)
HIRI-ViT: Scaling Vision Transformer with High Resolution Inputs
von: Yao, Ting, et al.
Veröffentlicht: (2024)
von: Yao, Ting, et al.
Veröffentlicht: (2024)
ViTA-Seg: Vision Transformer for Amodal Segmentation in Robotics
von: Caramia, Donato, et al.
Veröffentlicht: (2025)
von: Caramia, Donato, et al.
Veröffentlicht: (2025)
ViTmiX: Vision Transformer Explainability Augmented by Mixed Visualization Methods
von: Hogea, Eduard, et al.
Veröffentlicht: (2024)
von: Hogea, Eduard, et al.
Veröffentlicht: (2024)
Language-Unlocked ViT (LUViT): Empowering Self-Supervised Vision Transformers with LLMs
von: Kuzucu, Selim, et al.
Veröffentlicht: (2025)
von: Kuzucu, Selim, et al.
Veröffentlicht: (2025)
ADFQ-ViT: Activation-Distribution-Friendly Post-Training Quantization for Vision Transformers
von: Jiang, Yanfeng, et al.
Veröffentlicht: (2024)
von: Jiang, Yanfeng, et al.
Veröffentlicht: (2024)
Hyb-KAN ViT: Hybrid Kolmogorov-Arnold Networks Augmented Vision Transformer
von: Dey, Sainath, et al.
Veröffentlicht: (2025)
von: Dey, Sainath, et al.
Veröffentlicht: (2025)
ReViT: Enhancing Vision Transformers Feature Diversity with Attention Residual Connections
von: Diko, Anxhelo, et al.
Veröffentlicht: (2024)
von: Diko, Anxhelo, et al.
Veröffentlicht: (2024)
SVD-ViT: Does SVD Make Vision Transformers Attend More to the Foreground?
von: Murata, Haruhiko, et al.
Veröffentlicht: (2026)
von: Murata, Haruhiko, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
Towards Visual Syntactical Understanding
von: Chowdhury, Sayeed Shafayet, et al.
Veröffentlicht: (2024) -
REMAP: Regularized Matching and Partial Alignment of Video Embeddings
von: Chandra, Soumyadeep, et al.
Veröffentlicht: (2025) -
2D-ThermAl: Physics-Informed Framework for Thermal Analysis of Circuits using Generative AI
von: Chandra, Soumyadeep, et al.
Veröffentlicht: (2025) -
Loss Knows Best: Detecting Annotation Errors in Videos via Loss Trajectories
von: Alwis, Praditha, et al.
Veröffentlicht: (2026) -
LocalViT: Analyzing Locality in Vision Transformers
von: Li, Yawei, et al.
Veröffentlicht: (2021)