SF-TMN: SlowFast Temporal Modeling Network for Surgical Phase Recognition
Fuente:
arXiv
Saved in:
| Main Authors: | Zhang, Bokai, Sarhan, Mohammad Hasan, Goel, Bharti, Petculescu, Svetlana, Ghanem, Amer |
|---|---|
| Format: | Preprint |
| Published: |
2023
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Friends Across Time: Multi-Scale Action Segmentation Transformer for Surgical Phase Recognition
by: Zhang, Bokai, et al.
Published: (2024)
by: Zhang, Bokai, et al.
Published: (2024)
StreamVLN: Streaming Vision-and-Language Navigation via SlowFast Context Modeling
by: Wei, Meng, et al.
Published: (2025)
by: Wei, Meng, et al.
Published: (2025)
SlowFast-VGen: Slow-Fast Learning for Action-Driven Long Video Generation
by: Hong, Yining, et al.
Published: (2024)
by: Hong, Yining, et al.
Published: (2024)
Slot-VLM: SlowFast Slots for Video-Language Modeling
by: Xu, Jiaqi, et al.
Published: (2024)
by: Xu, Jiaqi, et al.
Published: (2024)
SlowFast-SCI: Slow-Fast Deep Unfolding Learning for Spectral Compressive Imaging
by: Zeng, Haijin, et al.
Published: (2025)
by: Zeng, Haijin, et al.
Published: (2025)
PhysMamba: Efficient Remote Physiological Measurement with SlowFast Temporal Difference Mamba
by: Luo, Chaoqi, et al.
Published: (2024)
by: Luo, Chaoqi, et al.
Published: (2024)
WhisperNetV2: SlowFast Siamese Network For Lip-Based Biometrics
by: Zakeri, Abdollah, et al.
Published: (2024)
by: Zakeri, Abdollah, et al.
Published: (2024)
SFMViT: SlowFast Meet ViT in Chaotic World
by: Lin, Jiaying, et al.
Published: (2024)
by: Lin, Jiaying, et al.
Published: (2024)
Enhancing Visual Place Recognition via Fast and Slow Adaptive Biasing in Event Cameras
by: Nair, Gokul B., et al.
Published: (2024)
by: Nair, Gokul B., et al.
Published: (2024)
Language-Conditioned Robotic Manipulation with Fast and Slow Thinking
by: Zhu, Minjie, et al.
Published: (2024)
by: Zhu, Minjie, et al.
Published: (2024)
SlowFast-LLaVA: A Strong Training-Free Baseline for Video Large Language Models
by: Xu, Mingze, et al.
Published: (2024)
by: Xu, Mingze, et al.
Published: (2024)
Geo-RepNet: Geometry-Aware Representation Learning for Surgical Phase Recognition in Endoscopic Submucosal Dissection
by: Tang, Rui, et al.
Published: (2025)
by: Tang, Rui, et al.
Published: (2025)
SpikePingpong: Spike Vision-based Fast-Slow Pingpong Robot System
by: Wang, Hao, et al.
Published: (2025)
by: Wang, Hao, et al.
Published: (2025)
AutoLayout: Closed-Loop Layout Synthesis via Slow-Fast Collaborative Reasoning
by: Chen, Weixing, et al.
Published: (2025)
by: Chen, Weixing, et al.
Published: (2025)
Advancing Autonomous Driving Perception: Analysis of Sensor Fusion and Computer Vision Techniques
by: Bharti, Urvishkumar, et al.
Published: (2024)
by: Bharti, Urvishkumar, et al.
Published: (2024)
Incremental Multimodal Surface Mapping via Self-Organizing Gaussian Mixture Models
by: Goel, Kshitij, et al.
Published: (2023)
by: Goel, Kshitij, et al.
Published: (2023)
Egocentric Instruction-oriented Affordance Prediction via Large Multimodal Model
by: Ji, Bokai, et al.
Published: (2025)
by: Ji, Bokai, et al.
Published: (2025)
Multimodal Graph Representation Learning for Robust Surgical Workflow Recognition with Adversarial Feature Disentanglement
by: Bai, Long, et al.
Published: (2025)
by: Bai, Long, et al.
Published: (2025)
Open-RadVLAD: Fast and Robust Radar Place Recognition
by: Gadd, Matthew, et al.
Published: (2024)
by: Gadd, Matthew, et al.
Published: (2024)
SlowFast-LLaVA-1.5: A Family of Token-Efficient Video Large Language Models for Long-Form Video Understanding
by: Xu, Mingze, et al.
Published: (2025)
by: Xu, Mingze, et al.
Published: (2025)
Cosmos-H-Surgical: Learning Surgical Robot Policies from Videos via World Modeling
by: He, Yufan, et al.
Published: (2025)
by: He, Yufan, et al.
Published: (2025)
LS-HAR: Language Supervised Human Action Recognition with Salient Fusion, Construction Sites as a Use-Case
by: Mahdavian, Mohammad, et al.
Published: (2024)
by: Mahdavian, Mohammad, et al.
Published: (2024)
SF-Loc: A Visual Mapping and Geo-Localization System based on Sparse Visual Structure Frames
by: Zhou, Yuxuan, et al.
Published: (2024)
by: Zhou, Yuxuan, et al.
Published: (2024)
CRASH: Crash Recognition and Anticipation System Harnessing with Context-Aware and Temporal Focus Attentions
by: Liao, Haicheng, et al.
Published: (2024)
by: Liao, Haicheng, et al.
Published: (2024)
Efficient Event Camera Volume System
by: Soto, Juan Camilo, et al.
Published: (2026)
by: Soto, Juan Camilo, et al.
Published: (2026)
Surgformer: Surgical Transformer with Hierarchical Temporal Attention for Surgical Phase Recognition
by: Yang, Shu, et al.
Published: (2024)
by: Yang, Shu, et al.
Published: (2024)
Applications of Spiking Neural Networks in Visual Place Recognition
by: Hussaini, Somayeh, et al.
Published: (2023)
by: Hussaini, Somayeh, et al.
Published: (2023)
OSSAR: Towards Open-Set Surgical Activity Recognition in Robot-assisted Surgery
by: Bai, Long, et al.
Published: (2024)
by: Bai, Long, et al.
Published: (2024)
LIDEA: Human-to-Robot Imitation Learning via Implicit Feature Distillation and Explicit Geometry Alignment
by: Xu, Yifu, et al.
Published: (2026)
by: Xu, Yifu, et al.
Published: (2026)
Distance and Collision Probability Estimation from Gaussian Surface Models
by: Goel, Kshitij, et al.
Published: (2024)
by: Goel, Kshitij, et al.
Published: (2024)
Open-Source Multi-Viewpoint Surgical Telerobotics
by: Caccianiga, Guido, et al.
Published: (2025)
by: Caccianiga, Guido, et al.
Published: (2025)
SLNet: A Super-Lightweight Geometry-Adaptive Network for 3D Point Cloud Recognition
by: Saeid, Mohammad, et al.
Published: (2026)
by: Saeid, Mohammad, et al.
Published: (2026)
STEP: Spatial Temporal Graph Convolutional Networks for Emotion Perception from Gaits
by: Bhattacharya, Uttaran, et al.
Published: (2019)
by: Bhattacharya, Uttaran, et al.
Published: (2019)
EDENet: Echo Direction Encoding Network for Place Recognition Based on Ground Penetrating Radar
by: Zhang, Pengyu, et al.
Published: (2025)
by: Zhang, Pengyu, et al.
Published: (2025)
FASIONAD : FAst and Slow FusION Thinking Systems for Human-Like Autonomous Driving with Adaptive Feedback
by: Qian, Kangan, et al.
Published: (2024)
by: Qian, Kangan, et al.
Published: (2024)
Fast LiDAR Upsampling using Conditional Diffusion Models
by: Helgesen, Sander Elias Magnussen, et al.
Published: (2024)
by: Helgesen, Sander Elias Magnussen, et al.
Published: (2024)
PRAM: Place Recognition Anywhere Model for Efficient Visual Localization
by: Xue, Fei, et al.
Published: (2024)
by: Xue, Fei, et al.
Published: (2024)
Understanding Spatio-Temporal Relations in Human-Object Interaction using Pyramid Graph Convolutional Network
by: Xing, Hao, et al.
Published: (2024)
by: Xing, Hao, et al.
Published: (2024)
MPTF-Net: Multi-view Pyramid Transformer Fusion Network for LiDAR-based Place Recognition
by: Li, Shuyuan, et al.
Published: (2026)
by: Li, Shuyuan, et al.
Published: (2026)
LCPR: A Multi-Scale Attention-Based LiDAR-Camera Fusion Network for Place Recognition
by: Zhou, Zijie, et al.
Published: (2023)
by: Zhou, Zijie, et al.
Published: (2023)
Similar Items
-
Friends Across Time: Multi-Scale Action Segmentation Transformer for Surgical Phase Recognition
by: Zhang, Bokai, et al.
Published: (2024) -
StreamVLN: Streaming Vision-and-Language Navigation via SlowFast Context Modeling
by: Wei, Meng, et al.
Published: (2025) -
SlowFast-VGen: Slow-Fast Learning for Action-Driven Long Video Generation
by: Hong, Yining, et al.
Published: (2024) -
Slot-VLM: SlowFast Slots for Video-Language Modeling
by: Xu, Jiaqi, et al.
Published: (2024) -
SlowFast-SCI: Slow-Fast Deep Unfolding Learning for Spectral Compressive Imaging
by: Zeng, Haijin, et al.
Published: (2025)