Watch and Learn: Leveraging Expert Knowledge and Language for Surgical Video Understanding
Fuente:
arXiv
Saved in:
| Main Authors: | Gastager, David, Ghazaei, Ghazal, Patsch, Constantin |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
SURGIVID: Annotation-Efficient Surgical Video Object Discovery
by: Köksal, Çağhan, et al.
Published: (2024)
by: Köksal, Çağhan, et al.
Published: (2024)
SANGRIA: Surgical Video Scene Graph Optimization for Surgical Workflow Prediction
by: Köksal, Çağhan, et al.
Published: (2024)
by: Köksal, Çağhan, et al.
Published: (2024)
ProtoFlow: Interpretable and Robust Surgical Workflow Modeling with Learned Dynamic Scene Graph Prototypes
by: Holm, Felix, et al.
Published: (2025)
by: Holm, Felix, et al.
Published: (2025)
Towards Comprehensive Real-Time Scene Understanding in Ophthalmic Surgery through Multimodal Image Fusion
by: Rohrmoser, Nikolo, et al.
Published: (2026)
by: Rohrmoser, Nikolo, et al.
Published: (2026)
SG2VID: Scene Graphs Enable Fine-Grained Control for Video Synthesis
by: Sivakumar, Ssharvien Kumar, et al.
Published: (2025)
by: Sivakumar, Ssharvien Kumar, et al.
Published: (2025)
Outside Knowledge Conversational Video (OKCV) Dataset -- Dialoguing over Videos
by: Reichman, Benjamin, et al.
Published: (2025)
by: Reichman, Benjamin, et al.
Published: (2025)
CAT-SG: A Large Dynamic Scene Graph Dataset for Fine-Grained Understanding of Cataract Surgery
by: Holm, Felix, et al.
Published: (2025)
by: Holm, Felix, et al.
Published: (2025)
Technical Report for Egocentric Mistake Detection for the HoloAssist Challenge
by: Patsch, Constantin, et al.
Published: (2025)
by: Patsch, Constantin, et al.
Published: (2025)
SurGrID: Controllable Surgical Simulation via Scene Graph to Image Diffusion
by: Frisch, Yannik, et al.
Published: (2025)
by: Frisch, Yannik, et al.
Published: (2025)
Learning Multi-modal Representations by Watching Hundreds of Surgical Video Lectures
by: Yuan, Kun, et al.
Published: (2023)
by: Yuan, Kun, et al.
Published: (2023)
Efficient Remote Sensing Change Detection with Change State Space Models
by: Ghazaei, Elman, et al.
Published: (2025)
by: Ghazaei, Elman, et al.
Published: (2025)
Text-conditioned State Space Model For Domain-generalized Change Detection Visual Question Answering
by: Ghazaei, Elman, et al.
Published: (2025)
by: Ghazaei, Elman, et al.
Published: (2025)
WAT: Online Video Understanding Needs Watching Before Thinking
by: Han, Zifan, et al.
Published: (2026)
by: Han, Zifan, et al.
Published: (2026)
SurgFed: Language-guided Multi-Task Federated Learning for Surgical Video Understanding
by: Fang, Zheng, et al.
Published: (2026)
by: Fang, Zheng, et al.
Published: (2026)
Surgical Video Understanding with Label Interpolation
by: Kim, Garam, et al.
Published: (2025)
by: Kim, Garam, et al.
Published: (2025)
How Well Can General Vision-Language Models Learn Medicine By Watching Public Educational Videos?
by: Thapa, Rahul, et al.
Published: (2025)
by: Thapa, Rahul, et al.
Published: (2025)
Speak While Watching: Unleashing TRUE Real-Time Video Understanding Capability of Multimodal Large Language Models
by: Lin, Junyan, et al.
Published: (2026)
by: Lin, Junyan, et al.
Published: (2026)
TennisExpert: Towards Expert-Level Analytical Sports Video Understanding
by: Liu, Zhaoyu, et al.
Published: (2026)
by: Liu, Zhaoyu, et al.
Published: (2026)
VideoExpert: Augmented LLM for Temporal-Sensitive Video Understanding
by: Zhao, Henghao, et al.
Published: (2025)
by: Zhao, Henghao, et al.
Published: (2025)
Watch, Listen, Understand, Mislead: Tri-modal Adversarial Attacks on Short Videos for Content Appropriateness Evaluation
by: Mustakim, Sahid Hossain, et al.
Published: (2025)
by: Mustakim, Sahid Hossain, et al.
Published: (2025)
Frame2Freq: Spectral Adapters for Fine-Grained Video Understanding
by: Ponbagavathi, Thinesh Thiyakesan, et al.
Published: (2026)
by: Ponbagavathi, Thinesh Thiyakesan, et al.
Published: (2026)
Video Streaming Thinking: VideoLLMs Can Watch and Think Simultaneously
by: Guan, Yiran, et al.
Published: (2026)
by: Guan, Yiran, et al.
Published: (2026)
Watch and Learn: Learning to Use Computers from Online Videos
by: Song, Chan Hee, et al.
Published: (2025)
by: Song, Chan Hee, et al.
Published: (2025)
Instrument-tissue Interaction Detection Framework for Surgical Video Understanding
by: Lin, Wenjun, et al.
Published: (2024)
by: Lin, Wenjun, et al.
Published: (2024)
An Approach to Enriching Surgical Video Datasets for Fine-Grained Spatial-Temporal Understanding of Vision-Language Models
by: Maack, Lennart, et al.
Published: (2026)
by: Maack, Lennart, et al.
Published: (2026)
SurgMotion: A Video-Native Foundation Model for Universal Understanding of Surgical Videos
by: Wu, Jinlin, et al.
Published: (2026)
by: Wu, Jinlin, et al.
Published: (2026)
Data-Efficient Learning for Generalizable Surgical Video Understanding
by: Nasirihaghighi, Sahar
Published: (2025)
by: Nasirihaghighi, Sahar
Published: (2025)
Magic Fixup: Streamlining Photo Editing by Watching Dynamic Videos
by: Alzayer, Hadi, et al.
Published: (2024)
by: Alzayer, Hadi, et al.
Published: (2024)
Token Merging via Spatiotemporal Information Mining for Surgical Video Understanding
by: Jiang, Xixi, et al.
Published: (2025)
by: Jiang, Xixi, et al.
Published: (2025)
HIPPO-Video: Simulating Watch Histories with Large Language Models for Personalized Video Highlighting
by: Lee, Jeongeun, et al.
Published: (2025)
by: Lee, Jeongeun, et al.
Published: (2025)
EgoExo-Gen: Ego-centric Video Prediction by Watching Exo-centric Videos
by: Xu, Jilan, et al.
Published: (2025)
by: Xu, Jilan, et al.
Published: (2025)
ADL4D: Towards A Contextually Rich Dataset for 4D Activities of Daily Living
by: Zakour, Marsil, et al.
Published: (2024)
by: Zakour, Marsil, et al.
Published: (2024)
CauCLIP: Bridging the Sim-to-Real Gap in Surgical Video Understanding via Causality-Inspired Vision-Language Modeling
by: He, Yuxin, et al.
Published: (2026)
by: He, Yuxin, et al.
Published: (2026)
OphCLIP: Hierarchical Retrieval-Augmented Learning for Ophthalmic Surgical Video-Language Pretraining
by: Hu, Ming, et al.
Published: (2024)
by: Hu, Ming, et al.
Published: (2024)
Watching Movies Like a Human: Egocentric Emotion Understanding for Embodied Companions
by: Dong, Ze, et al.
Published: (2026)
by: Dong, Ze, et al.
Published: (2026)
Distilling Expert Surgical Knowledge: How to train local surgical VLMs for anatomy explanation in Complete Mesocolic Excision
by: Maack, Lennart, et al.
Published: (2025)
by: Maack, Lennart, et al.
Published: (2025)
SurgViVQA: Temporally-Grounded Video Question Answering for Surgical Scene Understanding
by: Drago, Mauro Orazio, et al.
Published: (2025)
by: Drago, Mauro Orazio, et al.
Published: (2025)
OphNet: A Large-Scale Video Benchmark for Ophthalmic Surgical Workflow Understanding
by: Hu, Ming, et al.
Published: (2024)
by: Hu, Ming, et al.
Published: (2024)
Affordance-First Decomposition for Continual Learning in Video-Language Understanding
by: Xu, Mengzhu, et al.
Published: (2025)
by: Xu, Mengzhu, et al.
Published: (2025)
ExpertEdit: Learning Skill-Aware Motion Editing from Expert Videos
by: Somayazulu, Arjun, et al.
Published: (2026)
by: Somayazulu, Arjun, et al.
Published: (2026)
Similar Items
-
SURGIVID: Annotation-Efficient Surgical Video Object Discovery
by: Köksal, Çağhan, et al.
Published: (2024) -
SANGRIA: Surgical Video Scene Graph Optimization for Surgical Workflow Prediction
by: Köksal, Çağhan, et al.
Published: (2024) -
ProtoFlow: Interpretable and Robust Surgical Workflow Modeling with Learned Dynamic Scene Graph Prototypes
by: Holm, Felix, et al.
Published: (2025) -
Towards Comprehensive Real-Time Scene Understanding in Ophthalmic Surgery through Multimodal Image Fusion
by: Rohrmoser, Nikolo, et al.
Published: (2026) -
SG2VID: Scene Graphs Enable Fine-Grained Control for Video Synthesis
by: Sivakumar, Ssharvien Kumar, et al.
Published: (2025)