SurgTEMP: Temporal-Aware Surgical Video Question Answering with Text-guided Visual Memory for Laparoscopic Cholecystectomy
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Li, Shi, Srivastav, Vinkle, Chanel, Nicolas, Sharma, Saurav, Banik, Nabani, Arboit, Lorenzo, Yuan, Kun, Mascagni, Pietro, Padoy, Nicolas |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Endoshare: A Publicly Available, Surgeons-Friendly Solution to De-Identify and Manage Surgical Videos
von: Arboit, Lorenzo, et al.
Veröffentlicht: (2025)
von: Arboit, Lorenzo, et al.
Veröffentlicht: (2025)
Procedure-Aware Surgical Video-language Pretraining with Hierarchical Knowledge Augmentation
von: Yuan, Kun, et al.
Veröffentlicht: (2024)
von: Yuan, Kun, et al.
Veröffentlicht: (2024)
Jumpstarting Surgical Computer Vision
von: Alapatt, Deepak, et al.
Veröffentlicht: (2023)
von: Alapatt, Deepak, et al.
Veröffentlicht: (2023)
HecVL: Hierarchical Video-Language Pretraining for Zero-shot Surgical Phase Recognition
von: Yuan, Kun, et al.
Veröffentlicht: (2024)
von: Yuan, Kun, et al.
Veröffentlicht: (2024)
Multi-modal Representations for Fine-grained Multi-label Critical View of Safety Recognition
von: Baby, Britty, et al.
Veröffentlicht: (2025)
von: Baby, Britty, et al.
Veröffentlicht: (2025)
Early Operative Difficulty Assessment in Laparoscopic Cholecystectomy via Snapshot-Centric Video Analysis
von: Sharma, Saurav, et al.
Veröffentlicht: (2025)
von: Sharma, Saurav, et al.
Veröffentlicht: (2025)
Learning Multi-modal Representations by Watching Hundreds of Surgical Video Lectures
von: Yuan, Kun, et al.
Veröffentlicht: (2023)
von: Yuan, Kun, et al.
Veröffentlicht: (2023)
Adaptation of Multi-modal Representation Models for Multi-task Surgical Computer Vision
von: Walimbe, Soham, et al.
Veröffentlicht: (2025)
von: Walimbe, Soham, et al.
Veröffentlicht: (2025)
Text-driven Adaptation of Foundation Models for Few-shot Surgical Workflow Analysis
von: Chen, Tingxuan, et al.
Veröffentlicht: (2025)
von: Chen, Tingxuan, et al.
Veröffentlicht: (2025)
CliPPER: Contextual Video-Language Pretraining on Long-form Intraoperative Surgical Procedures for Event Recognition
von: Stilz, Florian, et al.
Veröffentlicht: (2026)
von: Stilz, Florian, et al.
Veröffentlicht: (2026)
SelfPose3d: Self-Supervised Multi-Person Multi-View 3d Pose Estimation
von: Srivastav, Vinkle, et al.
Veröffentlicht: (2024)
von: Srivastav, Vinkle, et al.
Veröffentlicht: (2024)
Surgical Text-to-Image Generation
von: Nwoye, Chinedu Innocent, et al.
Veröffentlicht: (2024)
von: Nwoye, Chinedu Innocent, et al.
Veröffentlicht: (2024)
Advancing Surgical VQA with Scene Graph Knowledge
von: Yuan, Kun, et al.
Veröffentlicht: (2023)
von: Yuan, Kun, et al.
Veröffentlicht: (2023)
Learning from Synchronization: Self-Supervised Uncalibrated Multi-View Person Association in Challenging Scenes
von: Chen, Keqi, et al.
Veröffentlicht: (2025)
von: Chen, Keqi, et al.
Veröffentlicht: (2025)
End-to-End Learning of Multi-Organ Implicit Surfaces from 3D Medical Imaging Data
von: Zarin, Farahdiba, et al.
Veröffentlicht: (2025)
von: Zarin, Farahdiba, et al.
Veröffentlicht: (2025)
Overcoming Dimensional Collapse in Self-supervised Contrastive Learning for Medical Image Segmentation
von: Hassanpour, Jamshid, et al.
Veröffentlicht: (2024)
von: Hassanpour, Jamshid, et al.
Veröffentlicht: (2024)
fine-CLIP: Enhancing Zero-Shot Fine-Grained Surgical Action Recognition with Vision-Language Models
von: Sharma, Saurav, et al.
Veröffentlicht: (2025)
von: Sharma, Saurav, et al.
Veröffentlicht: (2025)
Multi-view Video-Pose Pretraining for Operating Room Surgical Activity Recognition
von: Hamoud, Idris, et al.
Veröffentlicht: (2025)
von: Hamoud, Idris, et al.
Veröffentlicht: (2025)
Optimizing Latent Graph Representations of Surgical Scenes for Zero-Shot Domain Transfer
von: Satyanaik, Siddhant, et al.
Veröffentlicht: (2024)
von: Satyanaik, Siddhant, et al.
Veröffentlicht: (2024)
CycleSAM: Few-Shot Surgical Scene Segmentation with Cycle- and Scene-Consistent Feature Matching
von: Murali, Aditya, et al.
Veröffentlicht: (2024)
von: Murali, Aditya, et al.
Veröffentlicht: (2024)
Learning from Sparse Point Labels for Dense Carcinosis Localization in Advanced Ovarian Cancer Assessment
von: Zarin, Farahdiba, et al.
Veröffentlicht: (2025)
von: Zarin, Farahdiba, et al.
Veröffentlicht: (2025)
Self-Supervised Uncalibrated Multi-View Video Anonymization in the Operating Room
von: Chen, Keqi, et al.
Veröffentlicht: (2026)
von: Chen, Keqi, et al.
Veröffentlicht: (2026)
A Skull-Adaptive Framework for AI-Based 3D Transcranial Focused Ultrasound Simulation
von: Srivastav, Vinkle, et al.
Veröffentlicht: (2025)
von: Srivastav, Vinkle, et al.
Veröffentlicht: (2025)
Expert Consensus-based Video-Based Assessment Tool for Workflow Analysis in Minimally Invasive Colorectal Surgery: Development and Validation of ColoWorkflow
von: Jain, Pooja P, et al.
Veröffentlicht: (2025)
von: Jain, Pooja P, et al.
Veröffentlicht: (2025)
Surgeons Awareness, Expectations, and Involvement with Artificial Intelligence: a Survey Pre and Post the GPT Era
von: Arboit, Lorenzo, et al.
Veröffentlicht: (2025)
von: Arboit, Lorenzo, et al.
Veröffentlicht: (2025)
Artificial Intelligence for the Assessment of Peritoneal Carcinosis during Diagnostic Laparoscopy for Advanced Ovarian Cancer
von: Oliva, Riccardo, et al.
Veröffentlicht: (2025)
von: Oliva, Riccardo, et al.
Veröffentlicht: (2025)
SurgViVQA: Temporally-Grounded Video Question Answering for Surgical Scene Understanding
von: Drago, Mauro Orazio, et al.
Veröffentlicht: (2025)
von: Drago, Mauro Orazio, et al.
Veröffentlicht: (2025)
State-Change Learning for Prediction of Future Events in Endoscopic Videos
von: Sharma, Saurav, et al.
Veröffentlicht: (2025)
von: Sharma, Saurav, et al.
Veröffentlicht: (2025)
Recognizing Surgical Phases Anywhere: Few-Shot Test-time Adaptation and Task-graph Guided Refinement
von: Yuan, Kun, et al.
Veröffentlicht: (2025)
von: Yuan, Kun, et al.
Veröffentlicht: (2025)
Where are they looking in the operating room?
von: Chen, Keqi, et al.
Veröffentlicht: (2026)
von: Chen, Keqi, et al.
Veröffentlicht: (2026)
When do they StOP?: A First Step Towards Automatically Identifying Team Communication in the Operating Room
von: Chen, Keqi, et al.
Veröffentlicht: (2025)
von: Chen, Keqi, et al.
Veröffentlicht: (2025)
SurgTPGS: Semantic 3D Surgical Scene Understanding with Text Promptable Gaussian Splatting
von: Huang, Yiming, et al.
Veröffentlicht: (2025)
von: Huang, Yiming, et al.
Veröffentlicht: (2025)
S4M: 4-points to Segment Anything
von: Meyer, Adrien, et al.
Veröffentlicht: (2025)
von: Meyer, Adrien, et al.
Veröffentlicht: (2025)
SurgLQA: Scalable Long-Horizon Surgical Video Question Answering
von: Guo, Diandian, et al.
Veröffentlicht: (2026)
von: Guo, Diandian, et al.
Veröffentlicht: (2026)
SurgiTrack: Fine-Grained Multi-Class Multi-Tool Tracking in Surgical Videos
von: Nwoye, Chinedu Innocent, et al.
Veröffentlicht: (2024)
von: Nwoye, Chinedu Innocent, et al.
Veröffentlicht: (2024)
Laparoscopic Cholecystectomy in Situs Inversus Totalis: A Case Report
von: Prapti Lakhey, et al.
Veröffentlicht: (2025)
von: Prapti Lakhey, et al.
Veröffentlicht: (2025)
The Endoscapes Dataset for Surgical Scene Segmentation, Object Detection, and Critical View of Safety Assessment: Official Splits and Benchmark
von: Murali, Aditya, et al.
Veröffentlicht: (2023)
von: Murali, Aditya, et al.
Veröffentlicht: (2023)
Building Surgical Capacity Through Multidisciplinary Laparoscopic Cholecystectomy Training in Kampong Cham, Cambodia
von: Leif Manuel Sorensen, et al.
Veröffentlicht: (2026)
von: Leif Manuel Sorensen, et al.
Veröffentlicht: (2026)
SurgVidLM: Towards Multi-grained Surgical Video Understanding with Large Language Model
von: Wang, Guankun, et al.
Veröffentlicht: (2025)
von: Wang, Guankun, et al.
Veröffentlicht: (2025)
RoboSurg-VQA: A Multimodal Benchmark for Surgical Segmentation-Aware Visual Question Answering
von: Zhang, Chengyi, et al.
Veröffentlicht: (2026)
von: Zhang, Chengyi, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
Endoshare: A Publicly Available, Surgeons-Friendly Solution to De-Identify and Manage Surgical Videos
von: Arboit, Lorenzo, et al.
Veröffentlicht: (2025) -
Procedure-Aware Surgical Video-language Pretraining with Hierarchical Knowledge Augmentation
von: Yuan, Kun, et al.
Veröffentlicht: (2024) -
Jumpstarting Surgical Computer Vision
von: Alapatt, Deepak, et al.
Veröffentlicht: (2023) -
HecVL: Hierarchical Video-Language Pretraining for Zero-shot Surgical Phase Recognition
von: Yuan, Kun, et al.
Veröffentlicht: (2024) -
Multi-modal Representations for Fine-grained Multi-label Critical View of Safety Recognition
von: Baby, Britty, et al.
Veröffentlicht: (2025)