CAT-SG: A Large Dynamic Scene Graph Dataset for Fine-Grained Understanding of Cataract Surgery
Fuente:
arXiv
Saved in:
| Main Authors: | Holm, Felix, Ünver, Gözde, Ghazaei, Ghazal, Navab, Nassir |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
ProtoFlow: Interpretable and Robust Surgical Workflow Modeling with Learned Dynamic Scene Graph Prototypes
by: Holm, Felix, et al.
Published: (2025)
by: Holm, Felix, et al.
Published: (2025)
SANGRIA: Surgical Video Scene Graph Optimization for Surgical Workflow Prediction
by: Köksal, Çağhan, et al.
Published: (2024)
by: Köksal, Çağhan, et al.
Published: (2024)
Towards Comprehensive Real-Time Scene Understanding in Ophthalmic Surgery through Multimodal Image Fusion
by: Rohrmoser, Nikolo, et al.
Published: (2026)
by: Rohrmoser, Nikolo, et al.
Published: (2026)
SG2VID: Scene Graphs Enable Fine-Grained Control for Video Synthesis
by: Sivakumar, Ssharvien Kumar, et al.
Published: (2025)
by: Sivakumar, Ssharvien Kumar, et al.
Published: (2025)
SURGIVID: Annotation-Efficient Surgical Video Object Discovery
by: Köksal, Çağhan, et al.
Published: (2024)
by: Köksal, Çağhan, et al.
Published: (2024)
Advancing Surgical VQA with Scene Graph Knowledge
by: Yuan, Kun, et al.
Published: (2023)
by: Yuan, Kun, et al.
Published: (2023)
Prototype-Based Knowledge Guidance for Fine-Grained Structured Radiology Reporting
by: Pellegrini, Chantal, et al.
Published: (2026)
by: Pellegrini, Chantal, et al.
Published: (2026)
SG-Bot: Object Rearrangement via Coarse-to-Fine Robotic Imagination on Scene Graphs
by: Zhai, Guangyao, et al.
Published: (2023)
by: Zhai, Guangyao, et al.
Published: (2023)
VISAGE: Video Synthesis using Action Graphs for Surgery
by: Yeganeh, Yousef, et al.
Published: (2024)
by: Yeganeh, Yousef, et al.
Published: (2024)
EchoScene: Indoor Scene Generation via Information Echo over Scene Graph Diffusion
by: Zhai, Guangyao, et al.
Published: (2024)
by: Zhai, Guangyao, et al.
Published: (2024)
Semantic Scene Graph for Ultrasound Image Explanation and Scanning Guidance
by: Li, Xuesong, et al.
Published: (2025)
by: Li, Xuesong, et al.
Published: (2025)
Location-Free Scene Graph Generation
by: Özsoy, Ege, et al.
Published: (2023)
by: Özsoy, Ege, et al.
Published: (2023)
Watch and Learn: Leveraging Expert Knowledge and Language for Surgical Video Understanding
by: Gastager, David, et al.
Published: (2025)
by: Gastager, David, et al.
Published: (2025)
HecVL: Hierarchical Video-Language Pretraining for Zero-shot Surgical Phase Recognition
by: Yuan, Kun, et al.
Published: (2024)
by: Yuan, Kun, et al.
Published: (2024)
CliPPER: Contextual Video-Language Pretraining on Long-form Intraoperative Surgical Procedures for Event Recognition
by: Stilz, Florian, et al.
Published: (2026)
by: Stilz, Florian, et al.
Published: (2026)
Procedure-Aware Surgical Video-language Pretraining with Hierarchical Knowledge Augmentation
by: Yuan, Kun, et al.
Published: (2024)
by: Yuan, Kun, et al.
Published: (2024)
Fine-Grained Scene Graph Generation via Sample-Level Bias Prediction
by: Li, Yansheng, et al.
Published: (2024)
by: Li, Yansheng, et al.
Published: (2024)
LLaVA-SG: Leveraging Scene Graphs as Visual Semantic Expression in Vision-Language Models
by: Wang, Jingyi, et al.
Published: (2024)
by: Wang, Jingyi, et al.
Published: (2024)
From Linear Probing to Joint-Weighted Token Hierarchy: A Foundation Model Bridging Global and Cellular Representations in Biomarker Detection
by: Liu, Jingsong, et al.
Published: (2025)
by: Liu, Jingsong, et al.
Published: (2025)
Speckle2Self: Self-Supervised Ultrasound Speckle Reduction Without Clean Data
by: Li, Xuesong, et al.
Published: (2025)
by: Li, Xuesong, et al.
Published: (2025)
Text-driven Adaptation of Foundation Models for Few-shot Surgical Workflow Analysis
by: Chen, Tingxuan, et al.
Published: (2025)
by: Chen, Tingxuan, et al.
Published: (2025)
MEET: A Million-Scale Dataset for Fine-Grained Geospatial Scene Classification with Zoom-Free Remote Sensing Imagery
by: Li, Yansheng, et al.
Published: (2025)
by: Li, Yansheng, et al.
Published: (2025)
A Large-Scale Multimodal Dataset and Benchmarks for Human Activity Scene Understanding and Reasoning
by: Jiang, Siyang, et al.
Published: (2025)
by: Jiang, Siyang, et al.
Published: (2025)
SG-Tailor: Inter-Object Commonsense Relationship Reasoning for Scene Graph Manipulation
by: Shang, Haoliang, et al.
Published: (2025)
by: Shang, Haoliang, et al.
Published: (2025)
SurgMLLMBench: A Multimodal Large Language Model Benchmark Dataset for Surgical Scene Understanding
by: Choi, Tae-Min, et al.
Published: (2025)
by: Choi, Tae-Min, et al.
Published: (2025)
Robotic Ultrasound Makes CBCT Alive
by: Li, Feng, et al.
Published: (2026)
by: Li, Feng, et al.
Published: (2026)
SurgVidLM: Towards Multi-grained Surgical Video Understanding with Large Language Model
by: Wang, Guankun, et al.
Published: (2025)
by: Wang, Guankun, et al.
Published: (2025)
ReefNet: A Large-Scale Dataset and Benchmark for Fine-Grained Coral Reef Recognition
by: Felemban, Abdulwahab, et al.
Published: (2025)
by: Felemban, Abdulwahab, et al.
Published: (2025)
MMDocBench: Benchmarking Large Vision-Language Models for Fine-Grained Visual Document Understanding
by: Zhu, Fengbin, et al.
Published: (2024)
by: Zhu, Fengbin, et al.
Published: (2024)
SurgLLM: A Versatile Large Multimodal Model with Spatial Focus and Temporal Awareness for Surgical Video Understanding
by: Chen, Zhen, et al.
Published: (2025)
by: Chen, Zhen, et al.
Published: (2025)
Data-Efficient Surgical Phase Segmentation in Small-Incision Cataract Surgery: A Controlled Study of Vision Foundation Models
by: Spencer, Lincoln, et al.
Published: (2026)
by: Spencer, Lincoln, et al.
Published: (2026)
SkySenseGPT: A Fine-Grained Instruction Tuning Dataset and Model for Remote Sensing Vision-Language Understanding
by: Luo, Junwei, et al.
Published: (2024)
by: Luo, Junwei, et al.
Published: (2024)
Measuring Prediction Uncertainty in Neural Cellular Automata
by: Sadafi, Ario, et al.
Published: (2026)
by: Sadafi, Ario, et al.
Published: (2026)
DeepAf: One-Shot Spatiospectral Auto-Focus Model for Digital Pathology
by: Yeganeh, Yousef, et al.
Published: (2025)
by: Yeganeh, Yousef, et al.
Published: (2025)
Region-Grounded Report Generation for 3D Medical Imaging: A Fine-Grained Dataset and Graph-Enhanced Framework
by: Nguyen, Cong Huy, et al.
Published: (2026)
by: Nguyen, Cong Huy, et al.
Published: (2026)
FD$^2$: A Dedicated Framework for Fine-Grained Dataset Distillation
by: Ma, Hongxu, et al.
Published: (2026)
by: Ma, Hongxu, et al.
Published: (2026)
Adaptive Self-training Framework for Fine-grained Scene Graph Generation
by: Kim, Kibum, et al.
Published: (2024)
by: Kim, Kibum, et al.
Published: (2024)
STAR: A First-Ever Dataset and A Large-Scale Benchmark for Scene Graph Generation in Large-Size Satellite Imagery
by: Li, Yansheng, et al.
Published: (2024)
by: Li, Yansheng, et al.
Published: (2024)
VidEvent: A Large Dataset for Understanding Dynamic Evolution of Events in Videos
by: Liang, Baoyu, et al.
Published: (2025)
by: Liang, Baoyu, et al.
Published: (2025)
CGC: Compositional Grounded Contrast for Fine-Grained Multi-Image Understanding
by: Zheng, Lihao, et al.
Published: (2026)
by: Zheng, Lihao, et al.
Published: (2026)
Similar Items
-
ProtoFlow: Interpretable and Robust Surgical Workflow Modeling with Learned Dynamic Scene Graph Prototypes
by: Holm, Felix, et al.
Published: (2025) -
SANGRIA: Surgical Video Scene Graph Optimization for Surgical Workflow Prediction
by: Köksal, Çağhan, et al.
Published: (2024) -
Towards Comprehensive Real-Time Scene Understanding in Ophthalmic Surgery through Multimodal Image Fusion
by: Rohrmoser, Nikolo, et al.
Published: (2026) -
SG2VID: Scene Graphs Enable Fine-Grained Control for Video Synthesis
by: Sivakumar, Ssharvien Kumar, et al.
Published: (2025) -
SURGIVID: Annotation-Efficient Surgical Video Object Discovery
by: Köksal, Çağhan, et al.
Published: (2024)