VISAGE: Video Synthesis using Action Graphs for Surgery

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Yeganeh, Yousef, Lazuardi, Rachmadio, Shamseddin, Amir, Dari, Emine, Thirani, Yash, Navab, Nassir, Farshad, Azade
Formato: Preprint
Publicado: 2024
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866913567771459584
author Yeganeh, Yousef
Lazuardi, Rachmadio
Shamseddin, Amir
Dari, Emine
Thirani, Yash
Navab, Nassir
Farshad, Azade
author_facet Yeganeh, Yousef
Lazuardi, Rachmadio
Shamseddin, Amir
Dari, Emine
Thirani, Yash
Navab, Nassir
Farshad, Azade
contents Surgical data science (SDS) is a field that analyzes patient data before, during, and after surgery to improve surgical outcomes and skills. However, surgical data is scarce, heterogeneous, and complex, which limits the applicability of existing machine learning methods. In this work, we introduce the novel task of future video generation in laparoscopic surgery. This task can augment and enrich the existing surgical data and enable various applications, such as simulation, analysis, and robot-aided surgery. Ultimately, it involves not only understanding the current state of the operation but also accurately predicting the dynamic and often unpredictable nature of surgical procedures. Our proposed method, VISAGE (VIdeo Synthesis using Action Graphs for Surgery), leverages the power of action scene graphs to capture the sequential nature of laparoscopic procedures and utilizes diffusion models to synthesize temporally coherent video sequences. VISAGE predicts the future frames given only a single initial frame, and the action graph triplets. By incorporating domain-specific knowledge through the action graph, VISAGE ensures the generated videos adhere to the expected visual and motion patterns observed in real laparoscopic procedures. The results of our experiments demonstrate high-fidelity video generation for laparoscopy procedures, which enables various applications in SDS.
format Preprint
id arxiv_https___arxiv_org_abs_2410_17751
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle VISAGE: Video Synthesis using Action Graphs for Surgery
Yeganeh, Yousef
Lazuardi, Rachmadio
Shamseddin, Amir
Dari, Emine
Thirani, Yash
Navab, Nassir
Farshad, Azade
Computer Vision and Pattern Recognition
Artificial Intelligence
Machine Learning
Surgical data science (SDS) is a field that analyzes patient data before, during, and after surgery to improve surgical outcomes and skills. However, surgical data is scarce, heterogeneous, and complex, which limits the applicability of existing machine learning methods. In this work, we introduce the novel task of future video generation in laparoscopic surgery. This task can augment and enrich the existing surgical data and enable various applications, such as simulation, analysis, and robot-aided surgery. Ultimately, it involves not only understanding the current state of the operation but also accurately predicting the dynamic and often unpredictable nature of surgical procedures. Our proposed method, VISAGE (VIdeo Synthesis using Action Graphs for Surgery), leverages the power of action scene graphs to capture the sequential nature of laparoscopic procedures and utilizes diffusion models to synthesize temporally coherent video sequences. VISAGE predicts the future frames given only a single initial frame, and the action graph triplets. By incorporating domain-specific knowledge through the action graph, VISAGE ensures the generated videos adhere to the expected visual and motion patterns observed in real laparoscopic procedures. The results of our experiments demonstrate high-fidelity video generation for laparoscopy procedures, which enables various applications in SDS.
title VISAGE: Video Synthesis using Action Graphs for Surgery
topic Computer Vision and Pattern Recognition
Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2410.17751