ORacle: Large Vision-Language Models for Knowledge-Guided Holistic OR Domain Modeling
Fuente:
arXiv
Salvato in:
| Autori principali: | Özsoy, Ege, Pellegrini, Chantal, Keicher, Matthias, Navab, Nassir |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
RaDialog: A Large Vision-Language Model for Radiology Report Generation and Conversational Assistance
di: Pellegrini, Chantal, et al.
Pubblicazione: (2023)
di: Pellegrini, Chantal, et al.
Pubblicazione: (2023)
Prototype-Based Knowledge Guidance for Fine-Grained Structured Radiology Reporting
di: Pellegrini, Chantal, et al.
Pubblicazione: (2026)
di: Pellegrini, Chantal, et al.
Pubblicazione: (2026)
Specialized Foundation Models for Intelligent Operating Rooms
di: Özsoy, Ege, et al.
Pubblicazione: (2025)
di: Özsoy, Ege, et al.
Pubblicazione: (2025)
MM-OR: A Large Multimodal Operating Room Dataset for Semantic Understanding of High-Intensity Surgical Environments
di: Özsoy, Ege, et al.
Pubblicazione: (2025)
di: Özsoy, Ege, et al.
Pubblicazione: (2025)
Location-Free Scene Graph Generation
di: Özsoy, Ege, et al.
Pubblicazione: (2023)
di: Özsoy, Ege, et al.
Pubblicazione: (2023)
EgoExOR: An Ego-Exo-Centric Operating Room Dataset for Surgical Activity Understanding
di: Özsoy, Ege, et al.
Pubblicazione: (2025)
di: Özsoy, Ege, et al.
Pubblicazione: (2025)
EHR2Path: Scalable Modeling of Longitudinal Patient Pathways from Multimodal Electronic Health Records
di: Pellegrini, Chantal, et al.
Pubblicazione: (2025)
di: Pellegrini, Chantal, et al.
Pubblicazione: (2025)
Language Agents for Hypothesis-driven Clinical Decision Making with Reinforcement Learning
di: Bani-Harouni, David, et al.
Pubblicazione: (2025)
di: Bani-Harouni, David, et al.
Pubblicazione: (2025)
PanORama: Multiview Consistent Panoptic Segmentation in Operating Rooms
di: Gürbüz, Tuna, et al.
Pubblicazione: (2026)
di: Gürbüz, Tuna, et al.
Pubblicazione: (2026)
Rewarding Doubt: A Reinforcement Learning Approach to Calibrated Confidence Expression of Large Language Models
di: Bani-Harouni, David, et al.
Pubblicazione: (2025)
di: Bani-Harouni, David, et al.
Pubblicazione: (2025)
Language-Guided Open-World Anomaly Segmentation
di: Reichard, Klara, et al.
Pubblicazione: (2025)
di: Reichard, Klara, et al.
Pubblicazione: (2025)
Recognizing Surgical Phases Anywhere: Few-Shot Test-time Adaptation and Task-graph Guided Refinement
di: Yuan, Kun, et al.
Pubblicazione: (2025)
di: Yuan, Kun, et al.
Pubblicazione: (2025)
HieraSurg: Hierarchy-Aware Diffusion Model for Surgical Video Generation
di: Biagini, Diego, et al.
Pubblicazione: (2025)
di: Biagini, Diego, et al.
Pubblicazione: (2025)
Counterfactual Explanations for Medical Image Classification and Regression using Diffusion Autoencoder
di: Atad, Matan, et al.
Pubblicazione: (2024)
di: Atad, Matan, et al.
Pubblicazione: (2024)
Beyond Role-Based Surgical Domain Modeling: Generalizable Re-Identification in the Operating Room
di: Wang, Tony Danjun, et al.
Pubblicazione: (2025)
di: Wang, Tony Danjun, et al.
Pubblicazione: (2025)
SURGIVID: Annotation-Efficient Surgical Video Object Discovery
di: Köksal, Çağhan, et al.
Pubblicazione: (2024)
di: Köksal, Çağhan, et al.
Pubblicazione: (2024)
MonoGSDF: Exploring Monocular Geometric Cues for Gaussian Splatting-Guided Implicit Surface Reconstruction
di: Li, Kunyi, et al.
Pubblicazione: (2024)
di: Li, Kunyi, et al.
Pubblicazione: (2024)
Procedure-Aware Surgical Video-language Pretraining with Hierarchical Knowledge Augmentation
di: Yuan, Kun, et al.
Pubblicazione: (2024)
di: Yuan, Kun, et al.
Pubblicazione: (2024)
Temporal Differential Fields for 4D Motion Modeling via Image-to-Video Synthesis
di: You, Xin, et al.
Pubblicazione: (2025)
di: You, Xin, et al.
Pubblicazione: (2025)
Visual Autoregressive Modelling for Monocular Depth Estimation
di: El-Ghoussani, Amir, et al.
Pubblicazione: (2025)
di: El-Ghoussani, Amir, et al.
Pubblicazione: (2025)
VALOR-EVAL: Holistic Coverage and Faithfulness Evaluation of Large Vision-Language Models
di: Qiu, Haoyi, et al.
Pubblicazione: (2024)
di: Qiu, Haoyi, et al.
Pubblicazione: (2024)
GALA: Guided Attention with Language Alignment for Open Vocabulary Gaussian Splatting
di: Alegret, Elena, et al.
Pubblicazione: (2025)
di: Alegret, Elena, et al.
Pubblicazione: (2025)
Neural Semantic Map-Learning for Autonomous Vehicles
di: Herb, Markus, et al.
Pubblicazione: (2024)
di: Herb, Markus, et al.
Pubblicazione: (2024)
GLIMPSE: Holistic Cross-Modal Explainability for Large Vision-Language Models
di: Shen, Guanxi
Pubblicazione: (2025)
di: Shen, Guanxi
Pubblicazione: (2025)
HoVLE: Unleashing the Power of Monolithic Vision-Language Models with Holistic Vision-Language Embedding
di: Tao, Chenxin, et al.
Pubblicazione: (2024)
di: Tao, Chenxin, et al.
Pubblicazione: (2024)
ProtoFlow: Interpretable and Robust Surgical Workflow Modeling with Learned Dynamic Scene Graph Prototypes
di: Holm, Felix, et al.
Pubblicazione: (2025)
di: Holm, Felix, et al.
Pubblicazione: (2025)
HecVL: Hierarchical Video-Language Pretraining for Zero-shot Surgical Phase Recognition
di: Yuan, Kun, et al.
Pubblicazione: (2024)
di: Yuan, Kun, et al.
Pubblicazione: (2024)
OCRVerse: Towards Holistic OCR in End-to-End Vision-Language Models
di: Zhong, Yufeng, et al.
Pubblicazione: (2026)
di: Zhong, Yufeng, et al.
Pubblicazione: (2026)
GLOV: Guided Large Language Models as Implicit Optimizers for Vision Language Models
di: Mirza, M. Jehanzeb, et al.
Pubblicazione: (2024)
di: Mirza, M. Jehanzeb, et al.
Pubblicazione: (2024)
Deep Spectral Methods for Unsupervised Ultrasound Image Interpretation
di: Tmenova, Oleksandra, et al.
Pubblicazione: (2024)
di: Tmenova, Oleksandra, et al.
Pubblicazione: (2024)
ESCAPE: Equivariant Shape Completion via Anchor Point Encoding
di: Bekci, Burak, et al.
Pubblicazione: (2024)
di: Bekci, Burak, et al.
Pubblicazione: (2024)
CliPPER: Contextual Video-Language Pretraining on Long-form Intraoperative Surgical Procedures for Event Recognition
di: Stilz, Florian, et al.
Pubblicazione: (2026)
di: Stilz, Florian, et al.
Pubblicazione: (2026)
A Taxonomy and Library for Visualizing Learned Features in Convolutional Neural Networks
di: Grün, Felix, et al.
Pubblicazione: (2016)
di: Grün, Felix, et al.
Pubblicazione: (2016)
Hybrid Functional Maps for Crease-Aware Non-Isometric Shape Matching
di: Bastian, Lennart, et al.
Pubblicazione: (2023)
di: Bastian, Lennart, et al.
Pubblicazione: (2023)
Forecasting Continuous Non-Conservative Dynamical Systems in SO(3)
di: Bastian, Lennart, et al.
Pubblicazione: (2025)
di: Bastian, Lennart, et al.
Pubblicazione: (2025)
MS-CLR: Multi-Skeleton Contrastive Learning for Human Action Recognition
di: Kiray, Mert, et al.
Pubblicazione: (2025)
di: Kiray, Mert, et al.
Pubblicazione: (2025)
Towards Comprehensive Real-Time Scene Understanding in Ophthalmic Surgery through Multimodal Image Fusion
di: Rohrmoser, Nikolo, et al.
Pubblicazione: (2026)
di: Rohrmoser, Nikolo, et al.
Pubblicazione: (2026)
From Linear Probing to Joint-Weighted Token Hierarchy: A Foundation Model Bridging Global and Cellular Representations in Biomarker Detection
di: Liu, Jingsong, et al.
Pubblicazione: (2025)
di: Liu, Jingsong, et al.
Pubblicazione: (2025)
VHELM: A Holistic Evaluation of Vision Language Models
di: Lee, Tony, et al.
Pubblicazione: (2024)
di: Lee, Tony, et al.
Pubblicazione: (2024)
Text-driven Adaptation of Foundation Models for Few-shot Surgical Workflow Analysis
di: Chen, Tingxuan, et al.
Pubblicazione: (2025)
di: Chen, Tingxuan, et al.
Pubblicazione: (2025)
Documenti analoghi
-
RaDialog: A Large Vision-Language Model for Radiology Report Generation and Conversational Assistance
di: Pellegrini, Chantal, et al.
Pubblicazione: (2023) -
Prototype-Based Knowledge Guidance for Fine-Grained Structured Radiology Reporting
di: Pellegrini, Chantal, et al.
Pubblicazione: (2026) -
Specialized Foundation Models for Intelligent Operating Rooms
di: Özsoy, Ege, et al.
Pubblicazione: (2025) -
MM-OR: A Large Multimodal Operating Room Dataset for Semantic Understanding of High-Intensity Surgical Environments
di: Özsoy, Ege, et al.
Pubblicazione: (2025) -
Location-Free Scene Graph Generation
di: Özsoy, Ege, et al.
Pubblicazione: (2023)