PitVQA: Image-grounded Text Embedding LLM for Visual Question Answering in Pituitary Surgery
Fuente:
arXiv
Saved in:
| Main Authors: | He, Runlong, Xu, Mengya, Das, Adrito, Khan, Danyal Z., Bano, Sophia, Marcus, Hani J., Stoyanov, Danail, Clarkson, Matthew J., Islam, Mobarakol |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
PitVQA++: Vector Matrix-Low-Rank Adaptation for Open-Ended Visual Question Answering in Pituitary Surgery
by: He, Runlong, et al.
Published: (2025)
by: He, Runlong, et al.
Published: (2025)
PitRSDNet: Predicting Intra-operative Remaining Surgery Duration in Endoscopic Pituitary Surgery
by: Wijekoon, Anjana, et al.
Published: (2024)
by: Wijekoon, Anjana, et al.
Published: (2024)
Automated Surgical Skill Assessment in Endoscopic Pituitary Surgery using Real-time Instrument Tracking on a High-fidelity Bench-top Phantom
by: Das, Adrito, et al.
Published: (2024)
by: Das, Adrito, et al.
Published: (2024)
SurgAnt-ViVQA: Learning to Anticipate Surgical Events through GRU-Driven Temporal Cross-Attention
by: Dhake, Shreyas C., et al.
Published: (2025)
by: Dhake, Shreyas C., et al.
Published: (2025)
PitVis-2023 Challenge: Workflow Recognition in videos of Endoscopic Pituitary Surgery
by: Das, Adrito, et al.
Published: (2024)
by: Das, Adrito, et al.
Published: (2024)
Surgical-DeSAM: Decoupling SAM for Instrument Segmentation in Robotic Surgery
by: Sheng, Yuyang, et al.
Published: (2024)
by: Sheng, Yuyang, et al.
Published: (2024)
Surgical AI Copilot: Energy-Based Fourier Gradient Low-Rank Adaptation for Surgical LLM Agent Reasoning and Planning
by: Huang, Jiayuan, et al.
Published: (2025)
by: Huang, Jiayuan, et al.
Published: (2025)
SurgViVQA: Temporally-Grounded Video Question Answering for Surgical Scene Understanding
by: Drago, Mauro Orazio, et al.
Published: (2025)
by: Drago, Mauro Orazio, et al.
Published: (2025)
TemporalDoRA: Temporal PEFT for Robust Surgical Video Question Answering
by: Carlini, Luca, et al.
Published: (2026)
by: Carlini, Luca, et al.
Published: (2026)
When to Trust the Answer: Question-Aligned Semantic Nearest Neighbor Entropy for Safer Surgical VQA
by: Carlini, Luca, et al.
Published: (2025)
by: Carlini, Luca, et al.
Published: (2025)
Privacy-Preserving Synthetic Continual Semantic Segmentation for Robotic Surgery
by: Xu, Mengya, et al.
Published: (2024)
by: Xu, Mengya, et al.
Published: (2024)
DARES: Depth Anything in Robotic Endoscopic Surgery with Self-supervised Vector-LoRA of the Foundation Model
by: Zeinoddin, Mona Sheikh, et al.
Published: (2024)
by: Zeinoddin, Mona Sheikh, et al.
Published: (2024)
Multi-Modal Monocular Endoscopic Depth and Pose Estimation with Edge-Guided Self-Supervision
by: Ju, Xinwei, et al.
Published: (2026)
by: Ju, Xinwei, et al.
Published: (2026)
Artificial intelligence in histopathological image analysis of central nervous system tumours: A systematic review
by: Melanie P. Jensen, et al.
Published: (2024)
by: Melanie P. Jensen, et al.
Published: (2024)
Surgical-VQLA++: Adversarial Contrastive Learning for Calibrated Robust Visual Question-Localized Answering in Robotic Surgery
by: Bai, Long, et al.
Published: (2024)
by: Bai, Long, et al.
Published: (2024)
Gaussian Pancakes: Geometrically-Regularized 3D Gaussian Splatting for Realistic Endoscopic Reconstruction
by: Bonilla, Sierra, et al.
Published: (2024)
by: Bonilla, Sierra, et al.
Published: (2024)
SurgicalGS: Dynamic 3D Gaussian Splatting for Accurate Robotic-Assisted Surgical Scene Reconstruction
by: Chen, Jialei, et al.
Published: (2024)
by: Chen, Jialei, et al.
Published: (2024)
HyKey: Hyperspectral Keypoint Detection and Matching in Minimally Invasive Surgery
by: Saikia, Alexander, et al.
Published: (2026)
by: Saikia, Alexander, et al.
Published: (2026)
Illumination Histogram Consistency Metric for Quantitative Assessment of Video Sequences
by: Chen, Long, et al.
Published: (2024)
by: Chen, Long, et al.
Published: (2024)
Robotic Arm Platform for Multi-View Image Acquisition and 3D Reconstruction in Minimally Invasive Surgery
by: Saikia, Alexander, et al.
Published: (2024)
by: Saikia, Alexander, et al.
Published: (2024)
LLM-Assisted Multi-Teacher Continual Learning for Visual Question Answering in Robotic Surgery
by: Du, Yuyang, et al.
Published: (2024)
by: Du, Yuyang, et al.
Published: (2024)
RRT-GPMP2: A Motion Planner for Mobile Robots in Complex Maze Environments
by: Meng, Jiawei, et al.
Published: (2024)
by: Meng, Jiawei, et al.
Published: (2024)
Mismatched: Evaluating the Limits of Image Matching Approaches and Benchmarks
by: Bonilla, Sierra, et al.
Published: (2024)
by: Bonilla, Sierra, et al.
Published: (2024)
SAM 2 in Robotic Surgery: An Empirical Evaluation for Robustness and Generalization in Surgical Video Segmentation
by: Yu, Jieming, et al.
Published: (2024)
by: Yu, Jieming, et al.
Published: (2024)
Depth Augmented and FE Free 3D/2D Liver Registration for Laparoscopic Liver AR
by: Zhang, Hanyuan, et al.
Published: (2026)
by: Zhang, Hanyuan, et al.
Published: (2026)
Surgical-LVLM: Learning to Adapt Large Vision-Language Model for Grounded Visual Question Answering in Robotic Surgery
by: Wang, Guankun, et al.
Published: (2024)
by: Wang, Guankun, et al.
Published: (2024)
HUP-3D: A 3D multi-view synthetic dataset for assisted-egocentric hand-ultrasound pose estimation
by: Birlo, Manuel, et al.
Published: (2024)
by: Birlo, Manuel, et al.
Published: (2024)
BERT-VQA: Visual Question Answering on Plots
by: Vu, Tai, et al.
Published: (2025)
by: Vu, Tai, et al.
Published: (2025)
Mapping Pituitary Neuroendocrine Tumors: An Annotated MRI Dataset Profiling Tumor and Carotid Characteristics
by: Anand S. Pandit, et al.
Published: (2025)
by: Anand S. Pandit, et al.
Published: (2025)
Surgical-DINO: Adapter Learning of Foundation Models for Depth Estimation in Endoscopic Surgery
by: Cui, Beilei, et al.
Published: (2024)
by: Cui, Beilei, et al.
Published: (2024)
Visual Robustness Benchmark for Visual Question Answering (VQA)
by: Ishmam, Md Farhan, et al.
Published: (2024)
by: Ishmam, Md Farhan, et al.
Published: (2024)
Dynamic Obstacle Avoidance of Unmanned Surface Vehicles in Maritime Environments Using Gaussian Processes Based Motion Planning
by: Meng, Jiawei, et al.
Published: (2024)
by: Meng, Jiawei, et al.
Published: (2024)
SHADeS: Self-supervised Monocular Depth Estimation Through Non-Lambertian Image Decomposition
by: Daher, Rema, et al.
Published: (2025)
by: Daher, Rema, et al.
Published: (2025)
RAVEN: Query-Guided Representation Alignment for Question Answering over Audio, Video, Embedded Sensors, and Natural Language
by: Biswas, Subrata, et al.
Published: (2025)
by: Biswas, Subrata, et al.
Published: (2025)
CommVQA: Situating Visual Question Answering in Communicative Contexts
by: Naik, Nandita Shankar, et al.
Published: (2024)
by: Naik, Nandita Shankar, et al.
Published: (2024)
VQA$^2$: Visual Question Answering for Video Quality Assessment
by: Jia, Ziheng, et al.
Published: (2024)
by: Jia, Ziheng, et al.
Published: (2024)
Endo-FASt3r: Endoscopic Foundation model Adaptation for Structure from motion
by: Zeinoddin, Mona Sheikh, et al.
Published: (2025)
by: Zeinoddin, Mona Sheikh, et al.
Published: (2025)
Personalizing Federated Instrument Segmentation with Visual Trait Priors in Robotic Surgery
by: Xu, Jialang, et al.
Published: (2024)
by: Xu, Jialang, et al.
Published: (2024)
StackOverflowVQA: Stack Overflow Visual Question Answering Dataset
by: Mirzaei, Motahhare, et al.
Published: (2024)
by: Mirzaei, Motahhare, et al.
Published: (2024)
FlowVQA: Mapping Multimodal Logic in Visual Question Answering with Flowcharts
by: Singh, Shubhankar, et al.
Published: (2024)
by: Singh, Shubhankar, et al.
Published: (2024)
Similar Items
-
PitVQA++: Vector Matrix-Low-Rank Adaptation for Open-Ended Visual Question Answering in Pituitary Surgery
by: He, Runlong, et al.
Published: (2025) -
PitRSDNet: Predicting Intra-operative Remaining Surgery Duration in Endoscopic Pituitary Surgery
by: Wijekoon, Anjana, et al.
Published: (2024) -
Automated Surgical Skill Assessment in Endoscopic Pituitary Surgery using Real-time Instrument Tracking on a High-fidelity Bench-top Phantom
by: Das, Adrito, et al.
Published: (2024) -
SurgAnt-ViVQA: Learning to Anticipate Surgical Events through GRU-Driven Temporal Cross-Attention
by: Dhake, Shreyas C., et al.
Published: (2025) -
PitVis-2023 Challenge: Workflow Recognition in videos of Endoscopic Pituitary Surgery
by: Das, Adrito, et al.
Published: (2024)