A vision-language model and platform for temporally mapping surgery from video
Fuente:
arXiv
Guardado en:
| Autor principal: | Kiyasseh, Dani |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Openfly: A comprehensive platform for aerial vision-language navigation
por: Gao, Yunpeng, et al.
Publicado: (2025)
por: Gao, Yunpeng, et al.
Publicado: (2025)
FlySearch: Exploring how vision-language models explore
por: Pardyl, Adam, et al.
Publicado: (2025)
por: Pardyl, Adam, et al.
Publicado: (2025)
CarLLaVA: Vision language models for camera-only closed-loop driving
por: Renz, Katrin, et al.
Publicado: (2024)
por: Renz, Katrin, et al.
Publicado: (2024)
WHU-PCPR: A cross-platform heterogeneous point cloud dataset for place recognition in complex urban scenes
por: Zou, Xianghong, et al.
Publicado: (2026)
por: Zou, Xianghong, et al.
Publicado: (2026)
SLAM assisted 3D tracking system for laparoscopic surgery
por: Song, Jingwei, et al.
Publicado: (2024)
por: Song, Jingwei, et al.
Publicado: (2024)
Detecting spills using thermal imaging, pretrained deep learning models, and a robotic platform
por: Yeghiyan, Gregory, et al.
Publicado: (2025)
por: Yeghiyan, Gregory, et al.
Publicado: (2025)
NeuralLabeling: A versatile toolset for labeling vision datasets using Neural Radiance Fields
por: Erich, Floris, et al.
Publicado: (2023)
por: Erich, Floris, et al.
Publicado: (2023)
Object Depth and Size Estimation using Stereo-vision and Integration with SLAM
por: Hamad, Layth, et al.
Publicado: (2024)
por: Hamad, Layth, et al.
Publicado: (2024)
MoManipVLA: Transferring Vision-language-action Models for General Mobile Manipulation
por: Wu, Zhenyu, et al.
Publicado: (2025)
por: Wu, Zhenyu, et al.
Publicado: (2025)
Touch begins where vision ends: Generalizable policies for contact-rich manipulation
por: Zhao, Zifan, et al.
Publicado: (2025)
por: Zhao, Zifan, et al.
Publicado: (2025)
Multi-vision-based Picking Point Localisation of Target Fruit for Harvesting Robots
por: Beldek, C., et al.
Publicado: (2025)
por: Beldek, C., et al.
Publicado: (2025)
An indoor DSO-based ceiling-vision odometry system for indoor industrial environments
por: Bougouffa, Abdelhak, et al.
Publicado: (2024)
por: Bougouffa, Abdelhak, et al.
Publicado: (2024)
Quantitative evaluation of brain-inspired vision sensors in high-speed robotic perception
por: Wang, Taoyi, et al.
Publicado: (2025)
por: Wang, Taoyi, et al.
Publicado: (2025)
DSLO: Deep Sequence LiDAR Odometry Based on Inconsistent Spatio-temporal Propagation
por: Zhang, Huixin, et al.
Publicado: (2024)
por: Zhang, Huixin, et al.
Publicado: (2024)
AscDAMs: Advanced SLAM-based channel detection and mapping system
por: Wang, Tengfei, et al.
Publicado: (2024)
por: Wang, Tengfei, et al.
Publicado: (2024)
Self-localization on a 3D map by fusing global and local features from a monocular camera
por: Kikuchi, Satoshi, et al.
Publicado: (2025)
por: Kikuchi, Satoshi, et al.
Publicado: (2025)
SurgeMOD: Translating image-space tissue motions into vision-based surgical forces
por: Reyzabal, Mikel De Iturrate, et al.
Publicado: (2024)
por: Reyzabal, Mikel De Iturrate, et al.
Publicado: (2024)
Real-time Spatial-temporal Traversability Assessment via Feature-based Sparse Gaussian Process
por: Hou, Zhenyu, et al.
Publicado: (2025)
por: Hou, Zhenyu, et al.
Publicado: (2025)
Thermal Chameleon: Task-Adaptive Tone-mapping for Radiometric Thermal-Infrared images
por: Lee, Dong-Guw, et al.
Publicado: (2024)
por: Lee, Dong-Guw, et al.
Publicado: (2024)
COMPASS: COmpact Multi-channel Prior-map And Scene Signature for Floor-Plan-Based Visual Localization
por: Shaheer, Muhammad, et al.
Publicado: (2026)
por: Shaheer, Muhammad, et al.
Publicado: (2026)
General surgery vision transformer: A video pre-trained foundation model for general surgery
por: Schmidgall, Samuel, et al.
Publicado: (2024)
por: Schmidgall, Samuel, et al.
Publicado: (2024)
Monocular pose estimation of articulated open surgery tools -- in the wild
por: Spektor, Robert, et al.
Publicado: (2024)
por: Spektor, Robert, et al.
Publicado: (2024)
Multi-step manipulation task and motion planning guided by video demonstration
por: Zorina, Kateryna, et al.
Publicado: (2025)
por: Zorina, Kateryna, et al.
Publicado: (2025)
GP-VLS: A general-purpose vision language model for surgery
por: Schmidgall, Samuel, et al.
Publicado: (2024)
por: Schmidgall, Samuel, et al.
Publicado: (2024)
Free-form language-based robotic reasoning and grasping
por: Jiao, Runyu, et al.
Publicado: (2025)
por: Jiao, Runyu, et al.
Publicado: (2025)
Exploring Conditions for Diffusion models in Robotic Control
por: Shin, Heeseong, et al.
Publicado: (2025)
por: Shin, Heeseong, et al.
Publicado: (2025)
OmniNOCS: A unified NOCS dataset and model for 3D lifting of 2D objects
por: Krishnan, Akshay, et al.
Publicado: (2024)
por: Krishnan, Akshay, et al.
Publicado: (2024)
The Better You Learn, The Smarter You Prune: Towards Efficient Vision-language-action Models via Differentiable Token Pruning
por: Jiang, Titong, et al.
Publicado: (2025)
por: Jiang, Titong, et al.
Publicado: (2025)
IFFNeRF: Initialisation Free and Fast 6DoF pose estimation from a single image and a NeRF model
por: Bortolon, Matteo, et al.
Publicado: (2024)
por: Bortolon, Matteo, et al.
Publicado: (2024)
SAMPO:Scale-wise Autoregression with Motion PrOmpt for generative world models
por: Wang, Sen, et al.
Publicado: (2025)
por: Wang, Sen, et al.
Publicado: (2025)
Towards foundational LiDAR world models with efficient latent flow matching
por: Liu, Tianran, et al.
Publicado: (2025)
por: Liu, Tianran, et al.
Publicado: (2025)
4-Dimensional deformation part model for pose estimation using Kalman filter constraints
por: Martinez-Berti, Enrique, et al.
Publicado: (2024)
por: Martinez-Berti, Enrique, et al.
Publicado: (2024)
Event-based vision for egomotion estimation using precise event timing
por: Greatorex, Hugh, et al.
Publicado: (2025)
por: Greatorex, Hugh, et al.
Publicado: (2025)
Analyzing the impact of semantic LoD3 building models on image-based vehicle localization
por: Bieringer, Antonia, et al.
Publicado: (2024)
por: Bieringer, Antonia, et al.
Publicado: (2024)
Risk-aware Trajectory Prediction by Incorporating Spatio-temporal Traffic Interaction Analysis
por: Thuremella, Divya, et al.
Publicado: (2024)
por: Thuremella, Divya, et al.
Publicado: (2024)
An analysis of vision-language models for fabric retrieval
por: Giuliari, Francesco, et al.
Publicado: (2025)
por: Giuliari, Francesco, et al.
Publicado: (2025)
Robot Learning from Human Videos: A Survey
por: Ma, Junyi, et al.
Publicado: (2026)
por: Ma, Junyi, et al.
Publicado: (2026)
A Generalization of CLAP from 3D Localization to Image Processing, A Connection With RANSAC & Hough Transforms
por: Hou, Ruochen, et al.
Publicado: (2025)
por: Hou, Ruochen, et al.
Publicado: (2025)
DEF-oriCORN: efficient 3D scene understanding for robust language-directed manipulation without demonstrations
por: Son, Dongwon, et al.
Publicado: (2024)
por: Son, Dongwon, et al.
Publicado: (2024)
Video Annotator: A framework for efficiently building video classifiers using vision-language models and active learning
por: Ziai, Amir, et al.
Publicado: (2024)
por: Ziai, Amir, et al.
Publicado: (2024)
Ejemplares similares
-
Openfly: A comprehensive platform for aerial vision-language navigation
por: Gao, Yunpeng, et al.
Publicado: (2025) -
FlySearch: Exploring how vision-language models explore
por: Pardyl, Adam, et al.
Publicado: (2025) -
CarLLaVA: Vision language models for camera-only closed-loop driving
por: Renz, Katrin, et al.
Publicado: (2024) -
WHU-PCPR: A cross-platform heterogeneous point cloud dataset for place recognition in complex urban scenes
por: Zou, Xianghong, et al.
Publicado: (2026) -
SLAM assisted 3D tracking system for laparoscopic surgery
por: Song, Jingwei, et al.
Publicado: (2024)