VidLPRO: A $\underline{Vid}$eo-$\underline{L}$anguage $\underline{P}$re-training Framework for $\underline{Ro}$botic and Laparoscopic Surgery
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Honarmand, Mohammadmahdi, Jamal, Muhammad Abdullah, Mohareri, Omid |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
USDC: A Dataset of $\underline{U}$ser $\underline{S}$tance and $\underline{D}$ogmatism in Long $\underline{C}$onversations
von: Marreddy, Mounika, et al.
Veröffentlicht: (2024)
von: Marreddy, Mounika, et al.
Veröffentlicht: (2024)
FRAPPE: $\underline{\text{F}}$ast $\underline{\text{Ra}}$nk $\underline{\text{App}}$roximation with $\underline{\text{E}}$xplainable Features for Tensors
von: Shiao, William, et al.
Veröffentlicht: (2022)
von: Shiao, William, et al.
Veröffentlicht: (2022)
ALLMod: Exploring $\underline{\mathbf{A}}$rea-Efficiency of $\underline{\mathbf{L}}$UT-based $\underline{\mathbf{L}}$arge Number $\underline{\mathbf{Mod}}$ular Reduction via Hybrid Workloads
von: Liu, Fangxin, et al.
Veröffentlicht: (2025)
von: Liu, Fangxin, et al.
Veröffentlicht: (2025)
A Two-Stage Progressive Pre-training using Multi-Modal Contrastive Masked Autoencoders
von: Jamal, Muhammad Abdullah, et al.
Veröffentlicht: (2024)
von: Jamal, Muhammad Abdullah, et al.
Veröffentlicht: (2024)
Equivariant $H\underline{\mathbb{F}}_p$-modules are wild
von: Grevstad, Jacob Fjeld, et al.
Veröffentlicht: (2025)
von: Grevstad, Jacob Fjeld, et al.
Veröffentlicht: (2025)
The $RO(\mathcal{K})$-graded Coefficients of $H\underline{A}$
von: Keyes, Jesse
Veröffentlicht: (2025)
von: Keyes, Jesse
Veröffentlicht: (2025)
Rethinking RGB-D Fusion for Semantic Segmentation in Surgical Datasets
von: Jamal, Muhammad Abdullah, et al.
Veröffentlicht: (2024)
von: Jamal, Muhammad Abdullah, et al.
Veröffentlicht: (2024)
Mechanisms underlining Kelp (Saccharina japonica) adaptation to relative high seawater temperature.
von: Guo, Li, et al.
Veröffentlicht: (2025)
von: Guo, Li, et al.
Veröffentlicht: (2025)
On the structure of the $RO(G)$-graded homotopy of $H\underline{M}$ for cyclic $p$-groups
von: Sikora, Igor, et al.
Veröffentlicht: (2023)
von: Sikora, Igor, et al.
Veröffentlicht: (2023)
The $H \underline{\mathbb{F}}_2$-homology of $C_2$-equivariant Eilenberg-MacLane spaces
von: Petersen, Sarah
Veröffentlicht: (2022)
von: Petersen, Sarah
Veröffentlicht: (2022)
SurgLaVi: Large-Scale Hierarchical Dataset for Surgical Vision-Language Representation Learning
von: Perez, Alejandra, et al.
Veröffentlicht: (2025)
von: Perez, Alejandra, et al.
Veröffentlicht: (2025)
AdaEmbed: Semi-supervised Domain Adaptation in the Embedding Space
von: Mottaghi, Ali, et al.
Veröffentlicht: (2024)
von: Mottaghi, Ali, et al.
Veröffentlicht: (2024)
On the Role of Depth in Surgical Vision Foundation Models: An Empirical Study of RGB-D Pre-training
von: Han, John J., et al.
Veröffentlicht: (2026)
von: Han, John J., et al.
Veröffentlicht: (2026)
Multi-view Video-Pose Pretraining for Operating Room Surgical Activity Recognition
von: Hamoud, Idris, et al.
Veröffentlicht: (2025)
von: Hamoud, Idris, et al.
Veröffentlicht: (2025)
Anopheles mosquitoes in Morocco: implication for public health and underlined challenges for malaria re‐establishment prevention under current and future climate conditions
von: Outammassine Abdelkrim, et al.
Veröffentlicht: (2024)
von: Outammassine Abdelkrim, et al.
Veröffentlicht: (2024)
InstructVid2Vid: Controllable Video Editing with Natural Language Instructions
von: Qin, Bosheng, et al.
Veröffentlicht: (2023)
von: Qin, Bosheng, et al.
Veröffentlicht: (2023)
OmniVid: A Generative Framework for Universal Video Understanding
von: Wang, Junke, et al.
Veröffentlicht: (2024)
von: Wang, Junke, et al.
Veröffentlicht: (2024)
UniVid: Unifying Vision Tasks with Pre-trained Video Generation Models
von: Chen, Lan, et al.
Veröffentlicht: (2025)
von: Chen, Lan, et al.
Veröffentlicht: (2025)
HarmoVid: Relightful Video Portrait Harmonization
von: Choi, Jun Myeong, et al.
Veröffentlicht: (2026)
von: Choi, Jun Myeong, et al.
Veröffentlicht: (2026)
SUREON: A Benchmark and Vision-Language-Model for Surgical Reasoning
von: Perez, Alejandra, et al.
Veröffentlicht: (2026)
von: Perez, Alejandra, et al.
Veröffentlicht: (2026)
VidHal: Benchmarking Temporal Hallucinations in Vision LLMs
von: Choong, Wey Yeh, et al.
Veröffentlicht: (2024)
von: Choong, Wey Yeh, et al.
Veröffentlicht: (2024)
UniVid: The Open-Source Unified Video Model
von: Luo, Jiabin, et al.
Veröffentlicht: (2025)
von: Luo, Jiabin, et al.
Veröffentlicht: (2025)
AdaVid: Adaptive Video-Language Pretraining
von: Patel, Chaitanya, et al.
Veröffentlicht: (2025)
von: Patel, Chaitanya, et al.
Veröffentlicht: (2025)
IF-VidCap: Can Video Caption Models Follow Instructions?
von: Li, Shihao, et al.
Veröffentlicht: (2025)
von: Li, Shihao, et al.
Veröffentlicht: (2025)
VidText: Towards Comprehensive Evaluation for Video Text Understanding
von: Yang, Zhoufaran, et al.
Veröffentlicht: (2025)
von: Yang, Zhoufaran, et al.
Veröffentlicht: (2025)
TTA-Vid: Generalized Test-Time Adaptation for Video Reasoning
von: Jahagirdar, Soumya Shamarao, et al.
Veröffentlicht: (2026)
von: Jahagirdar, Soumya Shamarao, et al.
Veröffentlicht: (2026)
ExpVid: A Benchmark for Experiment Video Understanding & Reasoning
von: Xu, Yicheng, et al.
Veröffentlicht: (2025)
von: Xu, Yicheng, et al.
Veröffentlicht: (2025)
BachVid: Training-Free Video Generation with Consistent Background and Character
von: Yan, Han, et al.
Veröffentlicht: (2025)
von: Yan, Han, et al.
Veröffentlicht: (2025)
VidCLearn: A Continual Learning Approach for Text-to-Video Generation
von: Zanchetta, Luca, et al.
Veröffentlicht: (2025)
von: Zanchetta, Luca, et al.
Veröffentlicht: (2025)
VidLA: Video-Language Alignment at Scale
von: Rizve, Mamshad Nayeem, et al.
Veröffentlicht: (2024)
von: Rizve, Mamshad Nayeem, et al.
Veröffentlicht: (2024)
UniVidX: A Unified Multimodal Framework for Versatile Video Generation via Diffusion Priors
von: Chen, Houyuan, et al.
Veröffentlicht: (2026)
von: Chen, Houyuan, et al.
Veröffentlicht: (2026)
TextVidBench: A Benchmark for Long Video Scene Text Understanding
von: Zhong, Yangyang, et al.
Veröffentlicht: (2025)
von: Zhong, Yangyang, et al.
Veröffentlicht: (2025)
VidSplat: Gaussian Splatting Reconstruction with Geometry-Guided Video Diffusion Priors
von: Tang, Jimin, et al.
Veröffentlicht: (2026)
von: Tang, Jimin, et al.
Veröffentlicht: (2026)
PruneVid: Visual Token Pruning for Efficient Video Large Language Models
von: Huang, Xiaohu, et al.
Veröffentlicht: (2024)
von: Huang, Xiaohu, et al.
Veröffentlicht: (2024)
EvoVid: Temporal-Centric Self-Evolution for Video Large Language Models
von: Huang, Shiqi, et al.
Veröffentlicht: (2026)
von: Huang, Shiqi, et al.
Veröffentlicht: (2026)
VidEmo: Affective-Tree Reasoning for Emotion-Centric Video Foundation Models
von: Zhang, Zhicheng, et al.
Veröffentlicht: (2025)
von: Zhang, Zhicheng, et al.
Veröffentlicht: (2025)
VidEdit: Zero-Shot and Spatially Aware Text-Driven Video Editing
von: Couairon, Paul, et al.
Veröffentlicht: (2023)
von: Couairon, Paul, et al.
Veröffentlicht: (2023)
Vid-Morp: Video Moment Retrieval Pretraining from Unlabeled Videos in the Wild
von: Bao, Peijun, et al.
Veröffentlicht: (2024)
von: Bao, Peijun, et al.
Veröffentlicht: (2024)
VidComposition: Can MLLMs Analyze Compositions in Compiled Videos?
von: Tang, Yolo Y., et al.
Veröffentlicht: (2024)
von: Tang, Yolo Y., et al.
Veröffentlicht: (2024)
RelightVid: Temporal-Consistent Diffusion Model for Video Relighting
von: Fang, Ye, et al.
Veröffentlicht: (2025)
von: Fang, Ye, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
USDC: A Dataset of $\underline{U}$ser $\underline{S}$tance and $\underline{D}$ogmatism in Long $\underline{C}$onversations
von: Marreddy, Mounika, et al.
Veröffentlicht: (2024) -
FRAPPE: $\underline{\text{F}}$ast $\underline{\text{Ra}}$nk $\underline{\text{App}}$roximation with $\underline{\text{E}}$xplainable Features for Tensors
von: Shiao, William, et al.
Veröffentlicht: (2022) -
ALLMod: Exploring $\underline{\mathbf{A}}$rea-Efficiency of $\underline{\mathbf{L}}$UT-based $\underline{\mathbf{L}}$arge Number $\underline{\mathbf{Mod}}$ular Reduction via Hybrid Workloads
von: Liu, Fangxin, et al.
Veröffentlicht: (2025) -
A Two-Stage Progressive Pre-training using Multi-Modal Contrastive Masked Autoencoders
von: Jamal, Muhammad Abdullah, et al.
Veröffentlicht: (2024) -
Equivariant $H\underline{\mathbb{F}}_p$-modules are wild
von: Grevstad, Jacob Fjeld, et al.
Veröffentlicht: (2025)