Zero-shot Prompt-based Video Encoder for Surgical Gesture Recognition
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Rao, Mingxing, Qin, Yinhong, Kolouri, Soheil, Wu, Jie Ying, Moyer, Daniel |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Generalization and Memorization in Rectified Flow
von: Rao, Mingxing, et al.
Veröffentlicht: (2026)
von: Rao, Mingxing, et al.
Veröffentlicht: (2026)
Score-based Membership Inference on Diffusion Models
von: Rao, Mingxing, et al.
Veröffentlicht: (2025)
von: Rao, Mingxing, et al.
Veröffentlicht: (2025)
Training Noise Token Pruning
von: Rao, Mingxing, et al.
Veröffentlicht: (2024)
von: Rao, Mingxing, et al.
Veröffentlicht: (2024)
Latent Diffusion Inversion Requires Understanding the Latent Space
von: Rao, Mingxing, et al.
Veröffentlicht: (2025)
von: Rao, Mingxing, et al.
Veröffentlicht: (2025)
Multi-Modal Gesture Recognition from Video and Surgical Tool Pose Information via Motion Invariants
von: Atoum, Jumanh, et al.
Veröffentlicht: (2025)
von: Atoum, Jumanh, et al.
Veröffentlicht: (2025)
One Category One Prompt: Dataset Distillation using Diffusion Models
von: Abbasi, Ali, et al.
Veröffentlicht: (2024)
von: Abbasi, Ali, et al.
Veröffentlicht: (2024)
HecVL: Hierarchical Video-Language Pretraining for Zero-shot Surgical Phase Recognition
von: Yuan, Kun, et al.
Veröffentlicht: (2024)
von: Yuan, Kun, et al.
Veröffentlicht: (2024)
Zero-Shot Underwater Gesture Recognition
von: Sarma, Sandipan, et al.
Veröffentlicht: (2024)
von: Sarma, Sandipan, et al.
Veröffentlicht: (2024)
BackdoorIDS: Zero-shot Backdoor Detection for Pretrained Vision Encoder
von: Huang, Siquan, et al.
Veröffentlicht: (2026)
von: Huang, Siquan, et al.
Veröffentlicht: (2026)
Vector-Quantized Soft Label Compression for Dataset Distillation
von: Abbasi, Ali, et al.
Veröffentlicht: (2026)
von: Abbasi, Ali, et al.
Veröffentlicht: (2026)
Prompt-to-Gesture: Measuring the Capabilities of Image-to-Video Deictic Gesture Generation
von: Ali, Hassan, et al.
Veröffentlicht: (2026)
von: Ali, Hassan, et al.
Veröffentlicht: (2026)
Fine-gained Zero-shot Video Sampling
von: Chen, Dengsheng, et al.
Veröffentlicht: (2024)
von: Chen, Dengsheng, et al.
Veröffentlicht: (2024)
SCALE: Semantic- and Confidence-Aware Conditional Variational Autoencoder for Zero-shot Skeleton-based Action Recognition
von: Oraki, Soroush, et al.
Veröffentlicht: (2026)
von: Oraki, Soroush, et al.
Veröffentlicht: (2026)
VideoMaker: Zero-shot Customized Video Generation with the Inherent Force of Video Diffusion Models
von: Wu, Tao, et al.
Veröffentlicht: (2024)
von: Wu, Tao, et al.
Veröffentlicht: (2024)
SG-Nav: Online 3D Scene Graph Prompting for LLM-based Zero-shot Object Navigation
von: Yin, Hang, et al.
Veröffentlicht: (2024)
von: Yin, Hang, et al.
Veröffentlicht: (2024)
Think Step by Step: Chain-of-Gesture Prompting for Error Detection in Robotic Surgical Videos
von: Shao, Zhimin, et al.
Veröffentlicht: (2024)
von: Shao, Zhimin, et al.
Veröffentlicht: (2024)
Min Generalized Sliced Gromov Wasserstein: A Scalable Path to Gromov Wasserstein
von: Shahbazi, Ashkan, et al.
Veröffentlicht: (2026)
von: Shahbazi, Ashkan, et al.
Veröffentlicht: (2026)
Meta-Prompting for Automating Zero-shot Visual Recognition with LLMs
von: Mirza, M. Jehanzeb, et al.
Veröffentlicht: (2024)
von: Mirza, M. Jehanzeb, et al.
Veröffentlicht: (2024)
Zero-shot HOI Detection with MLLM-based Detector-agnostic Interaction Recognition
von: Xuan, Shiyu, et al.
Veröffentlicht: (2026)
von: Xuan, Shiyu, et al.
Veröffentlicht: (2026)
ZeroPose: CAD-Prompted Zero-shot Object 6D Pose Estimation in Cluttered Scenes
von: Chen, Jianqiu, et al.
Veröffentlicht: (2023)
von: Chen, Jianqiu, et al.
Veröffentlicht: (2023)
Prompt-based Visual Alignment for Zero-shot Policy Transfer
von: Gao, Haihan, et al.
Veröffentlicht: (2024)
von: Gao, Haihan, et al.
Veröffentlicht: (2024)
DiffTED: One-shot Audio-driven TED Talk Video Generation with Diffusion-based Co-speech Gestures
von: Hogue, Steven, et al.
Veröffentlicht: (2024)
von: Hogue, Steven, et al.
Veröffentlicht: (2024)
ActionHub: A Large-scale Action Video Description Dataset for Zero-shot Action Recognition
von: Zhou, Jiaming, et al.
Veröffentlicht: (2024)
von: Zhou, Jiaming, et al.
Veröffentlicht: (2024)
MM-Gesture: Towards Precise Micro-Gesture Recognition through Multimodal Fusion
von: Gu, Jihao, et al.
Veröffentlicht: (2025)
von: Gu, Jihao, et al.
Veröffentlicht: (2025)
Boosting Gesture Recognition with an Automatic Gesture Annotation Framework
von: Shen, Junxiao, et al.
Veröffentlicht: (2024)
von: Shen, Junxiao, et al.
Veröffentlicht: (2024)
Efficient Transferable Optimal Transport via Min-Sliced Transport Plans
von: Liu, Xinran, et al.
Veröffentlicht: (2025)
von: Liu, Xinran, et al.
Veröffentlicht: (2025)
NOLA: Compressing LoRA using Linear Combination of Random Basis
von: Koohpayegani, Soroush Abbasi, et al.
Veröffentlicht: (2023)
von: Koohpayegani, Soroush Abbasi, et al.
Veröffentlicht: (2023)
Zero-shot Compositional Action Recognition with Neural Logic Constraints
von: Ye, Gefan, et al.
Veröffentlicht: (2025)
von: Ye, Gefan, et al.
Veröffentlicht: (2025)
EndoPBR: Material and Lighting Estimation for Photorealistic Surgical Simulations via Physically-based Rendering
von: Han, John J., et al.
Veröffentlicht: (2025)
von: Han, John J., et al.
Veröffentlicht: (2025)
Hierarchical Compositional Representations for Few-shot Action Recognition
von: Li, Changzhen, et al.
Veröffentlicht: (2022)
von: Li, Changzhen, et al.
Veröffentlicht: (2022)
Forearm Ultrasound based Gesture Recognition on Edge
von: Bimbraw, Keshav, et al.
Veröffentlicht: (2024)
von: Bimbraw, Keshav, et al.
Veröffentlicht: (2024)
Equivariant vs. Invariant Layers: A Comparison of Backbone and Pooling for Point Cloud Classification
von: Kothapalli, Abihith, et al.
Veröffentlicht: (2023)
von: Kothapalli, Abihith, et al.
Veröffentlicht: (2023)
Exploring Conditional Multi-Modal Prompts for Zero-shot HOI Detection
von: Lei, Ting, et al.
Veröffentlicht: (2024)
von: Lei, Ting, et al.
Veröffentlicht: (2024)
Fine-grained Abnormality Prompt Learning for Zero-shot Anomaly Detection
von: Zhu, Jiawen, et al.
Veröffentlicht: (2024)
von: Zhu, Jiawen, et al.
Veröffentlicht: (2024)
GenCLIP: Generalizing CLIP Prompts for Zero-shot Anomaly Detection
von: Kim, Donghyeong, et al.
Veröffentlicht: (2025)
von: Kim, Donghyeong, et al.
Veröffentlicht: (2025)
PDV: Prompt Directional Vectors for Zero-shot Composed Image Retrieval
von: Tursun, Osman, et al.
Veröffentlicht: (2025)
von: Tursun, Osman, et al.
Veröffentlicht: (2025)
How Far Are Surgeons from Surgical World Models? A Pilot Study on Zero-shot Surgical Video Generation with Expert Assessment
von: Chen, Zhen, et al.
Veröffentlicht: (2025)
von: Chen, Zhen, et al.
Veröffentlicht: (2025)
An Evaluation of Large Pre-Trained Models for Gesture Recognition using Synthetic Videos
von: Reddy, Arun, et al.
Veröffentlicht: (2024)
von: Reddy, Arun, et al.
Veröffentlicht: (2024)
SHANDS: A Multi-View Dataset and Benchmark for Surgical Hand-Gesture and Error Recognition Toward Medical Training
von: Ma, Le, et al.
Veröffentlicht: (2026)
von: Ma, Le, et al.
Veröffentlicht: (2026)
ZeroShape: Regression-based Zero-shot Shape Reconstruction
von: Huang, Zixuan, et al.
Veröffentlicht: (2023)
von: Huang, Zixuan, et al.
Veröffentlicht: (2023)
Ähnliche Einträge
-
Generalization and Memorization in Rectified Flow
von: Rao, Mingxing, et al.
Veröffentlicht: (2026) -
Score-based Membership Inference on Diffusion Models
von: Rao, Mingxing, et al.
Veröffentlicht: (2025) -
Training Noise Token Pruning
von: Rao, Mingxing, et al.
Veröffentlicht: (2024) -
Latent Diffusion Inversion Requires Understanding the Latent Space
von: Rao, Mingxing, et al.
Veröffentlicht: (2025) -
Multi-Modal Gesture Recognition from Video and Surgical Tool Pose Information via Motion Invariants
von: Atoum, Jumanh, et al.
Veröffentlicht: (2025)