Recognize Any Surgical Object: Unleashing the Power of Weakly-Supervised Data
Fuente:
arXiv
Saved in:
| Main Authors: | Li, Jiajie, Quaranto, Brian R, Xu, Chenhui, Mishra, Ishan, Qin, Ruiyang, Liu, Dancheng, Kim, Peter C W, Xiong, Jinjun |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
LLaVA-Surg: Towards Multimodal Surgical Assistant via Structured Surgical Video Learning
by: Li, Jiajie, et al.
Published: (2024)
by: Li, Jiajie, et al.
Published: (2024)
Selective Prior Synchronization via SYNC Loss
by: Mishra, Ishan, et al.
Published: (2026)
by: Mishra, Ishan, et al.
Published: (2026)
Chain-of-Adaptation: Surgical Vision-Language Adaptation with Reinforcement Learning
by: Li, Jiajie, et al.
Published: (2026)
by: Li, Jiajie, et al.
Published: (2026)
Sub-Sequential Physics-Informed Learning with State Space Model
by: Xu, Chenhui, et al.
Published: (2025)
by: Xu, Chenhui, et al.
Published: (2025)
PI-Whisper: Designing an Adaptive and Incremental Automatic Speech Recognition System for Edge Devices
by: Nassereldine, Amir, et al.
Published: (2024)
by: Nassereldine, Amir, et al.
Published: (2024)
Ensembler: Protect Collaborative Inference Privacy from Model Inversion Attack via Selective Ensemble
by: Liu, Dancheng, et al.
Published: (2024)
by: Liu, Dancheng, et al.
Published: (2024)
Towards Pretraining Robust ASR Foundation Model with Acoustic-Aware Data Augmentation
by: Liu, Dancheng, et al.
Published: (2025)
by: Liu, Dancheng, et al.
Published: (2025)
Driving Through Uncertainty: Risk-Averse Control with LLM Commonsense for Autonomous Driving under Perception Deficits
by: Hu, Yuting, et al.
Published: (2025)
by: Hu, Yuting, et al.
Published: (2025)
Towards Understanding Multi-Round Large Language Model Reasoning: Approximability, Learnability and Generalizability
by: Xu, Chenhui, et al.
Published: (2025)
by: Xu, Chenhui, et al.
Published: (2025)
FP64 is All You Need: Rethinking Failure Modes in Physics-Informed Neural Networks
by: Xu, Chenhui, et al.
Published: (2025)
by: Xu, Chenhui, et al.
Published: (2025)
FASA: a Flexible and Automatic Speech Aligner for Extracting High-quality Aligned Children Speech Data
by: Liu, Dancheng, et al.
Published: (2024)
by: Liu, Dancheng, et al.
Published: (2024)
From Sight to Insight: Unleashing Eye-Tracking in Weakly Supervised Video Salient Object Detection
by: Qin, Qi, et al.
Published: (2025)
by: Qin, Qi, et al.
Published: (2025)
Distinguish Any Fake Videos: Unleashing the Power of Large-scale Data and Motion Features
by: Ji, Lichuan, et al.
Published: (2024)
by: Ji, Lichuan, et al.
Published: (2024)
Large Language Models have Intrinsic Self-Correction Ability
by: Liu, Dancheng, et al.
Published: (2024)
by: Liu, Dancheng, et al.
Published: (2024)
Recognize Any Regions
by: Yang, Haosen, et al.
Published: (2023)
by: Yang, Haosen, et al.
Published: (2023)
PartGLEE: A Foundation Model for Recognizing and Parsing Any Objects
by: Li, Junyi, et al.
Published: (2024)
by: Li, Junyi, et al.
Published: (2024)
Empirical Guidelines for Deploying LLMs onto Resource-constrained Edge Devices
by: Qin, Ruiyang, et al.
Published: (2024)
by: Qin, Ruiyang, et al.
Published: (2024)
Chain-of-Look Spatial Reasoning for Dense Surgical Instrument Counting
by: Bhyri, Rishikesh, et al.
Published: (2026)
by: Bhyri, Rishikesh, et al.
Published: (2026)
Tiny-Align: Bridging Automatic Speech Recognition and Large Language Model on the Edge
by: Qin, Ruiyang, et al.
Published: (2024)
by: Qin, Ruiyang, et al.
Published: (2024)
MorphAny3D: Unleashing the Power of Structured Latent in 3D Morphing
by: Sun, Xiaokun, et al.
Published: (2026)
by: Sun, Xiaokun, et al.
Published: (2026)
Track Any Peppers: Weakly Supervised Sweet Pepper Tracking Using VLMs
by: Lim, Jia Syuen, et al.
Published: (2024)
by: Lim, Jia Syuen, et al.
Published: (2024)
Depth Anything: Unleashing the Power of Large-Scale Unlabeled Data
by: Yang, Lihe, et al.
Published: (2024)
by: Yang, Lihe, et al.
Published: (2024)
Unleashing Guidance Without Classifiers for Human-Object Interaction Animation
by: Wang, Ziyin, et al.
Published: (2026)
by: Wang, Ziyin, et al.
Published: (2026)
xMLP: Revolutionizing Private Inference with Exclusive Square Activation
by: Li, Jiajie, et al.
Published: (2024)
by: Li, Jiajie, et al.
Published: (2024)
Unleashing the Power of Chain-of-Prediction for Monocular 3D Object Detection
by: Zhang, Zhihao, et al.
Published: (2025)
by: Zhang, Zhihao, et al.
Published: (2025)
Unleashing the Representational Power of Fourier Shapes for Attacking Infrared Object Detection
by: Yong, Yixing, et al.
Published: (2026)
by: Yong, Yixing, et al.
Published: (2026)
MorphOPC: Advancing Mask Optimization with Multi-scale Hierarchical Morphological Learning
by: Hu, Yuting, et al.
Published: (2026)
by: Hu, Yuting, et al.
Published: (2026)
SPWOOD: Sparse Partial Weakly-Supervised Oriented Object Detection
by: Zhang, Wei, et al.
Published: (2026)
by: Zhang, Wei, et al.
Published: (2026)
Partial Weakly-Supervised Oriented Object Detection
by: Liu, Mingxin, et al.
Published: (2025)
by: Liu, Mingxin, et al.
Published: (2025)
Unleashing the Power of Self-Supervised Image Denoising: A Comprehensive Review
by: Zhang, Dan, et al.
Published: (2023)
by: Zhang, Dan, et al.
Published: (2023)
Weakly-Supervised Referring Video Object Segmentation through Text Supervision
by: Shi, Miaojing, et al.
Published: (2026)
by: Shi, Miaojing, et al.
Published: (2026)
A Holistically Point-guided Text Framework for Weakly-Supervised Camouflaged Object Detection
by: Mok, Tsui Qin, et al.
Published: (2025)
by: Mok, Tsui Qin, et al.
Published: (2025)
Weakly Supervised Object Detection for Automatic Tooth-marked Tongue Recognition
by: Zhang, Yongcun, et al.
Published: (2024)
by: Zhang, Yongcun, et al.
Published: (2024)
Automatic Screening for Children with Speech Disorder using Automatic Speech Recognition: Opportunities and Challenges
by: Liu, Dancheng, et al.
Published: (2024)
by: Liu, Dancheng, et al.
Published: (2024)
Unleashing the Power of Intermediate Domains for Mixed Domain Semi-Supervised Medical Image Segmentation
by: Ma, Qinghe, et al.
Published: (2025)
by: Ma, Qinghe, et al.
Published: (2025)
EasyHOI: Unleashing the Power of Large Models for Reconstructing Hand-Object Interactions in the Wild
by: Liu, Yumeng, et al.
Published: (2024)
by: Liu, Yumeng, et al.
Published: (2024)
Weakly Supervised Point Clouds Transformer for 3D Object Detection
by: Tang, Zuojin, et al.
Published: (2023)
by: Tang, Zuojin, et al.
Published: (2023)
Weakly Supervised YOLO Network for Surgical Instrument Localization in Endoscopic Videos
by: Wei, Rongfeng, et al.
Published: (2023)
by: Wei, Rongfeng, et al.
Published: (2023)
Automating Intervention Discovery from Scientific Literature: A Progressive Ontology Prompting and Dual-LLM Framework
by: Hu, Yuting, et al.
Published: (2024)
by: Hu, Yuting, et al.
Published: (2024)
Unleashing the Power of Image-Tabular Self-Supervised Learning via Breaking Cross-Tabular Barriers
by: Fu, Yibing, et al.
Published: (2025)
by: Fu, Yibing, et al.
Published: (2025)
Similar Items
-
LLaVA-Surg: Towards Multimodal Surgical Assistant via Structured Surgical Video Learning
by: Li, Jiajie, et al.
Published: (2024) -
Selective Prior Synchronization via SYNC Loss
by: Mishra, Ishan, et al.
Published: (2026) -
Chain-of-Adaptation: Surgical Vision-Language Adaptation with Reinforcement Learning
by: Li, Jiajie, et al.
Published: (2026) -
Sub-Sequential Physics-Informed Learning with State Space Model
by: Xu, Chenhui, et al.
Published: (2025) -
PI-Whisper: Designing an Adaptive and Incremental Automatic Speech Recognition System for Edge Devices
by: Nassereldine, Amir, et al.
Published: (2024)