OphNet: A Large-Scale Video Benchmark for Ophthalmic Surgical Workflow Understanding
Fuente:
arXiv
Saved in:
| Main Authors: | Hu, Ming, Xia, Peng, Wang, Lin, Yan, Siyuan, Tang, Feilong, Xu, Zhongxing, Luo, Yimin, Song, Kaimin, Leitner, Jurgen, Cheng, Xuelian, Cheng, Jun, Liu, Chi, Zhou, Kaijing, Ge, Zongyuan |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
OphCLIP: Hierarchical Retrieval-Augmented Learning for Ophthalmic Surgical Video-Language Pretraining
by: Hu, Ming, et al.
Published: (2024)
by: Hu, Ming, et al.
Published: (2024)
Ophora: A Large-Scale Data-Driven Text-Guided Ophthalmic Surgical Video Generation Model
by: Li, Wei, et al.
Published: (2025)
by: Li, Wei, et al.
Published: (2025)
OphEdit: Training-Free Text-Guided Editing of Ophthalmic Surgical Videos
by: Jangir, Ritul, et al.
Published: (2026)
by: Jangir, Ritul, et al.
Published: (2026)
Towards Dynamic 3D Reconstruction of Hand-Instrument Interaction in Ophthalmic Surgery
by: Hu, Ming, et al.
Published: (2025)
by: Hu, Ming, et al.
Published: (2025)
OphIn-500K: Curating Web-Scale Visual Instructions for Scaling Ophthalmic Multimodal Large Language Models
by: Dong, Xuanzhao, et al.
Published: (2026)
by: Dong, Xuanzhao, et al.
Published: (2026)
ColonAdapter: Geometry Estimation Through Foundation Model Adaptation for Colonoscopy
by: Jiang, Zhiyi, et al.
Published: (2025)
by: Jiang, Zhiyi, et al.
Published: (2025)
Seeing Far and Clearly: Mitigating Hallucinations in MLLMs with Attention Causal Decoding
by: Tang, Feilong, et al.
Published: (2025)
by: Tang, Feilong, et al.
Published: (2025)
Hunting Attributes: Context Prototype-Aware Learning for Weakly Supervised Semantic Segmentation
by: Tang, Feilong, et al.
Published: (2024)
by: Tang, Feilong, et al.
Published: (2024)
Robust Multimodal Learning for Ophthalmic Disease Grading via Disentangled Representation
by: Wang, Xinkun, et al.
Published: (2025)
by: Wang, Xinkun, et al.
Published: (2025)
Neighbor Does Matter: Density-Aware Contrastive Learning for Medical Semi-supervised Segmentation
by: Tang, Feilong, et al.
Published: (2024)
by: Tang, Feilong, et al.
Published: (2024)
MMRC: A Large-Scale Benchmark for Understanding Multimodal Large Language Model in Real-World Conversation
by: Xue, Haochen, et al.
Published: (2025)
by: Xue, Haochen, et al.
Published: (2025)
RoomPlanner: Explicit Layout Planner for Easier LLM-Driven 3D Room Generation
by: Sun, Wenzhuo, et al.
Published: (2025)
by: Sun, Wenzhuo, et al.
Published: (2025)
Incomplete Modality Disentangled Representation for Ophthalmic Disease Grading and Diagnosis
by: Liu, Chengzhi, et al.
Published: (2025)
by: Liu, Chengzhi, et al.
Published: (2025)
Toward Modality Gap: Vision Prototype Learning for Weakly-supervised Semantic Segmentation with CLIP
by: Xu, Zhongxing, et al.
Published: (2024)
by: Xu, Zhongxing, et al.
Published: (2024)
Diffusion Model Driven Test-Time Image Adaptation for Robust Skin Lesion Classification
by: Hu, Ming, et al.
Published: (2024)
by: Hu, Ming, et al.
Published: (2024)
The inverse stability of Artin-Schreier polynomials over finite fields
by: Cheng, Kaimin
Published: (2024)
by: Cheng, Kaimin
Published: (2024)
The $3$-sparsity of $X^n-1$ over finite fields
by: Cheng, Kaimin
Published: (2025)
by: Cheng, Kaimin
Published: (2025)
The 3‐sparsity of Xn−1 over finite fields of characteristic 2
by: Kaimin Cheng
Published: (2026)
by: Kaimin Cheng
Published: (2026)
The $3$-sparsity of $X^n-1$ over finite fields, II
by: Cheng, Kaimin
Published: (2025)
by: Cheng, Kaimin
Published: (2025)
A class of ternary codes with few weights
by: Cheng, Kaimin
Published: (2024)
by: Cheng, Kaimin
Published: (2024)
Complete Walsh spectra for a permutation-inverse family of Boolean functions
by: Cheng, Kaimin
Published: (2026)
by: Cheng, Kaimin
Published: (2026)
Enhancing Vision-Language Models Generalization via Diversity-Driven Novel Feature Synthesis
by: Yan, Siyuan, et al.
Published: (2024)
by: Yan, Siyuan, et al.
Published: (2024)
Video-Thinker: Sparking "Thinking with Videos" via Reinforcement Learning
by: Wang, Shijian, et al.
Published: (2025)
by: Wang, Shijian, et al.
Published: (2025)
ScalingNoise: Scaling Inference-Time Search for Generating Infinite Videos
by: Yang, Haolin, et al.
Published: (2025)
by: Yang, Haolin, et al.
Published: (2025)
Surgical Ophthalmic Oncology
Published: (2020)
Published: (2020)
DermAgent: A Self-Reflective Agentic System for Dermatological Image Analysis with Multi-Tool Reasoning and Traceable Decision-Making
by: Liu, Yize, et al.
Published: (2026)
by: Liu, Yize, et al.
Published: (2026)
ReactDiff: Fundamental Multiple Appropriate Facial Reaction Diffusion Model
by: Cheng, Luo, et al.
Published: (2025)
by: Cheng, Luo, et al.
Published: (2025)
EH-Benchmark Ophthalmic Hallucination Benchmark and Agent-Driven Top-Down Traceable Reasoning Workflow
by: Pan, Xiaoyu, et al.
Published: (2025)
by: Pan, Xiaoyu, et al.
Published: (2025)
StreamAgent: Towards Anticipatory Agents for Streaming Video Understanding
by: Yang, Haolin, et al.
Published: (2025)
by: Yang, Haolin, et al.
Published: (2025)
MoRE: Unlocking Scalability in Reinforcement Learning for Quadruped Vision-Language-Action Models
by: Zhao, Han, et al.
Published: (2025)
by: Zhao, Han, et al.
Published: (2025)
Multiplicative character sums over two classes of subsets of quadratic extensions of finite fields
by: Cheng, Kaimin, et al.
Published: (2025)
by: Cheng, Kaimin, et al.
Published: (2025)
Weight distribution of a class of $p$-ary codes
by: Cheng, Kaimin, et al.
Published: (2025)
by: Cheng, Kaimin, et al.
Published: (2025)
On binomial Weil sums and an application
by: Cheng, Kaimin, et al.
Published: (2024)
by: Cheng, Kaimin, et al.
Published: (2024)
On the $2$-adic valuation of $σ_k(n)$
by: Cheng, Kaimin, et al.
Published: (2026)
by: Cheng, Kaimin, et al.
Published: (2026)
SurgicalPart-SAM: Part-to-Whole Collaborative Prompting for Surgical Instrument Segmentation
by: Yue, Wenxi, et al.
Published: (2023)
by: Yue, Wenxi, et al.
Published: (2023)
HGCLIP: Exploring Vision-Language Models with Graph Representations for Hierarchical Understanding
by: Xia, Peng, et al.
Published: (2023)
by: Xia, Peng, et al.
Published: (2023)
Infinitely Many Sign‐Changing Solutions for a Schrödinger Equation With Competing Potentials
by: Ke Wu, et al.
Published: (2025)
by: Ke Wu, et al.
Published: (2025)
Generalizing to Unseen Domains in Diabetic Retinopathy with Disentangled Representations
by: Xia, Peng, et al.
Published: (2024)
by: Xia, Peng, et al.
Published: (2024)
SANGRIA: Surgical Video Scene Graph Optimization for Surgical Workflow Prediction
by: Köksal, Çağhan, et al.
Published: (2024)
by: Köksal, Çağhan, et al.
Published: (2024)
MONICA: Benchmarking on Long-tailed Medical Image Classification
by: Ju, Lie, et al.
Published: (2024)
by: Ju, Lie, et al.
Published: (2024)
Similar Items
-
OphCLIP: Hierarchical Retrieval-Augmented Learning for Ophthalmic Surgical Video-Language Pretraining
by: Hu, Ming, et al.
Published: (2024) -
Ophora: A Large-Scale Data-Driven Text-Guided Ophthalmic Surgical Video Generation Model
by: Li, Wei, et al.
Published: (2025) -
OphEdit: Training-Free Text-Guided Editing of Ophthalmic Surgical Videos
by: Jangir, Ritul, et al.
Published: (2026) -
Towards Dynamic 3D Reconstruction of Hand-Instrument Interaction in Ophthalmic Surgery
by: Hu, Ming, et al.
Published: (2025) -
OphIn-500K: Curating Web-Scale Visual Instructions for Scaling Ophthalmic Multimodal Large Language Models
by: Dong, Xuanzhao, et al.
Published: (2026)