SurgiSAM2: Fine-tuning a foundational model for surgical video anatomy segmentation and detection
Fuente:
arXiv
Saved in:
| Main Authors: | Kamtam, Devanish N., Shrager, Joseph B., Malla, Satya Deepya, Wang, Xiaohan, Lin, Nicole, Cardona, Juan J., Yeung-Levy, Serena, Hu, Clarence |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Deep learning approaches to surgical video segmentation and object detection: A Scoping Review
by: Kamtam, Devanish N., et al.
Published: (2025)
by: Kamtam, Devanish N., et al.
Published: (2025)
Fine-tuning MLLMs Without Forgetting Is Easier Than You Think
by: Li, He, et al.
Published: (2026)
by: Li, He, et al.
Published: (2026)
Just Shift It: Test-Time Prototype Shifting for Zero-Shot Generalization with Vision-Language Models
by: Sui, Elaine, et al.
Published: (2024)
by: Sui, Elaine, et al.
Published: (2024)
Zero-shot Action Localization via the Confidence of Large Vision-Language Models
by: Aklilu, Josiah, et al.
Published: (2024)
by: Aklilu, Josiah, et al.
Published: (2024)
Feather the Throttle: Revisiting Visual Token Pruning for Vision-Language Model Acceleration
by: Endo, Mark, et al.
Published: (2024)
by: Endo, Mark, et al.
Published: (2024)
Fine-tuning vision foundation model for crack segmentation in civil infrastructures
by: Ge, Kang, et al.
Published: (2023)
by: Ge, Kang, et al.
Published: (2023)
SurgiColl
Published: (2026)
Published: (2026)
VideoAgent: Long-form Video Understanding with Large Language Model as Agent
by: Wang, Xiaohan, et al.
Published: (2024)
by: Wang, Xiaohan, et al.
Published: (2024)
Colony Grounded SAM2: Zero-shot detection and segmentation of bacterial colonies using foundation models
by: Korporaal, Daan, et al.
Published: (2026)
by: Korporaal, Daan, et al.
Published: (2026)
Disentangling spatio-temporal knowledge for weakly supervised object detection and segmentation in surgical video
by: Liao, Guiqiu, et al.
Published: (2024)
by: Liao, Guiqiu, et al.
Published: (2024)
ELIZA Reinterpreted: The world's first chatbot was not intended as a chatbot at all
by: Shrager, Jeff
Published: (2024)
by: Shrager, Jeff
Published: (2024)
Executable Archaeology: Reanimating the Logic Theorist from its IPL-V Source
by: Shrager, Jeff
Published: (2026)
by: Shrager, Jeff
Published: (2026)
Downscaling Intelligence: Exploring Perception and Reasoning Bottlenecks in Small Multimodal Models
by: Endo, Mark, et al.
Published: (2025)
by: Endo, Mark, et al.
Published: (2025)
Foundation Models Secretly Understand Neural Network Weights: Enhancing Hypernetwork Architectures with Foundation Models
by: Gu, Jeffrey, et al.
Published: (2025)
by: Gu, Jeffrey, et al.
Published: (2025)
ClickSAM: Fine-tuning Segment Anything Model using click prompts for ultrasound image segmentation
by: Guo, Aimee, et al.
Published: (2024)
by: Guo, Aimee, et al.
Published: (2024)
Video-STaR: Self-Training Enables Video Instruction Tuning with Any Supervision
by: Zohar, Orr, et al.
Published: (2024)
by: Zohar, Orr, et al.
Published: (2024)
Transductive Visual Programming: Evolving Tool Libraries from Experience for Spatial Reasoning
by: Wu, Shengguang, et al.
Published: (2025)
by: Wu, Shengguang, et al.
Published: (2025)
SurgiTrack: Fine-Grained Multi-Class Multi-Tool Tracking in Surgical Videos
by: Nwoye, Chinedu Innocent, et al.
Published: (2024)
by: Nwoye, Chinedu Innocent, et al.
Published: (2024)
Rule-based outlier detection of AI-generated anatomy segmentations
by: Krishnaswamy, Deepa, et al.
Published: (2024)
by: Krishnaswamy, Deepa, et al.
Published: (2024)
Fine-tuning foundational models to code diagnoses from veterinary health records
by: Boguslav, Mayla R., et al.
Published: (2024)
by: Boguslav, Mayla R., et al.
Published: (2024)
Adaptive transfer learning for surgical tool presence detection in laparoscopic videos through gradual freezing fine-tuning
by: Davila, Ana, et al.
Published: (2025)
by: Davila, Ana, et al.
Published: (2025)
GeoSAM: Fine-tuning SAM with Multi-Modal Prompts for Mobility Infrastructure Segmentation
by: Sultan, Rafi Ibn, et al.
Published: (2023)
by: Sultan, Rafi Ibn, et al.
Published: (2023)
V-GRPO: Online Reinforcement Learning for Denoising Generative Models Is Easier than You Think
by: Tang, Bingda, et al.
Published: (2026)
by: Tang, Bingda, et al.
Published: (2026)
Closing the Modality Gap for Mixed Modality Search
by: Li, Binxu, et al.
Published: (2025)
by: Li, Binxu, et al.
Published: (2025)
Temporal Preference Optimization for Long-Form Video Understanding
by: Li, Rui, et al.
Published: (2025)
by: Li, Rui, et al.
Published: (2025)
Continuous Perception Matters: Diagnosing Temporal Integration Failures in Multimodal Models
by: Wang, Zeyu, et al.
Published: (2024)
by: Wang, Zeyu, et al.
Published: (2024)
Multi-Human Mesh Recovery with Transformers
by: Wang, Zeyu, et al.
Published: (2024)
by: Wang, Zeyu, et al.
Published: (2024)
Connect, Collapse, Corrupt: Learning Cross-Modal Tasks with Uni-Modal Data
by: Zhang, Yuhui, et al.
Published: (2024)
by: Zhang, Yuhui, et al.
Published: (2024)
A Small Math Model: Recasting Strategy Choice Theory in an LLM-Inspired Architecture
by: Rahman, Roussel, et al.
Published: (2025)
by: Rahman, Roussel, et al.
Published: (2025)
Novel adaptation of video segmentation to 3D MRI: efficient zero-shot knee segmentation with SAM2
by: Yu, Andrew Seohwan, et al.
Published: (2024)
by: Yu, Andrew Seohwan, et al.
Published: (2024)
PlantSAM: An object detection‐driven segmentation pipeline for herbarium specimens
by: Youcef Sklab, et al.
Published: (2025)
by: Youcef Sklab, et al.
Published: (2025)
SAM-based instance segmentation models for the automation of structural damage detection
by: Ye, Zehao, et al.
Published: (2024)
by: Ye, Zehao, et al.
Published: (2024)
Diffusion-HPC: Synthetic Data Generation for Human Mesh Recovery in Challenging Domains
by: Weng, Zhenzhen, et al.
Published: (2023)
by: Weng, Zhenzhen, et al.
Published: (2023)
Viewpoint Textual Inversion: Discovering Scene Representations and 3D View Control in 2D Diffusion Models
by: Burgess, James, et al.
Published: (2023)
by: Burgess, James, et al.
Published: (2023)
Describing Differences in Image Sets with Natural Language
by: Dunlap, Lisa, et al.
Published: (2023)
by: Dunlap, Lisa, et al.
Published: (2023)
Why are Visually-Grounded Language Models Bad at Image Classification?
by: Zhang, Yuhui, et al.
Published: (2024)
by: Zhang, Yuhui, et al.
Published: (2024)
Tool Verification for Test-Time Reinforcement Learning
by: Liao, Ruotong, et al.
Published: (2026)
by: Liao, Ruotong, et al.
Published: (2026)
Fine-tuning foundation models of materials interatomic potentials with frozen transfer learning
by: Radova, Mariia, et al.
Published: (2025)
by: Radova, Mariia, et al.
Published: (2025)
RevSAM2: Prompt SAM2 for Medical Image Segmentation via Reverse-Propagation without Fine-tuning
by: Bai, Yunhao, et al.
Published: (2024)
by: Bai, Yunhao, et al.
Published: (2024)
NegVQA: Can Vision Language Models Understand Negation?
by: Zhang, Yuhui, et al.
Published: (2025)
by: Zhang, Yuhui, et al.
Published: (2025)
Similar Items
-
Deep learning approaches to surgical video segmentation and object detection: A Scoping Review
by: Kamtam, Devanish N., et al.
Published: (2025) -
Fine-tuning MLLMs Without Forgetting Is Easier Than You Think
by: Li, He, et al.
Published: (2026) -
Just Shift It: Test-Time Prototype Shifting for Zero-Shot Generalization with Vision-Language Models
by: Sui, Elaine, et al.
Published: (2024) -
Zero-shot Action Localization via the Confidence of Large Vision-Language Models
by: Aklilu, Josiah, et al.
Published: (2024) -
Feather the Throttle: Revisiting Visual Token Pruning for Vision-Language Model Acceleration
by: Endo, Mark, et al.
Published: (2024)