Unifying Segment Anything in Microscopy with Vision-Language Knowledge
Fuente:
arXiv
Saved in:
| Main Authors: | Li, Manyu, He, Ruian, Zhang, Zixian, Ma, Chenxi, Tan, Weimin, Yan, Bo |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
MicroVQA++: High-Quality Microscopy Reasoning Dataset with Weakly Supervised Graphs for Multimodal Large Language Model
by: Li, Manyu, et al.
Published: (2025)
by: Li, Manyu, et al.
Published: (2025)
MicroWorld: Empowering Multimodal Large Language Models to Bridge the Microscopic Domain Gap with Multimodal Attribute Graph
by: Li, Manyu, et al.
Published: (2026)
by: Li, Manyu, et al.
Published: (2026)
A Benchmarking Study of Vision-based Robotic Grasping Algorithms
by: Rameshbabu, Bharath K, et al.
Published: (2025)
by: Rameshbabu, Bharath K, et al.
Published: (2025)
RealVVT: Towards Photorealistic Video Virtual Try-on via Spatio-Temporal Consistency
by: Li, Siqi, et al.
Published: (2025)
by: Li, Siqi, et al.
Published: (2025)
A Generative Framework for Self-Supervised Facial Representation Learning
by: He, Ruian, et al.
Published: (2023)
by: He, Ruian, et al.
Published: (2023)
Facial Surgery Preview Based on the Orthognathic Treatment Prediction
by: Han, Huijun, et al.
Published: (2024)
by: Han, Huijun, et al.
Published: (2024)
Deep Learning-Based Automated Workflow for Accurate Segmentation and Measurement of Abdominal Organs in CT Scans
by: Shastry, Praveen, et al.
Published: (2025)
by: Shastry, Praveen, et al.
Published: (2025)
WAVE-DETR Multi-Modal Visible and Acoustic Real-Life Drone Detector
by: Stefanescu, Razvan, et al.
Published: (2025)
by: Stefanescu, Razvan, et al.
Published: (2025)
Parking Space Ground Truth Test Automation by Artificial Intelligence Using Convolutional Neural Networks
by: Rohe, Tony, et al.
Published: (2025)
by: Rohe, Tony, et al.
Published: (2025)
Muscles in Time: Learning to Understand Human Motion by Simulating Muscle Activations
by: Schneider, David, et al.
Published: (2024)
by: Schneider, David, et al.
Published: (2024)
FacialFlowNet: Advancing Facial Optical Flow Estimation with a Diverse Dataset and a Decomposed Model
by: Lu, Jianzhi, et al.
Published: (2024)
by: Lu, Jianzhi, et al.
Published: (2024)
Context-Aware Iteration Policy Network for Efficient Optical Flow Estimation
by: Cheng, Ri, et al.
Published: (2023)
by: Cheng, Ri, et al.
Published: (2023)
MM-Skin: Enhancing Dermatology Vision-Language Model with an Image-Text Dataset Derived from Textbooks
by: Zeng, Wenqi, et al.
Published: (2025)
by: Zeng, Wenqi, et al.
Published: (2025)
FRIEREN: Federated Learning with Vision-Language Regularization for Segmentation
by: Shen, Ding-Ruei
Published: (2025)
by: Shen, Ding-Ruei
Published: (2025)
Multimodal Urban Areas of Interest Generation via Remote Sensing Imagery and Geographical Prior
by: Shi, Chuanji, et al.
Published: (2024)
by: Shi, Chuanji, et al.
Published: (2024)
PlaneSAM: Multimodal Plane Instance Segmentation Using the Segment Anything Model
by: Deng, Zhongchen, et al.
Published: (2024)
by: Deng, Zhongchen, et al.
Published: (2024)
Multimodal Integration Challenges in Emotionally Expressive Child Avatars for Training Applications
by: Salehi, Pegah, et al.
Published: (2025)
by: Salehi, Pegah, et al.
Published: (2025)
The Batch Artifact Scanning Protocol: A new method using computed tomography (CT) to rapidly create three-dimensional models of objects from large collections en masse
by: Yezzi-Woodley, Katrina, et al.
Published: (2022)
by: Yezzi-Woodley, Katrina, et al.
Published: (2022)
The Repeated-Stimulus Confound in Electroencephalography
by: Kilgallen, Jack A., et al.
Published: (2025)
by: Kilgallen, Jack A., et al.
Published: (2025)
Sampling Strategies for Mitigating Bias in Face Synthesis Methods
by: Maragkoudakis, Emmanouil, et al.
Published: (2024)
by: Maragkoudakis, Emmanouil, et al.
Published: (2024)
2DSig-Detect: a semi-supervised framework for anomaly detection on image data using 2D-signatures
by: Xie, Xinheng, et al.
Published: (2024)
by: Xie, Xinheng, et al.
Published: (2024)
Automated Radiology Report Generation: A Review of Recent Advances
by: Sloan, Phillip, et al.
Published: (2024)
by: Sloan, Phillip, et al.
Published: (2024)
Remote SAMsing: From Segment Anything to Segment Everything
by: de Carvalho, Osmar Luiz Ferreira, et al.
Published: (2026)
by: de Carvalho, Osmar Luiz Ferreira, et al.
Published: (2026)
Segment Anything for Dendrites from Electron Microscopy
by: Zhuo, Zewen, et al.
Published: (2024)
by: Zhuo, Zewen, et al.
Published: (2024)
Revisiting Sampson Approximations for Geometric Estimation Problems
by: Rydell, Felix, et al.
Published: (2024)
by: Rydell, Felix, et al.
Published: (2024)
Isolated Sign Language Recognition with Segmentation and Pose Estimation
by: Perkins, Daniel, et al.
Published: (2025)
by: Perkins, Daniel, et al.
Published: (2025)
FastGS: Training 3D Gaussian Splatting in 100 Seconds
by: Ren, Shiwei, et al.
Published: (2025)
by: Ren, Shiwei, et al.
Published: (2025)
SEGS-SLAM: Structure-enhanced 3D Gaussian Splatting SLAM with Appearance Embedding
by: Wen, Tianci, et al.
Published: (2025)
by: Wen, Tianci, et al.
Published: (2025)
Geometric Artifact Correction for Symmetric Multi-Linear Trajectory CT: Theory, Method, and Generalization
by: Wang, Zhisheng, et al.
Published: (2024)
by: Wang, Zhisheng, et al.
Published: (2024)
How Well Do Vision-Language Models Understand Sequential Driving Scenes? A Sensitivity Study
by: Brusnicki, Roberto, et al.
Published: (2026)
by: Brusnicki, Roberto, et al.
Published: (2026)
Improve Vision Language Model Chain-of-thought Reasoning
by: Zhang, Ruohong, et al.
Published: (2024)
by: Zhang, Ruohong, et al.
Published: (2024)
Persona-aware and Explainable Bikeability Assessment: A Vision-Language Model Approach
by: Dai, Yilong, et al.
Published: (2026)
by: Dai, Yilong, et al.
Published: (2026)
Explainable embeddings with Distance Explainer
by: Meijer, Christiaan, et al.
Published: (2025)
by: Meijer, Christiaan, et al.
Published: (2025)
Real Time Human Detection by Unmanned Aerial Vehicles
by: Guettala, Walid, et al.
Published: (2024)
by: Guettala, Walid, et al.
Published: (2024)
Interactive Image Selection and Training for Brain Tumor Segmentation Network
by: Cerqueira, Matheus A., et al.
Published: (2024)
by: Cerqueira, Matheus A., et al.
Published: (2024)
Recent Advances and Challenges in Deep Audio-Visual Correlation Learning
by: Vilaça, Luís, et al.
Published: (2022)
by: Vilaça, Luís, et al.
Published: (2022)
Segment Anything in Medical Images
by: Ma, Jun, et al.
Published: (2023)
by: Ma, Jun, et al.
Published: (2023)
Biomechanical Constraints Assimilation in Deep-Learning Image Registration: Application to sliding and locally rigid deformations
by: Kheil, Ziad, et al.
Published: (2025)
by: Kheil, Ziad, et al.
Published: (2025)
Tri-Plane Mamba: Efficiently Adapting Segment Anything Model for 3D Medical Images
by: Wang, Hualiang, et al.
Published: (2024)
by: Wang, Hualiang, et al.
Published: (2024)
Matching Anything by Segmenting Anything
by: Li, Siyuan, et al.
Published: (2024)
by: Li, Siyuan, et al.
Published: (2024)
Similar Items
-
MicroVQA++: High-Quality Microscopy Reasoning Dataset with Weakly Supervised Graphs for Multimodal Large Language Model
by: Li, Manyu, et al.
Published: (2025) -
MicroWorld: Empowering Multimodal Large Language Models to Bridge the Microscopic Domain Gap with Multimodal Attribute Graph
by: Li, Manyu, et al.
Published: (2026) -
A Benchmarking Study of Vision-based Robotic Grasping Algorithms
by: Rameshbabu, Bharath K, et al.
Published: (2025) -
RealVVT: Towards Photorealistic Video Virtual Try-on via Spatio-Temporal Consistency
by: Li, Siqi, et al.
Published: (2025) -
A Generative Framework for Self-Supervised Facial Representation Learning
by: He, Ruian, et al.
Published: (2023)