VLM-Guided Iterative Refinement for Surgical Image Segmentation with Foundation Models
Fuente:
arXiv
Saved in:
| Main Authors: | Lou, Ange, Li, Yamin, Chang, Qi, Xi, Nan, Xie, Luyuan, Li, Zichao, Luan, Tianyu |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
AstroVLM: Expert Multi-agent Collaborative Reasoning for Astronomical Imaging Quality Diagnosis
by: Han, Yaohui, et al.
Published: (2026)
by: Han, Yaohui, et al.
Published: (2026)
Surgical Depth Anything: Depth Estimation for Surgical Scenes using Foundation Models
by: Lou, Ange, et al.
Published: (2024)
by: Lou, Ange, et al.
Published: (2024)
AdaptFly: Prompt-Guided Adaptation of Foundation Models for Low-Altitude UAV Networks
by: Chen, Jiao, et al.
Published: (2025)
by: Chen, Jiao, et al.
Published: (2025)
Unified End-to-End V2X Cooperative Autonomous Driving
by: Li, Zhiwei, et al.
Published: (2024)
by: Li, Zhiwei, et al.
Published: (2024)
ReCCur: A Recursive Corner-Case Curation Framework for Robust Vision-Language Understanding in Open and Edge Scenarios
by: Wei, Yihan, et al.
Published: (2026)
by: Wei, Yihan, et al.
Published: (2026)
What Makes Good Collaborative Views? Contrastive Mutual Information Maximization for Multi-Agent Perception
by: Su, Wanfang, et al.
Published: (2024)
by: Su, Wanfang, et al.
Published: (2024)
SAMSNeRF: Segment Anything Model (SAM) Guides Dynamic Surgical Scene Reconstruction by Neural Radiance Field (NeRF)
by: Lou, Ange, et al.
Published: (2023)
by: Lou, Ange, et al.
Published: (2023)
FetalAgents: A Multi-Agent System for Fetal Ultrasound Image and Video Analysis
by: Hu, Xiaotian, et al.
Published: (2026)
by: Hu, Xiaotian, et al.
Published: (2026)
ProCrit: Self-Elicited Multi-Perspective Reasoning with Critic-Guided Revision for Multimodal Sarcasm Detection
by: Xu, Yingjia, et al.
Published: (2026)
by: Xu, Yingjia, et al.
Published: (2026)
GenCellAgent: Generalizable, Training-Free Cellular Image Segmentation via Large Language Model Agents
by: Yu, Xi, et al.
Published: (2025)
by: Yu, Xi, et al.
Published: (2025)
Collaborative Memory: Multi-User Memory Sharing in LLM Agents with Dynamic Access Control
by: Rezazadeh, Alireza, et al.
Published: (2025)
by: Rezazadeh, Alireza, et al.
Published: (2025)
LogiStory: A Logic-Aware Framework for Multi-Image Story Visualization
by: Meng, Chutian, et al.
Published: (2026)
by: Meng, Chutian, et al.
Published: (2026)
CollaMamba: Efficient Collaborative Perception with Cross-Agent Spatial-Temporal State Space Model
by: Li, Yang, et al.
Published: (2024)
by: Li, Yang, et al.
Published: (2024)
AutoGen Driven Multi Agent Framework for Iterative Crime Data Analysis and Prediction
by: Fatima, Syeda Kisaa, et al.
Published: (2025)
by: Fatima, Syeda Kisaa, et al.
Published: (2025)
TraF-Align: Trajectory-aware Feature Alignment for Asynchronous Multi-agent Perception
by: Song, Zhiying, et al.
Published: (2025)
by: Song, Zhiying, et al.
Published: (2025)
AdaReasoner: Dynamic Tool Orchestration for Iterative Visual Reasoning
by: Song, Mingyang, et al.
Published: (2026)
by: Song, Mingyang, et al.
Published: (2026)
HiMemFormer: Hierarchical Memory-Aware Transformer for Multi-Agent Action Anticipation
by: Wang, Zirui, et al.
Published: (2024)
by: Wang, Zirui, et al.
Published: (2024)
VideoMultiAgents: A Multi-Agent Framework for Video Question Answering
by: Kugo, Noriyuki, et al.
Published: (2025)
by: Kugo, Noriyuki, et al.
Published: (2025)
AniMaker: Multi-Agent Animated Storytelling with MCTS-Driven Clip Generation
by: Shi, Haoyuan, et al.
Published: (2025)
by: Shi, Haoyuan, et al.
Published: (2025)
Hollywood Town: Long-Video Generation via Cross-Modal Multi-Agent Orchestration
by: Wei, Zheng, et al.
Published: (2025)
by: Wei, Zheng, et al.
Published: (2025)
SynWorld: Virtual Scenario Synthesis for Agentic Action Knowledge Refinement
by: Fang, Runnan, et al.
Published: (2025)
by: Fang, Runnan, et al.
Published: (2025)
VideoChat-M1: Collaborative Policy Planning for Video Understanding via Multi-Agent Reinforcement Learning
by: Chen, Boyu, et al.
Published: (2025)
by: Chen, Boyu, et al.
Published: (2025)
Towards Reliable Fetal Ultrasound Interpretation with Multi-Agent Collaboration
by: Hu, Xiaotian, et al.
Published: (2026)
by: Hu, Xiaotian, et al.
Published: (2026)
Zero-Shot Surgical Tool Segmentation in Monocular Video Using Segment Anything Model 2
by: Lou, Ange, et al.
Published: (2024)
by: Lou, Ange, et al.
Published: (2024)
EI-Drive: A Platform for Cooperative Perception with Realistic Communication Models
by: Zhou, Hanchu, et al.
Published: (2024)
by: Zhou, Hanchu, et al.
Published: (2024)
Knowledge-Informed Multi-Agent Trajectory Prediction at Signalized Intersections for Infrastructure-to-Everything
by: Yin, Huilin, et al.
Published: (2025)
by: Yin, Huilin, et al.
Published: (2025)
C$^2$T: Captioning-Structure and LLM-Aligned Common-Sense Reward Learning for Traffic--Vehicle Coordination
by: Chen, Yuyang, et al.
Published: (2026)
by: Chen, Yuyang, et al.
Published: (2026)
IMAGAgent: Orchestrating Multi-Turn Image Editing via Constraint-Aware Planning and Reflection
by: Shen, Fei, et al.
Published: (2026)
by: Shen, Fei, et al.
Published: (2026)
Think Before You Segment: An Object-aware Reasoning Agent for Referring Audio-Visual Segmentation
by: Zhou, Jinxing, et al.
Published: (2025)
by: Zhou, Jinxing, et al.
Published: (2025)
UNCAP: Uncertainty-Guided Neurosymbolic Planning Using Natural Language Communication for Cooperative Autonomous Vehicles
by: Bhatt, Neel P., et al.
Published: (2025)
by: Bhatt, Neel P., et al.
Published: (2025)
Active Scout: Multi-Target Tracking Using Neural Radiance Fields in Dense Urban Environments
by: Hsu, Christopher D., et al.
Published: (2024)
by: Hsu, Christopher D., et al.
Published: (2024)
Visual Sensor Pose Optimisation Using Visibility Models for Smart Cities
by: Arnold, Eduardo, et al.
Published: (2021)
by: Arnold, Eduardo, et al.
Published: (2021)
Sentinel: Embodied Cooperative Spatial Reasoning and Planning
by: Lin, Xiangye, et al.
Published: (2026)
by: Lin, Xiangye, et al.
Published: (2026)
Cascading multi-agent anomaly detection in surveillance systems via vision-language models and embedding-based classification
by: Rehman, Tayyab, et al.
Published: (2026)
by: Rehman, Tayyab, et al.
Published: (2026)
Visual Multi-Agent System: Mitigating Hallucination Snowballing via Visual Flow
by: Yu, Xinlei, et al.
Published: (2025)
by: Yu, Xinlei, et al.
Published: (2025)
V2X-DGPE: Addressing Domain Gaps and Pose Errors for Robust Collaborative 3D Object Detection
by: Wang, Sichao, et al.
Published: (2025)
by: Wang, Sichao, et al.
Published: (2025)
Fast2comm:Collaborative perception combined with prior knowledge
by: Zhang, Zhengbin, et al.
Published: (2025)
by: Zhang, Zhengbin, et al.
Published: (2025)
FootBots: A Transformer-based Architecture for Motion Prediction in Soccer
by: Capellera, Guillem, et al.
Published: (2024)
by: Capellera, Guillem, et al.
Published: (2024)
Enhancing CLIP Robustness via Cross-Modality Alignment
by: Zhu, Xingyu, et al.
Published: (2025)
by: Zhu, Xingyu, et al.
Published: (2025)
MAG-3D: Multi-Agent Grounded Reasoning for 3D Understanding
by: Zheng, Henry, et al.
Published: (2026)
by: Zheng, Henry, et al.
Published: (2026)
Similar Items
-
AstroVLM: Expert Multi-agent Collaborative Reasoning for Astronomical Imaging Quality Diagnosis
by: Han, Yaohui, et al.
Published: (2026) -
Surgical Depth Anything: Depth Estimation for Surgical Scenes using Foundation Models
by: Lou, Ange, et al.
Published: (2024) -
AdaptFly: Prompt-Guided Adaptation of Foundation Models for Low-Altitude UAV Networks
by: Chen, Jiao, et al.
Published: (2025) -
Unified End-to-End V2X Cooperative Autonomous Driving
by: Li, Zhiwei, et al.
Published: (2024) -
ReCCur: A Recursive Corner-Case Curation Framework for Robust Vision-Language Understanding in Open and Edge Scenarios
by: Wei, Yihan, et al.
Published: (2026)