Rethinking Interactive Image Segmentation with Low Latency, High Quality, and Diverse Prompts
Fuente:
arXiv
Salvato in:
| Autori principali: | Liu, Qin, Cho, Jaemin, Bansal, Mohit, Niethammer, Marc |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Ctrl-Adapter: An Efficient and Versatile Framework for Adapting Diverse Controls to Any Diffusion Model
di: Lin, Han, et al.
Pubblicazione: (2024)
di: Lin, Han, et al.
Pubblicazione: (2024)
RotBench: Evaluating Multimodal Large Language Models on Identifying Image Rotation
di: Niu, Tianyi, et al.
Pubblicazione: (2025)
di: Niu, Tianyi, et al.
Pubblicazione: (2025)
On The Robustness of Foundational 3D Medical Image Segmentation Models Against Imprecise Visual Prompts
di: Chattopadhyay, Soumitri, et al.
Pubblicazione: (2026)
di: Chattopadhyay, Soumitri, et al.
Pubblicazione: (2026)
CLaMR: Contextualized Late-Interaction for Multimodal Content Retrieval
di: Wan, David, et al.
Pubblicazione: (2025)
di: Wan, David, et al.
Pubblicazione: (2025)
Self-Correcting Text-to-Video Generation with Misalignment Detection and Localized Refinement
di: Lee, Daeun, et al.
Pubblicazione: (2024)
di: Lee, Daeun, et al.
Pubblicazione: (2024)
Video-Skill-CoT: Skill-based Chain-of-Thoughts for Domain-Adaptive Video Reasoning
di: Lee, Daeun, et al.
Pubblicazione: (2025)
di: Lee, Daeun, et al.
Pubblicazione: (2025)
Exploring Cycle Consistency Learning in Interactive Volume Segmentation
di: Liu, Qin, et al.
Pubblicazione: (2023)
di: Liu, Qin, et al.
Pubblicazione: (2023)
Zero-shot Domain Generalization of Foundational Models for 3D Medical Image Segmentation: An Experimental Study
di: Chattopadhyay, Soumitri, et al.
Pubblicazione: (2025)
di: Chattopadhyay, Soumitri, et al.
Pubblicazione: (2025)
CAPTURe: Evaluating Spatial Reasoning in Vision Language Models via Occluded Object Counting
di: Pothiraj, Atin, et al.
Pubblicazione: (2025)
di: Pothiraj, Atin, et al.
Pubblicazione: (2025)
DiagrammerGPT: Generating Open-Domain, Open-Platform Diagrams via LLM Planning
di: Zala, Abhay, et al.
Pubblicazione: (2023)
di: Zala, Abhay, et al.
Pubblicazione: (2023)
VideoDirectorGPT: Consistent Multi-scene Video Generation via LLM-Guided Planning
di: Lin, Han, et al.
Pubblicazione: (2023)
di: Lin, Han, et al.
Pubblicazione: (2023)
M3DocRAG: Multi-modal Retrieval is What You Need for Multi-page Multi-document Understanding
di: Cho, Jaemin, et al.
Pubblicazione: (2024)
di: Cho, Jaemin, et al.
Pubblicazione: (2024)
Bifrost-1: Bridging Multimodal LLMs and Diffusion Models with Patch-level CLIP Latents
di: Lin, Han, et al.
Pubblicazione: (2025)
di: Lin, Han, et al.
Pubblicazione: (2025)
SELMA: Learning and Merging Skill-Specific Text-to-Image Experts with Auto-Generated Data
di: Li, Jialu, et al.
Pubblicazione: (2024)
di: Li, Jialu, et al.
Pubblicazione: (2024)
Contrastive Region Guidance: Improving Grounding in Vision-Language Models without Training
di: Wan, David, et al.
Pubblicazione: (2024)
di: Wan, David, et al.
Pubblicazione: (2024)
Diagnostic Benchmark and Iterative Inpainting for Layout-Guided Image Generation
di: Cho, Jaemin, et al.
Pubblicazione: (2023)
di: Cho, Jaemin, et al.
Pubblicazione: (2023)
AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
di: Wang, Zun, et al.
Pubblicazione: (2026)
di: Wang, Zun, et al.
Pubblicazione: (2026)
LiVOS: Light Video Object Segmentation with Gated Linear Matching
di: Liu, Qin, et al.
Pubblicazione: (2024)
di: Liu, Qin, et al.
Pubblicazione: (2024)
Training-free Guidance in Text-to-Video Generation via Multimodal Planning and Structured Noise Initialization
di: Li, Jialu, et al.
Pubblicazione: (2025)
di: Li, Jialu, et al.
Pubblicazione: (2025)
Investigating Demographic Bias in Brain MRI Segmentation: A Comparative Study of Deep-Learning and Non-Deep-Learning Methods
di: Danaee, Ghazal, et al.
Pubblicazione: (2025)
di: Danaee, Ghazal, et al.
Pubblicazione: (2025)
BALD-SAM: Disagreement-based Active Prompting in Interactive Segmentation
di: Chowdhury, Prithwijit, et al.
Pubblicazione: (2026)
di: Chowdhury, Prithwijit, et al.
Pubblicazione: (2026)
EPiC: Efficient Video Camera Control Learning with Precise Anchor-Video Guidance
di: Wang, Zun, et al.
Pubblicazione: (2025)
di: Wang, Zun, et al.
Pubblicazione: (2025)
V-Co: A Closer Look at Visual Representation Alignment via Co-Denoising
di: Lin, Han, et al.
Pubblicazione: (2026)
di: Lin, Han, et al.
Pubblicazione: (2026)
PicoSAM2: Low-Latency Segmentation In-Sensor for Edge Vision Applications
di: Bonazzi, Pietro, et al.
Pubblicazione: (2025)
di: Bonazzi, Pietro, et al.
Pubblicazione: (2025)
Temporal Prompting Matters: Rethinking Referring Video Object Segmentation
di: Lin, Ci-Siang, et al.
Pubblicazione: (2025)
di: Lin, Ci-Siang, et al.
Pubblicazione: (2025)
SPT: Sequence Prompt Transformer for Interactive Image Segmentation
di: Cheng, Senlin, et al.
Pubblicazione: (2024)
di: Cheng, Senlin, et al.
Pubblicazione: (2024)
Semantic Prompting with Image-Token for Continual Learning
di: Han, Jisu, et al.
Pubblicazione: (2024)
di: Han, Jisu, et al.
Pubblicazione: (2024)
VisionCoach: Reinforcing Grounded Video Reasoning via Visual-Perception Prompting
di: Lee, Daeun, et al.
Pubblicazione: (2026)
di: Lee, Daeun, et al.
Pubblicazione: (2026)
PA-SAM: Prompt Adapter SAM for High-Quality Image Segmentation
di: Xie, Zhaozhi, et al.
Pubblicazione: (2024)
di: Xie, Zhaozhi, et al.
Pubblicazione: (2024)
Benchmarking Human and Automated Prompting in the Segment Anything Model
di: Quesada, Jorge, et al.
Pubblicazione: (2024)
di: Quesada, Jorge, et al.
Pubblicazione: (2024)
Guiding Registration with Emergent Similarity from Pre-Trained Diffusion Models
di: Tursynbek, Nurislam, et al.
Pubblicazione: (2025)
di: Tursynbek, Nurislam, et al.
Pubblicazione: (2025)
$\texttt{NePhi}$: Neural Deformation Fields for Approximately Diffeomorphic Medical Image Registration
di: Tian, Lin, et al.
Pubblicazione: (2023)
di: Tian, Lin, et al.
Pubblicazione: (2023)
ROS-SAM: High-Quality Interactive Segmentation for Remote Sensing Moving Object
di: Shan, Zhe, et al.
Pubblicazione: (2025)
di: Shan, Zhe, et al.
Pubblicazione: (2025)
PromptIQA: Boosting the Performance and Generalization for No-Reference Image Quality Assessment via Prompts
di: Chen, Zewen, et al.
Pubblicazione: (2024)
di: Chen, Zewen, et al.
Pubblicazione: (2024)
ESAM++: Efficient Online 3D Perception on the Edge
di: Liu, Qin, et al.
Pubblicazione: (2026)
di: Liu, Qin, et al.
Pubblicazione: (2026)
A Unified Model for Longitudinal Multi-Modal Multi-View Prediction with Missingness
di: Chen, Boqi, et al.
Pubblicazione: (2024)
di: Chen, Boqi, et al.
Pubblicazione: (2024)
SAM-REF: Introducing Image-Prompt Synergy during Interaction for Detail Enhancement in the Segment Anything Model
di: Yu, Chongkai, et al.
Pubblicazione: (2024)
di: Yu, Chongkai, et al.
Pubblicazione: (2024)
Davidsonian Scene Graph: Improving Reliability in Fine-grained Evaluation for Text-to-Image Generation
di: Cho, Jaemin, et al.
Pubblicazione: (2023)
di: Cho, Jaemin, et al.
Pubblicazione: (2023)
Dynamic Prompt Generation for Interactive 3D Medical Image Segmentation Training
di: Ndir, Tidiane Camaret, et al.
Pubblicazione: (2025)
di: Ndir, Tidiane Camaret, et al.
Pubblicazione: (2025)
PPBoost: Progressive Prompt Boosting for Text-Driven Medical Image Segmentation
di: Li, Xuchen, et al.
Pubblicazione: (2025)
di: Li, Xuchen, et al.
Pubblicazione: (2025)
Documenti analoghi
-
Ctrl-Adapter: An Efficient and Versatile Framework for Adapting Diverse Controls to Any Diffusion Model
di: Lin, Han, et al.
Pubblicazione: (2024) -
RotBench: Evaluating Multimodal Large Language Models on Identifying Image Rotation
di: Niu, Tianyi, et al.
Pubblicazione: (2025) -
On The Robustness of Foundational 3D Medical Image Segmentation Models Against Imprecise Visual Prompts
di: Chattopadhyay, Soumitri, et al.
Pubblicazione: (2026) -
CLaMR: Contextualized Late-Interaction for Multimodal Content Retrieval
di: Wan, David, et al.
Pubblicazione: (2025) -
Self-Correcting Text-to-Video Generation with Misalignment Detection and Localized Refinement
di: Lee, Daeun, et al.
Pubblicazione: (2024)