Text-guided Visual Prompt DINO for Generic Segmentation
Fuente:
arXiv
Salvato in:
| Autori principali: | Guan, Yuchen, Sun, Chong, Fu, Canmiao, Huang, Zhipeng, Yuan, Chun, Li, Chen |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
WinTok: A Win-Win Hybrid Tokenizer via Decomposing Visual Understanding and Generation with Transferable Tokens
di: Guo, Yiwei, et al.
Pubblicazione: (2026)
di: Guo, Yiwei, et al.
Pubblicazione: (2026)
PET-DINO: Unifying Visual Cues into Grounding DINO with Prompt-Enriched Training
di: Fu, Weifu, et al.
Pubblicazione: (2026)
di: Fu, Weifu, et al.
Pubblicazione: (2026)
WeGen: A Unified Model for Interactive Multimodal Generation as We Chat
di: Huang, Zhipeng, et al.
Pubblicazione: (2025)
di: Huang, Zhipeng, et al.
Pubblicazione: (2025)
DINO-Tok: Adapting DINO for Visual Tokenizers
di: Jia, Mingkai, et al.
Pubblicazione: (2025)
di: Jia, Mingkai, et al.
Pubblicazione: (2025)
WeTok: Powerful Discrete Tokenization for High-Fidelity Visual Reconstruction
di: Zhuang, Shaobin, et al.
Pubblicazione: (2025)
di: Zhuang, Shaobin, et al.
Pubblicazione: (2025)
Video-GPT via Next Clip Diffusion
di: Zhuang, Shaobin, et al.
Pubblicazione: (2025)
di: Zhuang, Shaobin, et al.
Pubblicazione: (2025)
Get In Video: Add Anything You Want to the Video
di: Zhuang, Shaobin, et al.
Pubblicazione: (2025)
di: Zhuang, Shaobin, et al.
Pubblicazione: (2025)
SegDINO: An Efficient Design for Medical and Natural Image Segmentation with DINO-V3
di: Yang, Sicheng, et al.
Pubblicazione: (2025)
di: Yang, Sicheng, et al.
Pubblicazione: (2025)
UniGlyph: Unified Segmentation-Conditioned Diffusion for Precise Visual Text Synthesis
di: Wang, Yuanrui, et al.
Pubblicazione: (2025)
di: Wang, Yuanrui, et al.
Pubblicazione: (2025)
ReferDINO: Referring Video Object Segmentation with Visual Grounding Foundations
di: Liang, Tianming, et al.
Pubblicazione: (2025)
di: Liang, Tianming, et al.
Pubblicazione: (2025)
Breaking Dual Bottlenecks: Evolving Unified Multimodal Models into Self-Adaptive Interleaved Visual Reasoners
di: Liu, Qingyang, et al.
Pubblicazione: (2026)
di: Liu, Qingyang, et al.
Pubblicazione: (2026)
Unlocking Generalization in Polyp Segmentation with DINO Self-Attention "keys"
di: Monteiro, Carla, et al.
Pubblicazione: (2025)
di: Monteiro, Carla, et al.
Pubblicazione: (2025)
FreqDINO: Frequency-Guided Adaptation for Generalized Boundary-Aware Ultrasound Image Segmentation
di: Zhang, Yixuan, et al.
Pubblicazione: (2025)
di: Zhang, Yixuan, et al.
Pubblicazione: (2025)
From CLIP to DINO: Visual Encoders Shout in Multi-modal Large Language Models
di: Jiang, Dongsheng, et al.
Pubblicazione: (2023)
di: Jiang, Dongsheng, et al.
Pubblicazione: (2023)
Empowering DINO Representations for Underwater Instance Segmentation via Aligner and Prompter
di: Chen, Zhiyang, et al.
Pubblicazione: (2025)
di: Chen, Zhiyang, et al.
Pubblicazione: (2025)
X-Prompt: Multi-modal Visual Prompt for Video Object Segmentation
di: Guo, Pinxue, et al.
Pubblicazione: (2024)
di: Guo, Pinxue, et al.
Pubblicazione: (2024)
UnicEdit-10M: A Dataset and Benchmark Breaking the Scale-Quality Barrier via Unified Verification for Reasoning-Enriched Edits
di: Ye, Keming, et al.
Pubblicazione: (2025)
di: Ye, Keming, et al.
Pubblicazione: (2025)
Topology-Agnostic Animal Motion Generation from Text Prompt
di: Chen, Keyi, et al.
Pubblicazione: (2025)
di: Chen, Keyi, et al.
Pubblicazione: (2025)
DINO-Foresight: Looking into the Future with DINO
di: Karypidis, Efstathios, et al.
Pubblicazione: (2024)
di: Karypidis, Efstathios, et al.
Pubblicazione: (2024)
Grounding DINO-US-SAM: Text-Prompted Multi-Organ Segmentation in Ultrasound with LoRA-Tuned Vision-Language Models
di: Rasaee, Hamza, et al.
Pubblicazione: (2025)
di: Rasaee, Hamza, et al.
Pubblicazione: (2025)
Human-Free Automated Prompting for Vision-Language Anomaly Detection: Prompt Optimization with Meta-guiding Prompt Scheme
di: Chen, Pi-Wei, et al.
Pubblicazione: (2024)
di: Chen, Pi-Wei, et al.
Pubblicazione: (2024)
Text-driven Multiplanar Visual Interaction for Semi-supervised Medical Image Segmentation
di: Huang, Kaiwen, et al.
Pubblicazione: (2025)
di: Huang, Kaiwen, et al.
Pubblicazione: (2025)
PrAda: Few-Shot Visual Adaptation for Text-Prompted Segmentation
di: Rosi, Gabriele, et al.
Pubblicazione: (2026)
di: Rosi, Gabriele, et al.
Pubblicazione: (2026)
GuiDINO: Rethinking Vision Foundation Model in Medical Image Segmentation
di: Liang, Zhuonan, et al.
Pubblicazione: (2026)
di: Liang, Zhuonan, et al.
Pubblicazione: (2026)
Glyph-ByT5: A Customized Text Encoder for Accurate Visual Text Rendering
di: Liu, Zeyu, et al.
Pubblicazione: (2024)
di: Liu, Zeyu, et al.
Pubblicazione: (2024)
Enhancing Medical Visual Grounding via Knowledge-guided Spatial Prompts
di: Gao, Yifan, et al.
Pubblicazione: (2026)
di: Gao, Yifan, et al.
Pubblicazione: (2026)
DI-MaskDINO: A Joint Object Detection and Instance Segmentation Model
di: Nan, Zhixiong, et al.
Pubblicazione: (2024)
di: Nan, Zhixiong, et al.
Pubblicazione: (2024)
DINO-MVR: Multi-View Readout of Frozen DINOv3 for Annotation-Efficient Medical Segmentation
di: Jiang, Wei, et al.
Pubblicazione: (2026)
di: Jiang, Wei, et al.
Pubblicazione: (2026)
Anatomy-Aware Text-Visual Fusion with Dual-Perspective Prompts for Fine-Grained Lumbar Spine Segmentation
di: Lian, Sheng, et al.
Pubblicazione: (2025)
di: Lian, Sheng, et al.
Pubblicazione: (2025)
Vision-guided and Mask-enhanced Adaptive Denoising for Prompt-based Image Editing
di: Wang, Kejie, et al.
Pubblicazione: (2024)
di: Wang, Kejie, et al.
Pubblicazione: (2024)
OVS-DINO: Open-Vocabulary Segmentation via Structure-Aligned SAM-DINO with Language Guidance
di: Zeng, Haoxi, et al.
Pubblicazione: (2026)
di: Zeng, Haoxi, et al.
Pubblicazione: (2026)
Enhancing Logits Distillation with Plug\&Play Kendall's $τ$ Ranking Loss
di: Guan, Yuchen, et al.
Pubblicazione: (2024)
di: Guan, Yuchen, et al.
Pubblicazione: (2024)
Evaluating Stenosis Detection with Grounding DINO, YOLO, and DINO-DETR
di: Ansari, Muhammad Musab
Pubblicazione: (2025)
di: Ansari, Muhammad Musab
Pubblicazione: (2025)
TV-3DG: Mastering Text-to-3D Customized Generation with Visual Prompt
di: Yang, Jiahui, et al.
Pubblicazione: (2024)
di: Yang, Jiahui, et al.
Pubblicazione: (2024)
PixelDINO: Semi-Supervised Semantic Segmentation for Detecting Permafrost Disturbances
di: Heidler, Konrad, et al.
Pubblicazione: (2024)
di: Heidler, Konrad, et al.
Pubblicazione: (2024)
Grounding DINO: Marrying DINO with Grounded Pre-Training for Open-Set Object Detection
di: Liu, Shilong, et al.
Pubblicazione: (2023)
di: Liu, Shilong, et al.
Pubblicazione: (2023)
Enhancing Label-efficient Medical Image Segmentation with Text-guided Diffusion Models
di: Feng, Chun-Mei
Pubblicazione: (2024)
di: Feng, Chun-Mei
Pubblicazione: (2024)
DINO-SLAM: DINO-informed RGB-D SLAM for Neural Implicit and Explicit Representations
di: Gong, Ziren, et al.
Pubblicazione: (2025)
di: Gong, Ziren, et al.
Pubblicazione: (2025)
Char-SAM: Turning Segment Anything Model into Scene Text Segmentation Annotator with Character-level Visual Prompts
di: Xie, Enze, et al.
Pubblicazione: (2024)
di: Xie, Enze, et al.
Pubblicazione: (2024)
TextInVision: Text and Prompt Complexity Driven Visual Text Generation Benchmark
di: Fallah, Forouzan, et al.
Pubblicazione: (2025)
di: Fallah, Forouzan, et al.
Pubblicazione: (2025)
Documenti analoghi
-
WinTok: A Win-Win Hybrid Tokenizer via Decomposing Visual Understanding and Generation with Transferable Tokens
di: Guo, Yiwei, et al.
Pubblicazione: (2026) -
PET-DINO: Unifying Visual Cues into Grounding DINO with Prompt-Enriched Training
di: Fu, Weifu, et al.
Pubblicazione: (2026) -
WeGen: A Unified Model for Interactive Multimodal Generation as We Chat
di: Huang, Zhipeng, et al.
Pubblicazione: (2025) -
DINO-Tok: Adapting DINO for Visual Tokenizers
di: Jia, Mingkai, et al.
Pubblicazione: (2025) -
WeTok: Powerful Discrete Tokenization for High-Fidelity Visual Reconstruction
di: Zhuang, Shaobin, et al.
Pubblicazione: (2025)