Moondream Segmentation: From Words to Masks
Fuente:
arXiv
Saved in:
| Main Author: | Reid, Ethan |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Segmenting Visuals With Querying Words: Language Anchors For Semi-Supervised Image Segmentation
by: Nadeem, Numair, et al.
Published: (2025)
by: Nadeem, Numair, et al.
Published: (2025)
Bridging Pixels and Words: Mask-Aware Local Semantic Fusion for Multimodal Media Verification
by: Chen, Zizhao, et al.
Published: (2026)
by: Chen, Zizhao, et al.
Published: (2026)
Open-Vocabulary Segmentation with Unpaired Mask-Text Supervision
by: Wang, Zhaoqing, et al.
Published: (2024)
by: Wang, Zhaoqing, et al.
Published: (2024)
AMLRIS: Alignment-aware Masked Learning for Referring Image Segmentation
by: Chen, Tongfei, et al.
Published: (2026)
by: Chen, Tongfei, et al.
Published: (2026)
Refer to Any Segmentation Mask Group With Vision-Language Prompts
by: Cao, Shengcao, et al.
Published: (2025)
by: Cao, Shengcao, et al.
Published: (2025)
PEM: Prototype-based Efficient MaskFormer for Image Segmentation
by: Cavagnero, Niccolò, et al.
Published: (2024)
by: Cavagnero, Niccolò, et al.
Published: (2024)
MaskUno: Switch-Split Block For Enhancing Instance Segmentation
by: Haidar, Jawad, et al.
Published: (2024)
by: Haidar, Jawad, et al.
Published: (2024)
Dual form Complementary Masking for Domain-Adaptive Image Segmentation
by: Wang, Jiawen, et al.
Published: (2025)
by: Wang, Jiawen, et al.
Published: (2025)
Efficient Masked Attention Transformer for Few-Shot Classification and Segmentation
by: Carrión-Ojeda, Dustin, et al.
Published: (2025)
by: Carrión-Ojeda, Dustin, et al.
Published: (2025)
Instance-aware Image Colorization with Controllable Textual Descriptions and Segmentation Masks
by: An, Yanru, et al.
Published: (2025)
by: An, Yanru, et al.
Published: (2025)
Complementary Random Masking for RGB-Thermal Semantic Segmentation
by: Shin, Ukcheol, et al.
Published: (2023)
by: Shin, Ukcheol, et al.
Published: (2023)
Foundation Model-Driven Framework for Human-Object Interaction Prediction with Segmentation Mask Integration
by: Park, Juhan, et al.
Published: (2025)
by: Park, Juhan, et al.
Published: (2025)
From Pixels to Words -- Towards Native Vision-Language Primitives at Scale
by: Diao, Haiwen, et al.
Published: (2025)
by: Diao, Haiwen, et al.
Published: (2025)
PR-MaGIC: Prompt Refinement Via Mask Decoder Gradient Flow For In-Context Segmentation
by: Lee, Minjae, et al.
Published: (2026)
by: Lee, Minjae, et al.
Published: (2026)
Class-agnostic 3D Segmentation by Granularity-Consistent Automatic 2D Mask Tracking
by: Wang, Juan, et al.
Published: (2025)
by: Wang, Juan, et al.
Published: (2025)
Boosting Semi-Supervised Medical Image Segmentation via Masked Image Consistency and Discrepancy Learning
by: Zhou, Pengcheng, et al.
Published: (2025)
by: Zhou, Pengcheng, et al.
Published: (2025)
Free-Mask: A Novel Paradigm of Integration Between the Segmentation Diffusion Model and Image Editing
by: Gao, Bo, et al.
Published: (2024)
by: Gao, Bo, et al.
Published: (2024)
Unlocking the Potential of MLLMs in Referring Expression Segmentation via a Light-weight Mask Decoder
by: Wang, Jingchao, et al.
Published: (2025)
by: Wang, Jingchao, et al.
Published: (2025)
X-SAM: From Segment Anything to Any Segmentation
by: Wang, Hao, et al.
Published: (2025)
by: Wang, Hao, et al.
Published: (2025)
WordVIS: A Color Worth A Thousand Words
by: Khan, Umar, et al.
Published: (2024)
by: Khan, Umar, et al.
Published: (2024)
XMask3D: Cross-modal Mask Reasoning for Open Vocabulary 3D Semantic Segmentation
by: Wang, Ziyi, et al.
Published: (2024)
by: Wang, Ziyi, et al.
Published: (2024)
Self Pre-training with Topology- and Spatiality-aware Masked Autoencoders for 3D Medical Image Segmentation
by: Gu, Pengfei, et al.
Published: (2024)
by: Gu, Pengfei, et al.
Published: (2024)
Synthetic Data-Driven Multi-Architecture Framework for Automated Polyp Segmentation Through Integrated Detection and Mask Generation
by: Peter, Ojonugwa Oluwafemi Ejiga, et al.
Published: (2025)
by: Peter, Ojonugwa Oluwafemi Ejiga, et al.
Published: (2025)
SegGen: Supercharging Segmentation Models with Text2Mask and Mask2Img Synthesis
by: Ye, Hanrong, et al.
Published: (2023)
by: Ye, Hanrong, et al.
Published: (2023)
Cognitive Alignment At No Cost: Inducing Human Attention Biases For Interpretable Vision Transformers
by: Knights, Ethan
Published: (2026)
by: Knights, Ethan
Published: (2026)
Sim4Seg: Boosting Multimodal Multi-disease Medical Diagnosis Segmentation with Region-Aware Vision-Language Similarity Masks
by: Song, Lingran, et al.
Published: (2025)
by: Song, Lingran, et al.
Published: (2025)
Unleashing the Potential of Vision-Language Pre-Training for 3D Zero-Shot Lesion Segmentation via Mask-Attribute Alignment
by: Jiang, Yankai, et al.
Published: (2024)
by: Jiang, Yankai, et al.
Published: (2024)
Enhancing Autonomous Vehicle Perception in Adverse Weather through Image Augmentation during Semantic Segmentation Training
by: Kou, Ethan, et al.
Published: (2024)
by: Kou, Ethan, et al.
Published: (2024)
Gaussian Masked Autoencoders
by: Rajasegaran, Jathushan, et al.
Published: (2025)
by: Rajasegaran, Jathushan, et al.
Published: (2025)
Triple Point Masking
by: Liu, Jiaming, et al.
Published: (2024)
by: Liu, Jiaming, et al.
Published: (2024)
SyncMask: Synchronized Attentional Masking for Fashion-centric Vision-Language Pretraining
by: Song, Chull Hwan, et al.
Published: (2024)
by: Song, Chull Hwan, et al.
Published: (2024)
Masked Latent Transformer with the Random Masking Ratio to Advance the Diagnosis of Dental Fluorosis
by: Wu, Yun, et al.
Published: (2024)
by: Wu, Yun, et al.
Published: (2024)
HU-based Foreground Masking for 3D Medical Masked Image Modeling
by: Lee, Jin, et al.
Published: (2025)
by: Lee, Jin, et al.
Published: (2025)
Visual Words Meet BM25: Sparse Auto-Encoder Visual Word Scoring for Image Retrieval
by: Han, Donghoon, et al.
Published: (2026)
by: Han, Donghoon, et al.
Published: (2026)
MaskAnyNet: Rethinking Masked Image Regions as Valuable Information in Supervised Learning
by: Hong, Jingshan, et al.
Published: (2025)
by: Hong, Jingshan, et al.
Published: (2025)
Mask-aware Text-to-Image Retrieval: Referring Expression Segmentation Meets Cross-modal Retrieval
by: Shen, Li-Cheng, et al.
Published: (2025)
by: Shen, Li-Cheng, et al.
Published: (2025)
CertMask: Certifiable Defense Against Adversarial Patches via Theoretically Optimal Mask Coverage
by: Lyu, Xuntao, et al.
Published: (2025)
by: Lyu, Xuntao, et al.
Published: (2025)
From Diffusion to Rectified Flow: Rethinking Text-Based Segmentation
by: Qu, Zishen, et al.
Published: (2026)
by: Qu, Zishen, et al.
Published: (2026)
From Question to Exploration: Test-Time Adaptation in Semantic Segmentation?
by: Yi, Chang'an, et al.
Published: (2023)
by: Yi, Chang'an, et al.
Published: (2023)
Ukrainian Visual Word Sense Disambiguation Benchmark
by: Laba, Yurii, et al.
Published: (2026)
by: Laba, Yurii, et al.
Published: (2026)
Similar Items
-
Segmenting Visuals With Querying Words: Language Anchors For Semi-Supervised Image Segmentation
by: Nadeem, Numair, et al.
Published: (2025) -
Bridging Pixels and Words: Mask-Aware Local Semantic Fusion for Multimodal Media Verification
by: Chen, Zizhao, et al.
Published: (2026) -
Open-Vocabulary Segmentation with Unpaired Mask-Text Supervision
by: Wang, Zhaoqing, et al.
Published: (2024) -
AMLRIS: Alignment-aware Masked Learning for Referring Image Segmentation
by: Chen, Tongfei, et al.
Published: (2026) -
Refer to Any Segmentation Mask Group With Vision-Language Prompts
by: Cao, Shengcao, et al.
Published: (2025)