Unified Open-World Segmentation with Multi-Modal Prompts
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Liu, Yang, Yin, Yufei, Jing, Chenchen, Zhu, Muzhi, Chen, Hao, Xi, Yuling, Feng, Bo, Wang, Hao, Li, Shiyu, Shen, Chunhua |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
A Simple Image Segmentation Framework via In-Context Examples
von: Liu, Yang, et al.
Veröffentlicht: (2024)
von: Liu, Yang, et al.
Veröffentlicht: (2024)
OmniJigsaw: Enhancing Omni-Modal Reasoning via Modality-Orchestrated Reordering
von: Jia, Yiduo, et al.
Veröffentlicht: (2026)
von: Jia, Yiduo, et al.
Veröffentlicht: (2026)
Preserving Source Video Realism: High-Fidelity Face Swapping for Cinematic Quality
von: Luo, Zekai, et al.
Veröffentlicht: (2025)
von: Luo, Zekai, et al.
Veröffentlicht: (2025)
Unleashing the Potential of the Diffusion Model in Few-shot Semantic Segmentation
von: Zhu, Muzhi, et al.
Veröffentlicht: (2024)
von: Zhu, Muzhi, et al.
Veröffentlicht: (2024)
Matcher: Segment Anything with One Shot Using All-Purpose Feature Matching
von: Liu, Yang, et al.
Veröffentlicht: (2023)
von: Liu, Yang, et al.
Veröffentlicht: (2023)
Generative Active Learning for Long-tailed Instance Segmentation
von: Zhu, Muzhi, et al.
Veröffentlicht: (2024)
von: Zhu, Muzhi, et al.
Veröffentlicht: (2024)
DiverGen: Improving Instance Segmentation by Learning Wider Data Distribution with More Diverse Generative Data
von: Fan, Chengxiang, et al.
Veröffentlicht: (2024)
von: Fan, Chengxiang, et al.
Veröffentlicht: (2024)
MMControl: Unified Multi-Modal Control for Joint Audio-Video Generation
von: Li, Liyang, et al.
Veröffentlicht: (2026)
von: Li, Liyang, et al.
Veröffentlicht: (2026)
CoLeCLIP: Open-Domain Continual Learning via Joint Task Prompt and Vocabulary Learning
von: Li, Yukun, et al.
Veröffentlicht: (2024)
von: Li, Yukun, et al.
Veröffentlicht: (2024)
Multi-Modal Prototypes for Open-World Semantic Segmentation
von: Yang, Yuhuan, et al.
Veröffentlicht: (2023)
von: Yang, Yuhuan, et al.
Veröffentlicht: (2023)
SegAgent: Exploring Pixel Understanding Capabilities in MLLMs by Imitating Human Annotator Trajectories
von: Zhu, Muzhi, et al.
Veröffentlicht: (2025)
von: Zhu, Muzhi, et al.
Veröffentlicht: (2025)
Modality-Agnostic Prompt Learning for Multi-Modal Camouflaged Object Detection
von: Wang, Hao, et al.
Veröffentlicht: (2026)
von: Wang, Hao, et al.
Veröffentlicht: (2026)
Omni-R1: Reinforcement Learning for Omnimodal Reasoning via Two-System Collaboration
von: Zhong, Hao, et al.
Veröffentlicht: (2025)
von: Zhong, Hao, et al.
Veröffentlicht: (2025)
Synergistic Prompting for Robust Visual Recognition with Missing Modalities
von: Zhang, Zhihui, et al.
Veröffentlicht: (2025)
von: Zhang, Zhihui, et al.
Veröffentlicht: (2025)
Where to Look: Can Foundation Models Reach a Target Viewpoint Through Active Exploration?
von: Li, Liyang, et al.
Veröffentlicht: (2026)
von: Li, Liyang, et al.
Veröffentlicht: (2026)
Reducing Semantic Ambiguity In Domain Adaptive Semantic Segmentation Via Probabilistic Prototypical Pixel Contrast
von: Hao, Xiaoke, et al.
Veröffentlicht: (2024)
von: Hao, Xiaoke, et al.
Veröffentlicht: (2024)
Towards Training-free Open-world Segmentation via Image Prompt Foundation Models
von: Tang, Lv, et al.
Veröffentlicht: (2023)
von: Tang, Lv, et al.
Veröffentlicht: (2023)
HieraTok: Multi-Scale Visual Tokenizer Improves Image Reconstruction and Generation
von: Chen, Cong, et al.
Veröffentlicht: (2025)
von: Chen, Cong, et al.
Veröffentlicht: (2025)
DICEPTION: A Generalist Diffusion Model for Visual Perceptual Tasks
von: Zhao, Canyu, et al.
Veröffentlicht: (2025)
von: Zhao, Canyu, et al.
Veröffentlicht: (2025)
PerturboLLaVA: Reducing Multimodal Hallucinations with Perturbative Visual Training
von: Chen, Cong, et al.
Veröffentlicht: (2025)
von: Chen, Cong, et al.
Veröffentlicht: (2025)
UniGeoSeg: Towards Unified Open-World Segmentation for Geospatial Scenes
von: Ni, Shuo, et al.
Veröffentlicht: (2025)
von: Ni, Shuo, et al.
Veröffentlicht: (2025)
OmniWorld: A Multi-Domain and Multi-Modal Dataset for 4D World Modeling
von: Zhou, Yang, et al.
Veröffentlicht: (2025)
von: Zhou, Yang, et al.
Veröffentlicht: (2025)
Scaling up Multi-domain Semantic Segmentation with Sentence Embeddings
von: Yin, Wei, et al.
Veröffentlicht: (2022)
von: Yin, Wei, et al.
Veröffentlicht: (2022)
Exploring Spatial Intelligence from a Generative Perspective
von: Zhu, Muzhi, et al.
Veröffentlicht: (2026)
von: Zhu, Muzhi, et al.
Veröffentlicht: (2026)
CLAPS: A CLIP-Unified Auto-Prompt Segmentation for Multi-Modal Retinal Imaging
von: Zhao, Zhihao, et al.
Veröffentlicht: (2025)
von: Zhao, Zhihao, et al.
Veröffentlicht: (2025)
Unified Embedding Alignment for Open-Vocabulary Video Instance Segmentation
von: Fang, Hao, et al.
Veröffentlicht: (2024)
von: Fang, Hao, et al.
Veröffentlicht: (2024)
VisionLLaMA: A Unified LLaMA Backbone for Vision Tasks
von: Chu, Xiangxiang, et al.
Veröffentlicht: (2024)
von: Chu, Xiangxiang, et al.
Veröffentlicht: (2024)
Segment then Splat: Unified 3D Open-Vocabulary Segmentation via Gaussian Splatting
von: Lu, Yiren, et al.
Veröffentlicht: (2025)
von: Lu, Yiren, et al.
Veröffentlicht: (2025)
Bridge Thinking and Acting: Unleashing Physical Potential of VLM with Generalizable Action Expert
von: Liu, Mingyu, et al.
Veröffentlicht: (2025)
von: Liu, Mingyu, et al.
Veröffentlicht: (2025)
Active-O3: Empowering Multimodal Large Language Models with Active Perception via GRPO
von: Zhu, Muzhi, et al.
Veröffentlicht: (2025)
von: Zhu, Muzhi, et al.
Veröffentlicht: (2025)
Exploring Conditional Multi-Modal Prompts for Zero-shot HOI Detection
von: Lei, Ting, et al.
Veröffentlicht: (2024)
von: Lei, Ting, et al.
Veröffentlicht: (2024)
CountGD: Multi-Modal Open-World Counting
von: Amini-Naieni, Niki, et al.
Veröffentlicht: (2024)
von: Amini-Naieni, Niki, et al.
Veröffentlicht: (2024)
OmniDFA: A Unified Framework for Open Set Synthesis Image Detection and Few-Shot Attribution
von: Wu, Shiyu, et al.
Veröffentlicht: (2025)
von: Wu, Shiyu, et al.
Veröffentlicht: (2025)
Tinker: Diffusion's Gift to 3D--Multi-View Consistent Editing From Sparse Inputs without Per-Scene Optimization
von: Zhao, Canyu, et al.
Veröffentlicht: (2025)
von: Zhao, Canyu, et al.
Veröffentlicht: (2025)
Recurrent Generic Contour-based Instance Segmentation with Progressive Learning
von: Feng, Hao, et al.
Veröffentlicht: (2023)
von: Feng, Hao, et al.
Veröffentlicht: (2023)
RGM: A Robust Generalizable Matching Model
von: Zhang, Songyan, et al.
Veröffentlicht: (2023)
von: Zhang, Songyan, et al.
Veröffentlicht: (2023)
Fine-grained Abnormality Prompt Learning for Zero-shot Anomaly Detection
von: Zhu, Jiawen, et al.
Veröffentlicht: (2024)
von: Zhu, Jiawen, et al.
Veröffentlicht: (2024)
RESAnything: Attribute Prompting for Arbitrary Referring Segmentation
von: Wang, Ruiqi, et al.
Veröffentlicht: (2025)
von: Wang, Ruiqi, et al.
Veröffentlicht: (2025)
ODYSSEY: Open-World Quadrupeds Exploration and Manipulation for Long-Horizon Tasks
von: Wang, Kaijun, et al.
Veröffentlicht: (2025)
von: Wang, Kaijun, et al.
Veröffentlicht: (2025)
Sub-Region-Aware Modality Fusion and Adaptive Prompting for Multi-Modal Brain Tumor Segmentation
von: Alijani, Shadi, et al.
Veröffentlicht: (2026)
von: Alijani, Shadi, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
A Simple Image Segmentation Framework via In-Context Examples
von: Liu, Yang, et al.
Veröffentlicht: (2024) -
OmniJigsaw: Enhancing Omni-Modal Reasoning via Modality-Orchestrated Reordering
von: Jia, Yiduo, et al.
Veröffentlicht: (2026) -
Preserving Source Video Realism: High-Fidelity Face Swapping for Cinematic Quality
von: Luo, Zekai, et al.
Veröffentlicht: (2025) -
Unleashing the Potential of the Diffusion Model in Few-shot Semantic Segmentation
von: Zhu, Muzhi, et al.
Veröffentlicht: (2024) -
Matcher: Segment Anything with One Shot Using All-Purpose Feature Matching
von: Liu, Yang, et al.
Veröffentlicht: (2023)