OMG-Seg: Is One Model Good Enough For All Segmentation?
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Li, Xiangtai, Yuan, Haobo, Li, Wei, Ding, Henghui, Wu, Size, Zhang, Wenwei, Li, Yining, Chen, Kai, Loy, Chen Change |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Open-Vocabulary SAM: Segment and Recognize Twenty-thousand Classes Interactively
von: Yuan, Haobo, et al.
Veröffentlicht: (2024)
von: Yuan, Haobo, et al.
Veröffentlicht: (2024)
Transformer-Based Visual Segmentation: A Survey
von: Li, Xiangtai, et al.
Veröffentlicht: (2023)
von: Li, Xiangtai, et al.
Veröffentlicht: (2023)
OMG-LLaVA: Bridging Image-level, Object-level, Pixel-level Reasoning and Understanding
von: Zhang, Tao, et al.
Veröffentlicht: (2024)
von: Zhang, Tao, et al.
Veröffentlicht: (2024)
DST-Det: Simple Dynamic Self-Training for Open-Vocabulary Object Detection
von: Xu, Shilin, et al.
Veröffentlicht: (2023)
von: Xu, Shilin, et al.
Veröffentlicht: (2023)
CLIPSelf: Vision Transformer Distills Itself for Open-Vocabulary Dense Prediction
von: Wu, Size, et al.
Veröffentlicht: (2023)
von: Wu, Size, et al.
Veröffentlicht: (2023)
Explore In-Context Segmentation via Latent Diffusion Models
von: Wang, Chaoyang, et al.
Veröffentlicht: (2024)
von: Wang, Chaoyang, et al.
Veröffentlicht: (2024)
F-LMM: Grounding Frozen Large Multimodal Models
von: Wu, Size, et al.
Veröffentlicht: (2024)
von: Wu, Size, et al.
Veröffentlicht: (2024)
Mamba or RWKV: Exploring High-Quality and High-Efficiency Segment Anything Model
von: Yuan, Haobo, et al.
Veröffentlicht: (2024)
von: Yuan, Haobo, et al.
Veröffentlicht: (2024)
Controllable Human-centric Keyframe Interpolation with Generative Prior
von: Guo, Zujin, et al.
Veröffentlicht: (2025)
von: Guo, Zujin, et al.
Veröffentlicht: (2025)
EdgeSAM: Prompt-In-the-Loop Distillation for SAM
von: Zhou, Chong, et al.
Veröffentlicht: (2023)
von: Zhou, Chong, et al.
Veröffentlicht: (2023)
MosaicFusion: Diffusion Models as Data Augmenters for Large Vocabulary Instance Segmentation
von: Xie, Jiahao, et al.
Veröffentlicht: (2023)
von: Xie, Jiahao, et al.
Veröffentlicht: (2023)
Harmonizing Visual Representations for Unified Multimodal Understanding and Generation
von: Wu, Size, et al.
Veröffentlicht: (2025)
von: Wu, Size, et al.
Veröffentlicht: (2025)
Generalizable Implicit Motion Modeling for Video Frame Interpolation
von: Guo, Zujin, et al.
Veröffentlicht: (2024)
von: Guo, Zujin, et al.
Veröffentlicht: (2024)
Towards Language-Driven Video Inpainting via Multimodal Large Language Models
von: Wu, Jianzong, et al.
Veröffentlicht: (2024)
von: Wu, Jianzong, et al.
Veröffentlicht: (2024)
MOWA: Multiple-in-One Image Warping Model
von: Liao, Kang, et al.
Veröffentlicht: (2024)
von: Liao, Kang, et al.
Veröffentlicht: (2024)
OpenUni: A Simple Baseline for Unified Multimodal Understanding and Generation
von: Wu, Size, et al.
Veröffentlicht: (2025)
von: Wu, Size, et al.
Veröffentlicht: (2025)
MeViS: A Multi-Modal Dataset for Referring Motion Expression Video Segmentation
von: Ding, Henghui, et al.
Veröffentlicht: (2025)
von: Ding, Henghui, et al.
Veröffentlicht: (2025)
Thinking with Camera: A Unified Multimodal Model for Camera-Centric Understanding and Generation
von: Liao, Kang, et al.
Veröffentlicht: (2025)
von: Liao, Kang, et al.
Veröffentlicht: (2025)
Enhanced Generative Structure Prior for Chinese Text Image Super-resolution
von: Li, Xiaoming, et al.
Veröffentlicht: (2025)
von: Li, Xiaoming, et al.
Veröffentlicht: (2025)
Kalman-Inspired Feature Propagation for Video Face Super-Resolution
von: Feng, Ruicheng, et al.
Veröffentlicht: (2024)
von: Feng, Ruicheng, et al.
Veröffentlicht: (2024)
AITTI: Learning Adaptive Inclusive Token for Text-to-Image Generation
von: Hou, Xinyu, et al.
Veröffentlicht: (2024)
von: Hou, Xinyu, et al.
Veröffentlicht: (2024)
SegPoint: Segment Any Point Cloud via Large Language Model
von: He, Shuting, et al.
Veröffentlicht: (2024)
von: He, Shuting, et al.
Veröffentlicht: (2024)
RTMO: Towards High-Performance One-Stage Real-Time Multi-Person Pose Estimation
von: Lu, Peng, et al.
Veröffentlicht: (2023)
von: Lu, Peng, et al.
Veröffentlicht: (2023)
Point-In-Context: Understanding Point Cloud via In-Context Learning
von: Liu, Mengyuan, et al.
Veröffentlicht: (2024)
von: Liu, Mengyuan, et al.
Veröffentlicht: (2024)
UnSeg: One Universal Unlearnable Example Generator is Enough against All Image Segmentation
von: Sun, Ye, et al.
Veröffentlicht: (2024)
von: Sun, Ye, et al.
Veröffentlicht: (2024)
Eliminating Feature Ambiguity for Few-Shot Segmentation
von: Xu, Qianxiong, et al.
Veröffentlicht: (2024)
von: Xu, Qianxiong, et al.
Veröffentlicht: (2024)
DifFace: Blind Face Restoration with Diffused Error Contraction
von: Yue, Zongsheng, et al.
Veröffentlicht: (2022)
von: Yue, Zongsheng, et al.
Veröffentlicht: (2022)
MVIP-NeRF: Multi-view 3D Inpainting on NeRF Scenes via Diffusion Prior
von: Chen, Honghua, et al.
Veröffentlicht: (2024)
von: Chen, Honghua, et al.
Veröffentlicht: (2024)
SemFlow: Binding Semantic Segmentation and Image Synthesis via Rectified Flow
von: Wang, Chaoyang, et al.
Veröffentlicht: (2024)
von: Wang, Chaoyang, et al.
Veröffentlicht: (2024)
Omegance: A Single Parameter for Various Granularities in Diffusion-Based Synthesis
von: Hou, Xinyu, et al.
Veröffentlicht: (2024)
von: Hou, Xinyu, et al.
Veröffentlicht: (2024)
Trans-Adapter: A Plug-and-Play Framework for Transparent Image Inpainting
von: Dai, Yuekun, et al.
Veröffentlicht: (2025)
von: Dai, Yuekun, et al.
Veröffentlicht: (2025)
Contextual Object Detection with Multimodal Large Language Models
von: Zang, Yuhang, et al.
Veröffentlicht: (2023)
von: Zang, Yuhang, et al.
Veröffentlicht: (2023)
RMP-SAM: Towards Real-Time Multi-Purpose Segment Anything
von: Xu, Shilin, et al.
Veröffentlicht: (2024)
von: Xu, Shilin, et al.
Veröffentlicht: (2024)
Efficient Diffusion Model for Image Restoration by Residual Shifting
von: Yue, Zongsheng, et al.
Veröffentlicht: (2024)
von: Yue, Zongsheng, et al.
Veröffentlicht: (2024)
Control Color: Multimodal Diffusion-based Interactive Image Colorization
von: Liang, Zhexin, et al.
Veröffentlicht: (2024)
von: Liang, Zhexin, et al.
Veröffentlicht: (2024)
Learning Inclusion Matching for Animation Paint Bucket Colorization
von: Dai, Yuekun, et al.
Veröffentlicht: (2024)
von: Dai, Yuekun, et al.
Veröffentlicht: (2024)
Duolando: Follower GPT with Off-Policy Reinforcement Learning for Dance Accompaniment
von: Siyao, Li, et al.
Veröffentlicht: (2024)
von: Siyao, Li, et al.
Veröffentlicht: (2024)
Partial CLIP is Enough: Chimera-Seg for Zero-shot Semantic Segmentation
von: Chen, Jialei, et al.
Veröffentlicht: (2025)
von: Chen, Jialei, et al.
Veröffentlicht: (2025)
Auto Cherry-Picker: Learning from High-quality Generative Data Driven by Language
von: Chen, Yicheng, et al.
Veröffentlicht: (2024)
von: Chen, Yicheng, et al.
Veröffentlicht: (2024)
Learning 3D Garment Animation from Trajectories of A Piece of Cloth
von: Shao, Yidi, et al.
Veröffentlicht: (2025)
von: Shao, Yidi, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Open-Vocabulary SAM: Segment and Recognize Twenty-thousand Classes Interactively
von: Yuan, Haobo, et al.
Veröffentlicht: (2024) -
Transformer-Based Visual Segmentation: A Survey
von: Li, Xiangtai, et al.
Veröffentlicht: (2023) -
OMG-LLaVA: Bridging Image-level, Object-level, Pixel-level Reasoning and Understanding
von: Zhang, Tao, et al.
Veröffentlicht: (2024) -
DST-Det: Simple Dynamic Self-Training for Open-Vocabulary Object Detection
von: Xu, Shilin, et al.
Veröffentlicht: (2023) -
CLIPSelf: Vision Transformer Distills Itself for Open-Vocabulary Dense Prediction
von: Wu, Size, et al.
Veröffentlicht: (2023)