Universal Segmentation at Arbitrary Granularity with Language Instruction
Fuente:
arXiv
Saved in:
| Main Authors: | Liu, Yong, Zhang, Cairong, Wang, Yitong, Wang, Jiahao, Yang, Yujiu, Tang, Yansong |
|---|---|
| Format: | Preprint |
| Published: |
2023
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Stepping Out of Similar Semantic Space for Open-Vocabulary Segmentation
by: Liu, Yong, et al.
Published: (2025)
by: Liu, Yong, et al.
Published: (2025)
1st Place Solution for 5th LSVOS Challenge: Referring Video Object Segmentation
by: Luo, Zhuoyan, et al.
Published: (2024)
by: Luo, Zhuoyan, et al.
Published: (2024)
Open-Vocabulary Segmentation with Semantic-Assisted Calibration
by: Liu, Yong, et al.
Published: (2023)
by: Liu, Yong, et al.
Published: (2023)
HyperSeg: Towards Universal Visual Segmentation with Large Language Model
by: Wei, Cong, et al.
Published: (2024)
by: Wei, Cong, et al.
Published: (2024)
Fully Aligned Network for Referring Image Segmentation
by: Liu, Yong, et al.
Published: (2024)
by: Liu, Yong, et al.
Published: (2024)
SAM2-LOVE: Segment Anything Model 2 in Language-aided Audio-Visual Scenes
by: Wang, Yuji, et al.
Published: (2025)
by: Wang, Yuji, et al.
Published: (2025)
HPT++: Hierarchically Prompting Vision-Language Models with Multi-Granularity Knowledge Generation and Improved Structure Modeling
by: Wang, Yubin, et al.
Published: (2024)
by: Wang, Yubin, et al.
Published: (2024)
IteRPrimE: Zero-shot Referring Image Segmentation with Iterative Grad-CAM Refinement and Primary Word Emphasis
by: Wang, Yuji, et al.
Published: (2025)
by: Wang, Yuji, et al.
Published: (2025)
DreamLight: Towards Harmonious and Consistent Image Relighting
by: Liu, Yong, et al.
Published: (2025)
by: Liu, Yong, et al.
Published: (2025)
InstructSeg: Unifying Instructed Visual Segmentation with Multi-modal Large Language Models
by: Wei, Cong, et al.
Published: (2024)
by: Wei, Cong, et al.
Published: (2024)
RESAnything: Attribute Prompting for Arbitrary Referring Segmentation
by: Wang, Ruiqi, et al.
Published: (2025)
by: Wang, Ruiqi, et al.
Published: (2025)
Towards Unified Referring Expression Segmentation Across Omni-Level Visual Target Granularities
by: Liu, Jing, et al.
Published: (2025)
by: Liu, Jing, et al.
Published: (2025)
Segment and Caption Anything
by: Huang, Xiaoke, et al.
Published: (2023)
by: Huang, Xiaoke, et al.
Published: (2023)
Segment Anything with Motion, Geometry, and Semantic Adaptation for Complex Nonlinear Visual Object Tracking
by: Zhu, Deyi, et al.
Published: (2026)
by: Zhu, Deyi, et al.
Published: (2026)
Instruction-guided Multi-Granularity Segmentation and Captioning with Large Multimodal Model
by: Zhou, Li, et al.
Published: (2024)
by: Zhou, Li, et al.
Published: (2024)
FDDet: Achieving Data-Efficient Food Defect Detection Under Real-World Scenarios
by: Xu, Ruihao, et al.
Published: (2026)
by: Xu, Ruihao, et al.
Published: (2026)
LaSagnA: Language-based Segmentation Assistant for Complex Queries
by: Wei, Cong, et al.
Published: (2024)
by: Wei, Cong, et al.
Published: (2024)
Self-Calibrated CLIP for Training-Free Open-Vocabulary Segmentation
by: Bai, Sule, et al.
Published: (2024)
by: Bai, Sule, et al.
Published: (2024)
PixelRefer: A Unified Framework for Spatio-Temporal Object Referring with Arbitrary Granularity
by: Yuan, Yuqian, et al.
Published: (2025)
by: Yuan, Yuqian, et al.
Published: (2025)
Unveiling Parts Beyond Objects:Towards Finer-Granularity Referring Expression Segmentation
by: Wang, Wenxuan, et al.
Published: (2023)
by: Wang, Wenxuan, et al.
Published: (2023)
HOI-aware Adaptive Network for Weakly-supervised Action Segmentation
by: Zhang, Runzhong, et al.
Published: (2026)
by: Zhang, Runzhong, et al.
Published: (2026)
Arbitrary Reading Order Scene Text Spotter with Local Semantics Guidance
by: Lyu, Jiahao, et al.
Published: (2024)
by: Lyu, Jiahao, et al.
Published: (2024)
UniVG-R1: Reasoning Guided Universal Visual Grounding with Reinforcement Learning
by: Bai, Sule, et al.
Published: (2025)
by: Bai, Sule, et al.
Published: (2025)
Mitigating the Reasoning Tax in Vision-Language Fine-Tuning with Input-Adaptive Depth Aggregation
by: Ren, Yiming, et al.
Published: (2026)
by: Ren, Yiming, et al.
Published: (2026)
Flash-VStream: Efficient Real-Time Understanding for Long Video Streams
by: Zhang, Haoji, et al.
Published: (2025)
by: Zhang, Haoji, et al.
Published: (2025)
GAGS: Granularity-Aware Feature Distillation for Language Gaussian Splatting
by: Peng, Yuning, et al.
Published: (2024)
by: Peng, Yuning, et al.
Published: (2024)
Learning Dual-Level Deformable Implicit Representation for Real-World Scale Arbitrary Super-Resolution
by: Li, Zhiheng, et al.
Published: (2024)
by: Li, Zhiheng, et al.
Published: (2024)
CoHD: A Counting-Aware Hierarchical Decoding Framework for Generalized Referring Expression Segmentation
by: Luo, Zhuoyan, et al.
Published: (2024)
by: Luo, Zhuoyan, et al.
Published: (2024)
WildSeg3D: Segment Any 3D Objects in the Wild from 2D Images
by: Guo, Yansong, et al.
Published: (2025)
by: Guo, Yansong, et al.
Published: (2025)
GraCo: Granularity-Controllable Interactive Segmentation
by: Zhao, Yian, et al.
Published: (2024)
by: Zhao, Yian, et al.
Published: (2024)
Language-free Compositional Action Generation via Decoupling Refinement
by: Liu, Xiao, et al.
Published: (2023)
by: Liu, Xiao, et al.
Published: (2023)
E4S: Fine-grained Face Swapping via Editing With Regional GAN Inversion
by: Li, Maomao, et al.
Published: (2023)
by: Li, Maomao, et al.
Published: (2023)
MGIMM: Multi-Granularity Instruction Multimodal Model for Attribute-Guided Remote Sensing Image Detailed Description
by: Yang, Cong, et al.
Published: (2024)
by: Yang, Cong, et al.
Published: (2024)
MSP-MVS: Multi-Granularity Segmentation Prior Guided Multi-View Stereo
by: Yuan, Zhenlong, et al.
Published: (2024)
by: Yuan, Zhenlong, et al.
Published: (2024)
Target Refocusing via Attention Redistribution for Open-Vocabulary Semantic Segmentation: An Explainability Perspective
by: Li, Jiahao, et al.
Published: (2025)
by: Li, Jiahao, et al.
Published: (2025)
Towards Effective Multi-Moving-Camera Tracking: A New Dataset and Lightweight Link Model
by: Zhang, Yanting, et al.
Published: (2023)
by: Zhang, Yanting, et al.
Published: (2023)
Ponder & Press: Advancing Visual GUI Agent towards General Computer Control
by: Wang, Yiqin, et al.
Published: (2024)
by: Wang, Yiqin, et al.
Published: (2024)
Multi-Granularity Video Object Segmentation
by: Lim, Sangbeom, et al.
Published: (2024)
by: Lim, Sangbeom, et al.
Published: (2024)
Hierarchical Memory for Long Video QA
by: Wang, Yiqin, et al.
Published: (2024)
by: Wang, Yiqin, et al.
Published: (2024)
Flash-VStream: Memory-Based Real-Time Understanding for Long Video Streams
by: Zhang, Haoji, et al.
Published: (2024)
by: Zhang, Haoji, et al.
Published: (2024)
Similar Items
-
Stepping Out of Similar Semantic Space for Open-Vocabulary Segmentation
by: Liu, Yong, et al.
Published: (2025) -
1st Place Solution for 5th LSVOS Challenge: Referring Video Object Segmentation
by: Luo, Zhuoyan, et al.
Published: (2024) -
Open-Vocabulary Segmentation with Semantic-Assisted Calibration
by: Liu, Yong, et al.
Published: (2023) -
HyperSeg: Towards Universal Visual Segmentation with Large Language Model
by: Wei, Cong, et al.
Published: (2024) -
Fully Aligned Network for Referring Image Segmentation
by: Liu, Yong, et al.
Published: (2024)