PixelRefer: A Unified Framework for Spatio-Temporal Object Referring with Arbitrary Granularity
Fuente:
arXiv
Saved in:
| Main Authors: | Yuan, Yuqian, Zhang, Wenqiao, Li, Xin, Wang, Shihao, Li, Kehan, Li, Wentong, Xiao, Jun, Zhang, Lei, Ooi, Beng Chin |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
VideoRefer Suite: Advancing Spatial-Temporal Object Understanding with Video LLM
by: Yuan, Yuqian, et al.
Published: (2024)
by: Yuan, Yuqian, et al.
Published: (2024)
LMMs Meet Object-Centric Vision: Understanding, Segmentation, Editing and Generation
by: Yuan, Yuqian, et al.
Published: (2026)
by: Yuan, Yuqian, et al.
Published: (2026)
METER: A Dynamic Concept Adaptation Framework for Online Anomaly Detection
by: Zhu, Jiaqi, et al.
Published: (2023)
by: Zhu, Jiaqi, et al.
Published: (2023)
Distribution-aware Online Continual Learning for Urban Spatio-Temporal Forecasting
by: Wang, Chengxin, et al.
Published: (2024)
by: Wang, Chengxin, et al.
Published: (2024)
CtrlCoT: Dual-Granularity Chain-of-Thought Compression for Controllable Reasoning
by: Fan, Zhenxuan, et al.
Published: (2026)
by: Fan, Zhenxuan, et al.
Published: (2026)
EOC-Bench: Can MLLMs Identify, Recall, and Forecast Objects in an Egocentric World?
by: Yuan, Yuqian, et al.
Published: (2025)
by: Yuan, Yuqian, et al.
Published: (2025)
HealthGPT: A Medical Large Vision-Language Model for Unifying Comprehension and Generation via Heterogeneous Knowledge Adaptation
by: Lin, Tianwei, et al.
Published: (2025)
by: Lin, Tianwei, et al.
Published: (2025)
InstructSAM: Segment Any Instance with Any Instructions
by: Yuan, Yuqian, et al.
Published: (2026)
by: Yuan, Yuqian, et al.
Published: (2026)
Catching Every Ripple: Enhanced Anomaly Awareness via Dynamic Concept Adaptation
by: Zhu, Jiaqi, et al.
Published: (2026)
by: Zhu, Jiaqi, et al.
Published: (2026)
Osprey: Pixel Understanding with Visual Instruction Tuning
by: Yuan, Yuqian, et al.
Published: (2023)
by: Yuan, Yuqian, et al.
Published: (2023)
UniPixel: Unified Object Referring and Segmentation for Pixel-Level Visual Reasoning
by: Liu, Ye, et al.
Published: (2025)
by: Liu, Ye, et al.
Published: (2025)
HeartcareGPT: A Unified Multimodal ECG Suite for Dual Signal-Image Modeling and Understanding
by: Xie, Yihan, et al.
Published: (2025)
by: Xie, Yihan, et al.
Published: (2025)
IQA-Spider: Unifying Multi-Granularity Image Quality Assessment with Reasoning, Grounding and Referring
by: Peng, Xinge, et al.
Published: (2026)
by: Peng, Xinge, et al.
Published: (2026)
Toward Robust Signed Graph Learning through Joint Input-Target Denoising
by: Wu, Junran, et al.
Published: (2025)
by: Wu, Junran, et al.
Published: (2025)
EyecareGPT: Boosting Comprehensive Ophthalmology Understanding with Tailored Dataset, Benchmark and Model
by: Li, Sijing, et al.
Published: (2025)
by: Li, Sijing, et al.
Published: (2025)
CLEAR-Mamba:Towards Accurate, Adaptive and Trustworthy Multi-Sequence Ophthalmic Angiography Classification
by: Wang, Zhuonan, et al.
Published: (2026)
by: Wang, Zhuonan, et al.
Published: (2026)
AllSpark: A Multimodal Spatio-Temporal General Intelligence Model with Ten Modalities via Language as a Reference Framework
by: Shao, Run, et al.
Published: (2023)
by: Shao, Run, et al.
Published: (2023)
U-STS-LLM A Unified Spatio-Temporal Steered Large Language Model for Traffic Prediction and Imputation
by: Zhang, Yichen, et al.
Published: (2026)
by: Zhang, Yichen, et al.
Published: (2026)
OmniCT: Towards a Unified Slice-Volume LVLM for Comprehensive CT Analysis
by: Lin, Tianwei, et al.
Published: (2026)
by: Lin, Tianwei, et al.
Published: (2026)
Harnessing Vision-Language Pretrained Models with Temporal-Aware Adaptation for Referring Video Object Segmentation
by: Zhou, Zikun, et al.
Published: (2024)
by: Zhou, Zikun, et al.
Published: (2024)
DUET: A Tuning-Free Device-Cloud Collaborative Parameters Generation Framework for Efficient Device Model Generalization
by: Lv, Zheqi, et al.
Published: (2022)
by: Lv, Zheqi, et al.
Published: (2022)
IBISAgent: Reinforcing Pixel-Level Visual Reasoning in MLLMs for Universal Biomedical Object Referring and Segmentation
by: Jiang, Yankai, et al.
Published: (2026)
by: Jiang, Yankai, et al.
Published: (2026)
Unveiling Parts Beyond Objects:Towards Finer-Granularity Referring Expression Segmentation
by: Wang, Wenxuan, et al.
Published: (2023)
by: Wang, Wenxuan, et al.
Published: (2023)
RESAnything: Attribute Prompting for Arbitrary Referring Segmentation
by: Wang, Ruiqi, et al.
Published: (2025)
by: Wang, Ruiqi, et al.
Published: (2025)
Temporal-Enhanced Multimodal Transformer for Referring Multi-Object Tracking and Segmentation
by: Xiao, Changcheng, et al.
Published: (2024)
by: Xiao, Changcheng, et al.
Published: (2024)
Spatio-Temporal Hierarchical Causal Models
by: Li, Xintong, et al.
Published: (2025)
by: Li, Xintong, et al.
Published: (2025)
Rethinking Reference Trajectories in Agile Drone Racing: A Unified Reference-Free Model-Based Controller via MPPI
by: Zhao, Fangguo, et al.
Published: (2025)
by: Zhao, Fangguo, et al.
Published: (2025)
Towards Unified Referring Expression Segmentation Across Omni-Level Visual Target Granularities
by: Liu, Jing, et al.
Published: (2025)
by: Liu, Jing, et al.
Published: (2025)
Generative AI for Healthcare: Fundamentals, Challenges, and Perspectives
by: Chen, Gang, et al.
Published: (2025)
by: Chen, Gang, et al.
Published: (2025)
Temporal-Conditional Referring Video Object Segmentation with Noise-Free Text-to-Video Diffusion Model
by: Zhang, Ruixin, et al.
Published: (2025)
by: Zhang, Ruixin, et al.
Published: (2025)
A$^2$-Edit: Precise Reference-Guided Image Editing of Arbitrary Objects and Ambiguous Masks
by: Zheng, Huayu, et al.
Published: (2026)
by: Zheng, Huayu, et al.
Published: (2026)
PILOT: Planning via Internalized Latent Optimization Trajectories for Large Language Models
by: Zheng, Haoyu, et al.
Published: (2026)
by: Zheng, Haoyu, et al.
Published: (2026)
AnyMS: Bottom-up Attention Decoupling for Layout-guided and Training-free Multi-subject Customization
by: Yu, Binhe, et al.
Published: (2025)
by: Yu, Binhe, et al.
Published: (2025)
Pixel-SAIL: Single Transformer For Pixel-Grounded Understanding
by: Zhang, Tao, et al.
Published: (2025)
by: Zhang, Tao, et al.
Published: (2025)
VecAug: Unveiling Camouflaged Frauds with Cohort Augmentation for Enhanced Detection
by: Xiao, Fei, et al.
Published: (2024)
by: Xiao, Fei, et al.
Published: (2024)
SVAC: Scaling Is All You Need For Referring Video Object Segmentation
by: Zhang, Li, et al.
Published: (2025)
by: Zhang, Li, et al.
Published: (2025)
A Unified Framework for Modeling Heterogeneous Financial Data via Dual-Granularity Prompting
by: Lei, Yu, et al.
Published: (2024)
by: Lei, Yu, et al.
Published: (2024)
STER-VLM: Spatio-Temporal With Enhanced Reference Vision-Language Models
by: Nguyen-Nhu, Tinh-Anh, et al.
Published: (2025)
by: Nguyen-Nhu, Tinh-Anh, et al.
Published: (2025)
A Unified Model for Spatio-Temporal Prediction Queries with Arbitrary Modifiable Areal Units
by: Chen, Liyue, et al.
Published: (2024)
by: Chen, Liyue, et al.
Published: (2024)
CrossView Suite: Harnessing Cross-view Spatial Intelligence of MLLMs with Dataset, Model and Benchmark
by: Wang, Wei, et al.
Published: (2026)
by: Wang, Wei, et al.
Published: (2026)
Similar Items
-
VideoRefer Suite: Advancing Spatial-Temporal Object Understanding with Video LLM
by: Yuan, Yuqian, et al.
Published: (2024) -
LMMs Meet Object-Centric Vision: Understanding, Segmentation, Editing and Generation
by: Yuan, Yuqian, et al.
Published: (2026) -
METER: A Dynamic Concept Adaptation Framework for Online Anomaly Detection
by: Zhu, Jiaqi, et al.
Published: (2023) -
Distribution-aware Online Continual Learning for Urban Spatio-Temporal Forecasting
by: Wang, Chengxin, et al.
Published: (2024) -
CtrlCoT: Dual-Granularity Chain-of-Thought Compression for Controllable Reasoning
by: Fan, Zhenxuan, et al.
Published: (2026)