UMG-CLIP: A Unified Multi-Granularity Vision Generalist for Open-World Understanding
Fuente:
arXiv
Saved in:
| Main Authors: | Shi, Bowen, Zhao, Peisen, Wang, Zichen, Zhang, Yuhang, Wang, Yaoming, Li, Jin, Dai, Wenrui, Zou, Junni, Xiong, Hongkai, Tian, Qi, Zhang, Xiaopeng |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
GranViT: A Fine-Grained Vision Model With Autoregressive Perception For MLLMs
by: Zheng, Guanghao, et al.
Published: (2025)
by: Zheng, Guanghao, et al.
Published: (2025)
METEOR: Multi-Encoder Collaborative Token Pruning for Efficient Vision Language Models
by: Liu, Yuchen, et al.
Published: (2025)
by: Liu, Yuchen, et al.
Published: (2025)
Diffusion-Driven Progressive Target Manipulation for Source-Free Domain Adaptation
by: Huang, Yuyang, et al.
Published: (2025)
by: Huang, Yuyang, et al.
Published: (2025)
Point Cloud Denoising With Fine-Granularity Dynamic Graph Convolutional Networks
by: Xu, Wenqiang, et al.
Published: (2024)
by: Xu, Wenqiang, et al.
Published: (2024)
MimicNorm: Weight Mean and Last BN Layer Mimic the Dynamic of Batch Normalization
by: Fei, Wen, et al.
Published: (2020)
by: Fei, Wen, et al.
Published: (2020)
Noise Conditional Variational Score Distillation
by: Peng, Xinyu, et al.
Published: (2025)
by: Peng, Xinyu, et al.
Published: (2025)
Frequency-Aware Transformer for Learned Image Compression
by: Li, Han, et al.
Published: (2023)
by: Li, Han, et al.
Published: (2023)
Image Compression for Machine and Human Vision with Spatial-Frequency Adaptation
by: Li, Han, et al.
Published: (2024)
by: Li, Han, et al.
Published: (2024)
Towards Holistic Modeling for Video Frame Interpolation with Auto-regressive Diffusion Transformers
by: Peng, Xinyu, et al.
Published: (2026)
by: Peng, Xinyu, et al.
Published: (2026)
OneCAT: Decoder-Only Auto-Regressive Model for Unified Understanding and Generation
by: Li, Han, et al.
Published: (2025)
by: Li, Han, et al.
Published: (2025)
Point Cloud Resampling with Learnable Heat Diffusion
by: Xu, Wenqiang, et al.
Published: (2024)
by: Xu, Wenqiang, et al.
Published: (2024)
HiPART: Hierarchical Pose AutoRegressive Transformer for Occluded 3D Human Pose Estimation
by: Zheng, Hongwei, et al.
Published: (2025)
by: Zheng, Hongwei, et al.
Published: (2025)
Improving Diffusion Models for Inverse Problems Using Optimal Posterior Covariance
by: Peng, Xinyu, et al.
Published: (2024)
by: Peng, Xinyu, et al.
Published: (2024)
Error-Propagation-Free Learned Video Compression With Dual-Domain Progressive Temporal Alignment
by: Li, Han, et al.
Published: (2025)
by: Li, Han, et al.
Published: (2025)
Information-Theoretic Optimization for Task-Adapted Compressed Sensing Magnetic Resonance Imaging
by: Peng, Xinyu, et al.
Published: (2026)
by: Peng, Xinyu, et al.
Published: (2026)
3DGabSplat: 3D Gabor Splatting for Frequency-adaptive Radiance Field Rendering
by: Zhou, Junyu, et al.
Published: (2025)
by: Zhou, Junyu, et al.
Published: (2025)
On Disentangled Training for Nonlinear Transform in Learned Image Compression
by: Li, Han, et al.
Published: (2025)
by: Li, Han, et al.
Published: (2025)
LiftImage3D: Lifting Any Single Image to 3D Gaussians with Video Generation Priors
by: Chen, Yabo, et al.
Published: (2024)
by: Chen, Yabo, et al.
Published: (2024)
Cascade-Zero123: One Image to Highly Consistent 3D with Self-Prompted Nearby Views
by: Chen, Yabo, et al.
Published: (2023)
by: Chen, Yabo, et al.
Published: (2023)
CEIA: CLIP-Based Event-Image Alignment for Open-World Event-Based Understanding
by: Xu, Wenhao, et al.
Published: (2024)
by: Xu, Wenhao, et al.
Published: (2024)
DINO-X: A Unified Vision Model for Open-World Object Detection and Understanding
by: Ren, Tianhe, et al.
Published: (2024)
by: Ren, Tianhe, et al.
Published: (2024)
Temporal2Seq: A Unified Framework for Temporal Video Understanding Tasks
by: Yang, Min, et al.
Published: (2024)
by: Yang, Min, et al.
Published: (2024)
Unifying Biomedical Vision-Language Expertise: Towards a Generalist Foundation Model via Multi-CLIP Knowledge Distillation
by: Wang, Shansong, et al.
Published: (2025)
by: Wang, Shansong, et al.
Published: (2025)
UMind-VL: A Generalist Ultrasound Vision-Language Model for Unified Grounded Perception and Comprehensive Interpretation
by: Chen, Dengbo, et al.
Published: (2025)
by: Chen, Dengbo, et al.
Published: (2025)
From CLIP to DINO: Visual Encoders Shout in Multi-modal Large Language Models
by: Jiang, Dongsheng, et al.
Published: (2023)
by: Jiang, Dongsheng, et al.
Published: (2023)
SpikeMba: Multi-Modal Spiking Saliency Mamba for Temporal Video Grounding
by: Li, Wenrui, et al.
Published: (2024)
by: Li, Wenrui, et al.
Published: (2024)
Long-CLIP: Unlocking the Long-Text Capability of CLIP
by: Zhang, Beichen, et al.
Published: (2024)
by: Zhang, Beichen, et al.
Published: (2024)
Spiking Tucker Fusion Transformer for Audio-Visual Zero-Shot Learning
by: Li, Wenrui, et al.
Published: (2024)
by: Li, Wenrui, et al.
Published: (2024)
Towards Building Specialized Generalist AI with System 1 and System 2 Fusion
by: Zhang, Kaiyan, et al.
Published: (2024)
by: Zhang, Kaiyan, et al.
Published: (2024)
EV-NVC: Efficient Variable bitrate Neural Video Compression
by: Hu, Yongcun, et al.
Published: (2025)
by: Hu, Yongcun, et al.
Published: (2025)
Res$^2$CLIP: Few-Shot Generalist Anomaly Detection with Residual-to-Residual Alignment
by: Liu, Xinyue, et al.
Published: (2026)
by: Liu, Xinyue, et al.
Published: (2026)
Grounded Vision-Language Navigation for UAVs with Open-Vocabulary Goal Understanding
by: Zhang, Yuhang, et al.
Published: (2025)
by: Zhang, Yuhang, et al.
Published: (2025)
Medical Vision Generalist: Unifying Medical Imaging Tasks in Context
by: Ren, Sucheng, et al.
Published: (2024)
by: Ren, Sucheng, et al.
Published: (2024)
MME-Unify: A Comprehensive Benchmark for Unified Multimodal Understanding and Generation Models
by: Xie, Wulin, et al.
Published: (2025)
by: Xie, Wulin, et al.
Published: (2025)
A Low-Rank Method for Vision Language Model Hallucination Mitigation in Autonomous Driving
by: Long, Keke, et al.
Published: (2025)
by: Long, Keke, et al.
Published: (2025)
OpenSDI: Spotting Diffusion-Generated Images in the Open World
by: Wang, Yabin, et al.
Published: (2025)
by: Wang, Yabin, et al.
Published: (2025)
Lingshu: A Generalist Foundation Model for Unified Multimodal Medical Understanding and Reasoning
by: LASA Team, et al.
Published: (2025)
by: LASA Team, et al.
Published: (2025)
Refining Alignment Framework for Diffusion Models with Intermediate-Step Preference Ranking
by: Ren, Jie, et al.
Published: (2025)
by: Ren, Jie, et al.
Published: (2025)
MolSight: Optical Chemical Structure Recognition with SMILES Pretraining, Multi-Granularity Learning and Reinforcement Learning
by: Zhang, Wenrui, et al.
Published: (2025)
by: Zhang, Wenrui, et al.
Published: (2025)
FineViT: Progressively Unlocking Fine-Grained Perception with Dense Recaptions
by: Zhao, Peisen, et al.
Published: (2026)
by: Zhao, Peisen, et al.
Published: (2026)
Similar Items
-
GranViT: A Fine-Grained Vision Model With Autoregressive Perception For MLLMs
by: Zheng, Guanghao, et al.
Published: (2025) -
METEOR: Multi-Encoder Collaborative Token Pruning for Efficient Vision Language Models
by: Liu, Yuchen, et al.
Published: (2025) -
Diffusion-Driven Progressive Target Manipulation for Source-Free Domain Adaptation
by: Huang, Yuyang, et al.
Published: (2025) -
Point Cloud Denoising With Fine-Granularity Dynamic Graph Convolutional Networks
by: Xu, Wenqiang, et al.
Published: (2024) -
MimicNorm: Weight Mean and Last BN Layer Mimic the Dynamic of Batch Normalization
by: Fei, Wen, et al.
Published: (2020)