Unbiased Region-Language Alignment for Open-Vocabulary Dense Prediction
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Li, Yunheng, Li, Yuxuan, Zeng, Quansheng, Wang, Wenhai, Hou, Qibin, Cheng, Ming-Ming |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Cascade-CLIP: Cascaded Vision-Language Embeddings Alignment for Zero-Shot Semantic Segmentation
von: Li, Yunheng, et al.
Veröffentlicht: (2024)
von: Li, Yunheng, et al.
Veröffentlicht: (2024)
High-Quality Mask Tuning Matters for Open-Vocabulary Segmentation
von: Zeng, Quan-Sheng, et al.
Veröffentlicht: (2024)
von: Zeng, Quan-Sheng, et al.
Veröffentlicht: (2024)
Revisiting Efficient Semantic Segmentation: Learning Offsets for Better Spatial and Class Feature Alignment
von: Zhang, Shi-Chen, et al.
Veröffentlicht: (2025)
von: Zhang, Shi-Chen, et al.
Veröffentlicht: (2025)
A Glimpse to Compress: Dynamic Visual Token Pruning for Large Vision-Language Models
von: Zeng, Quan-Sheng, et al.
Veröffentlicht: (2025)
von: Zeng, Quan-Sheng, et al.
Veröffentlicht: (2025)
SM3Det: A Unified Model for Multi-Modal Remote Sensing Object Detection
von: Li, Yuxuan, et al.
Veröffentlicht: (2024)
von: Li, Yuxuan, et al.
Veröffentlicht: (2024)
TempSamp-R1: Effective Temporal Sampling with Reinforcement Fine-Tuning for Video LLMs
von: Li, Yunheng, et al.
Veröffentlicht: (2025)
von: Li, Yunheng, et al.
Veröffentlicht: (2025)
Unifying Heterogeneous Multi-Modal Remote Sensing Detection Via Language-Pivoted Pretraining
von: Li, Yuxuan, et al.
Veröffentlicht: (2026)
von: Li, Yuxuan, et al.
Veröffentlicht: (2026)
Rethinking Token-Level Policy Optimization for Multimodal Chain-of-Thought
von: Li, Yunheng, et al.
Veröffentlicht: (2026)
von: Li, Yunheng, et al.
Veröffentlicht: (2026)
Towards Universal Video MLLMs with Attribute-Structured and Quality-Verified Instructions
von: Li, Yunheng, et al.
Veröffentlicht: (2026)
von: Li, Yunheng, et al.
Veröffentlicht: (2026)
Dense Multimodal Alignment for Open-Vocabulary 3D Scene Understanding
von: Li, Ruihuang, et al.
Veröffentlicht: (2024)
von: Li, Ruihuang, et al.
Veröffentlicht: (2024)
Strip R-CNN: Large Strip Convolution for Remote Sensing Object Detection
von: Yuan, Xinbin, et al.
Veröffentlicht: (2025)
von: Yuan, Xinbin, et al.
Veröffentlicht: (2025)
PR-MIM: Delving Deeper into Partial Reconstruction in Masked Image Modeling
von: Li, Zhong-Yu, et al.
Veröffentlicht: (2024)
von: Li, Zhong-Yu, et al.
Veröffentlicht: (2024)
CLIPSelf: Vision Transformer Distills Itself for Open-Vocabulary Dense Prediction
von: Wu, Size, et al.
Veröffentlicht: (2023)
von: Wu, Size, et al.
Veröffentlicht: (2023)
MCANet: Medical Image Segmentation with Multi-Scale Cross-Axis Attention
von: Shao, Hao, et al.
Veröffentlicht: (2023)
von: Shao, Hao, et al.
Veröffentlicht: (2023)
SARDet-100K: Towards Open-Source Benchmark and ToolKit for Large-Scale SAR Object Detection
von: Li, Yuxuan, et al.
Veröffentlicht: (2024)
von: Li, Yuxuan, et al.
Veröffentlicht: (2024)
Open-Vocabulary Object Detection via Neighboring Region Attention Alignment
von: Qiang, Sunyuan, et al.
Veröffentlicht: (2024)
von: Qiang, Sunyuan, et al.
Veröffentlicht: (2024)
DFormer: Rethinking RGBD Representation Learning for Semantic Segmentation
von: Yin, Bowen, et al.
Veröffentlicht: (2023)
von: Yin, Bowen, et al.
Veröffentlicht: (2023)
Zone Evaluation: Revealing Spatial Bias in Object Detection
von: Zheng, Zhaohui, et al.
Veröffentlicht: (2023)
von: Zheng, Zhaohui, et al.
Veröffentlicht: (2023)
CrossKD: Cross-Head Knowledge Distillation for Object Detection
von: Wang, Jiabao, et al.
Veröffentlicht: (2023)
von: Wang, Jiabao, et al.
Veröffentlicht: (2023)
The Consistency Critic: Correcting Inconsistencies in Generated Images via Reference-Guided Attentive Alignment
von: Ouyang, Ziheng, et al.
Veröffentlicht: (2025)
von: Ouyang, Ziheng, et al.
Veröffentlicht: (2025)
DeCLIP: Decoupled Learning for Open-Vocabulary Dense Perception
von: Wang, Junjie, et al.
Veröffentlicht: (2025)
von: Wang, Junjie, et al.
Veröffentlicht: (2025)
GALA: Guided Attention with Language Alignment for Open Vocabulary Gaussian Splatting
von: Alegret, Elena, et al.
Veröffentlicht: (2025)
von: Alegret, Elena, et al.
Veröffentlicht: (2025)
Multi-Task Dense Prediction via Mixture of Low-Rank Experts
von: Yang, Yuqi, et al.
Veröffentlicht: (2024)
von: Yang, Yuqi, et al.
Veröffentlicht: (2024)
SRFormerV2: Taking a Closer Look at Permuted Self-Attention for Image Super-Resolution
von: Zhou, Yupeng, et al.
Veröffentlicht: (2023)
von: Zhou, Yupeng, et al.
Veröffentlicht: (2023)
Sora Generates Videos with Stunning Geometrical Consistency
von: Li, Xuanyi, et al.
Veröffentlicht: (2024)
von: Li, Xuanyi, et al.
Veröffentlicht: (2024)
Exploring Hierarchical Consistency and Unbiased Objectness for Open-Vocabulary Object Detection
von: Lee, Sanghoon, et al.
Veröffentlicht: (2026)
von: Lee, Sanghoon, et al.
Veröffentlicht: (2026)
LSKNet: A Foundation Lightweight Backbone for Remote Sensing
von: Li, Yuxuan, et al.
Veröffentlicht: (2024)
von: Li, Yuxuan, et al.
Veröffentlicht: (2024)
AR-1-to-3: Single Image to Consistent 3D Object Generation via Next-View Prediction
von: Zhang, Xuying, et al.
Veröffentlicht: (2025)
von: Zhang, Xuying, et al.
Veröffentlicht: (2025)
StoryDiffusion: Consistent Self-Attention for Long-Range Image and Video Generation
von: Zhou, Yupeng, et al.
Veröffentlicht: (2024)
von: Zhou, Yupeng, et al.
Veröffentlicht: (2024)
DFormerv2: Geometry Self-Attention for RGBD Semantic Segmentation
von: Yin, Bo-Wen, et al.
Veröffentlicht: (2025)
von: Yin, Bo-Wen, et al.
Veröffentlicht: (2025)
YOLO-MS: Rethinking Multi-Scale Representation Learning for Real-time Object Detection
von: Chen, Yuming, et al.
Veröffentlicht: (2023)
von: Chen, Yuming, et al.
Veröffentlicht: (2023)
ATAS: Any-to-Any Self-Distillation for Enhanced Open-Vocabulary Dense Prediction
von: Yeo, Juan, et al.
Veröffentlicht: (2025)
von: Yeo, Juan, et al.
Veröffentlicht: (2025)
OPUS: Occupancy Prediction Using a Sparse Set
von: Wang, Jiabao, et al.
Veröffentlicht: (2024)
von: Wang, Jiabao, et al.
Veröffentlicht: (2024)
ExpAlign: Expectation-Guided Vision-Language Alignment for Open-Vocabulary Grounding
von: Hu, Junyi, et al.
Veröffentlicht: (2026)
von: Hu, Junyi, et al.
Veröffentlicht: (2026)
WOW-Seg: A Word-free Open World Segmentation Model
von: Li, Danyang, et al.
Veröffentlicht: (2026)
von: Li, Danyang, et al.
Veröffentlicht: (2026)
Generalized Decoupled Learning for Enhancing Open-Vocabulary Dense Perception
von: Wang, Junjie, et al.
Veröffentlicht: (2025)
von: Wang, Junjie, et al.
Veröffentlicht: (2025)
Unified Embedding Alignment for Open-Vocabulary Video Instance Segmentation
von: Fang, Hao, et al.
Veröffentlicht: (2024)
von: Fang, Hao, et al.
Veröffentlicht: (2024)
Unified Dense Prediction of Video Diffusion
von: Yang, Lehan, et al.
Veröffentlicht: (2025)
von: Yang, Lehan, et al.
Veröffentlicht: (2025)
Towards Stable 3D Object Detection
von: Wang, Jiabao, et al.
Veröffentlicht: (2024)
von: Wang, Jiabao, et al.
Veröffentlicht: (2024)
Mixture of Style Experts for Diverse Image Stylization
von: Zhu, Shihao, et al.
Veröffentlicht: (2026)
von: Zhu, Shihao, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
Cascade-CLIP: Cascaded Vision-Language Embeddings Alignment for Zero-Shot Semantic Segmentation
von: Li, Yunheng, et al.
Veröffentlicht: (2024) -
High-Quality Mask Tuning Matters for Open-Vocabulary Segmentation
von: Zeng, Quan-Sheng, et al.
Veröffentlicht: (2024) -
Revisiting Efficient Semantic Segmentation: Learning Offsets for Better Spatial and Class Feature Alignment
von: Zhang, Shi-Chen, et al.
Veröffentlicht: (2025) -
A Glimpse to Compress: Dynamic Visual Token Pruning for Large Vision-Language Models
von: Zeng, Quan-Sheng, et al.
Veröffentlicht: (2025) -
SM3Det: A Unified Model for Multi-Modal Remote Sensing Object Detection
von: Li, Yuxuan, et al.
Veröffentlicht: (2024)