Vision-Language Model for Object Detection and Segmentation: A Review and Evaluation
Fuente:
arXiv
Guardado en:
| Autores principales: | Feng, Yongchao, Liu, Yajie, Yang, Shuai, Cai, Wenrui, Zhang, Jinqing, Zhan, Qiqi, Huang, Ziyue, Yan, Hongxi, Wan, Qiao, Liu, Chenguang, Wang, Junzhe, Lv, Jiahui, Liu, Ziqi, Shi, Tengyuan, Liu, Qingjie, Wang, Yunhong |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
OpenRSD: Towards Open-prompts for Object Detection in Remote Sensing Images
por: Huang, Ziyue, et al.
Publicado: (2025)
por: Huang, Ziyue, et al.
Publicado: (2025)
MutDet: Mutually Optimizing Pre-training for Remote Sensing Object Detection
por: Huang, Ziyue, et al.
Publicado: (2024)
por: Huang, Ziyue, et al.
Publicado: (2024)
PACF: Prototype Augmented Compact Features for Improving Domain Adaptive Object Detection
por: Liu, Chenguang, et al.
Publicado: (2025)
por: Liu, Chenguang, et al.
Publicado: (2025)
Uni-MDTrack: Learning Decoupled Memory and Dynamic States for Parameter-Efficient Visual Tracking in All Modality
por: Cai, Wenrui, et al.
Publicado: (2026)
por: Cai, Wenrui, et al.
Publicado: (2026)
HIPTrack: Visual Tracking with Historical Prompts
por: Cai, Wenrui, et al.
Publicado: (2023)
por: Cai, Wenrui, et al.
Publicado: (2023)
SPMTrack: Spatio-Temporal Parameter-Efficient Fine-Tuning with Mixture of Experts for Scalable Visual Tracking
por: Cai, Wenrui, et al.
Publicado: (2025)
por: Cai, Wenrui, et al.
Publicado: (2025)
A Survey on Remote Sensing Foundation Models: From Vision to Multimodality
por: Huang, Ziyue, et al.
Publicado: (2025)
por: Huang, Ziyue, et al.
Publicado: (2025)
Lightweight Spatial Embedding for Vision-based 3D Occupancy Prediction
por: Zhang, Jinqing, et al.
Publicado: (2024)
por: Zhang, Jinqing, et al.
Publicado: (2024)
EntroCut: Entropy-Guided Adaptive Truncation for Efficient Chain-of-Thought Reasoning in Small-scale Large Reasoning Models
por: Yan, Hongxi, et al.
Publicado: (2026)
por: Yan, Hongxi, et al.
Publicado: (2026)
YOLC: You Only Look Clusters for Tiny Object Detection in Aerial Images
por: Liu, Chenguang, et al.
Publicado: (2024)
por: Liu, Chenguang, et al.
Publicado: (2024)
Incremental Object Detection with CLIP
por: Huang, Ziyue, et al.
Publicado: (2023)
por: Huang, Ziyue, et al.
Publicado: (2023)
AttriPrompt: Dynamic Prompt Composition Learning for CLIP
por: Zhan, Qiqi, et al.
Publicado: (2025)
por: Zhan, Qiqi, et al.
Publicado: (2025)
Beyond Open Vocabulary: Multimodal Prompting for Object Detection in Remote Sensing Images
por: Yang, Shuai, et al.
Publicado: (2026)
por: Yang, Shuai, et al.
Publicado: (2026)
SkeletonX: Data-Efficient Skeleton-based Action Recognition via Cross-sample Feature Aggregation
por: Zhang, Zongye, et al.
Publicado: (2025)
por: Zhang, Zongye, et al.
Publicado: (2025)
DSD-DA: Distillation-based Source Debiasing for Domain Adaptive Object Detection
por: Feng, Yongchao, et al.
Publicado: (2023)
por: Feng, Yongchao, et al.
Publicado: (2023)
De-Simplifying Pseudo Labels to Enhancing Domain Adaptive Object Detection
por: Fu, Zehua, et al.
Publicado: (2025)
por: Fu, Zehua, et al.
Publicado: (2025)
GeoBEV: Learning Geometric BEV Representation for Multi-view 3D Object Detection
por: Zhang, Jinqing, et al.
Publicado: (2024)
por: Zhang, Jinqing, et al.
Publicado: (2024)
FSD-BEV: Foreground Self-Distillation for Multi-view 3D Object Detection
por: Jiang, Zheng, et al.
Publicado: (2024)
por: Jiang, Zheng, et al.
Publicado: (2024)
Semantic Enhanced Few-shot Object Detection
por: Wang, Zheng, et al.
Publicado: (2024)
por: Wang, Zheng, et al.
Publicado: (2024)
ResWorld: Temporal Residual World Model for End-to-End Autonomous Driving
por: Zhang, Jinqing, et al.
Publicado: (2026)
por: Zhang, Jinqing, et al.
Publicado: (2026)
Generic Knowledge Boosted Pre-training For Remote Sensing Images
por: Huang, Ziyue, et al.
Publicado: (2024)
por: Huang, Ziyue, et al.
Publicado: (2024)
Learn More, Forget Less: A Gradient-Aware Data Selection Approach for LLM
por: Liu, Yibai, et al.
Publicado: (2025)
por: Liu, Yibai, et al.
Publicado: (2025)
CtxMIM: Context-Enhanced Masked Image Modeling for Remote Sensing Image Understanding
por: Zhang, Mingming, et al.
Publicado: (2023)
por: Zhang, Mingming, et al.
Publicado: (2023)
HiT: Building Mapping with Hierarchical Transformers
por: Zhang, Mingming, et al.
Publicado: (2023)
por: Zhang, Mingming, et al.
Publicado: (2023)
Context-Enhanced Detector For Building Detection From Remote Sensing Images
por: Huang, Ziyue, et al.
Publicado: (2023)
por: Huang, Ziyue, et al.
Publicado: (2023)
Multi-Grained Cross-modal Alignment for Learning Open-vocabulary Semantic Segmentation from Text Supervision
por: Liu, Yajie, et al.
Publicado: (2024)
por: Liu, Yajie, et al.
Publicado: (2024)
A Survey on Data Synthesis and Augmentation for Large Language Models
por: Wang, Ke, et al.
Publicado: (2024)
por: Wang, Ke, et al.
Publicado: (2024)
Towards Robust and Controllable Text-to-Motion via Masked Autoregressive Diffusion
por: Zhang, Zongye, et al.
Publicado: (2025)
por: Zhang, Zongye, et al.
Publicado: (2025)
SeeDNorm: Self-Rescaled Dynamic Normalization
por: Cai, Wenrui, et al.
Publicado: (2025)
por: Cai, Wenrui, et al.
Publicado: (2025)
ONER: Online Experience Replay for Incremental Anomaly Detection
por: Jin, Yizhou, et al.
Publicado: (2024)
por: Jin, Yizhou, et al.
Publicado: (2024)
Reasoning-Driven Anomaly Detection and Localization with Image-Level Supervision
por: Jin, Yizhou, et al.
Publicado: (2026)
por: Jin, Yizhou, et al.
Publicado: (2026)
LIBERO-X: Robustness Litmus for Vision-Language-Action Models
por: Wang, Guodong, et al.
Publicado: (2026)
por: Wang, Guodong, et al.
Publicado: (2026)
TimeGMM: Single-Pass Probabilistic Forecasting via Adaptive Gaussian Mixture Models with Reversible Normalization
por: Liu, Lei, et al.
Publicado: (2026)
por: Liu, Lei, et al.
Publicado: (2026)
Diffusion Trajectory-guided Policy for Long-horizon Robot Manipulation
por: Fan, Shichao, et al.
Publicado: (2025)
por: Fan, Shichao, et al.
Publicado: (2025)
ActiveDC: Distribution Calibration for Active Finetuning
por: Xu, Wenshuai, et al.
Publicado: (2023)
por: Xu, Wenshuai, et al.
Publicado: (2023)
ContextQFormer: A New Context Modeling Method for Multi-Turn Multi-Modal Conversations
por: Lei, Yiming, et al.
Publicado: (2025)
por: Lei, Yiming, et al.
Publicado: (2025)
SeriesBench: A Benchmark for Narrative-Driven Drama Series Understanding
por: Zhang, Chenkai, et al.
Publicado: (2025)
por: Zhang, Chenkai, et al.
Publicado: (2025)
GODBench: A Benchmark for Multimodal Large Language Models in Video Comment Art
por: Lei, Yiming, et al.
Publicado: (2025)
por: Lei, Yiming, et al.
Publicado: (2025)
On the pancyclicity of $2$-connected $[5,3]$-graphs
por: Liu, Feng, et al.
Publicado: (2025)
por: Liu, Feng, et al.
Publicado: (2025)
Phys-Diff: A Physics-Inspired Latent Diffusion Model for Tropical Cyclone Forecasting
por: Liu, Lei, et al.
Publicado: (2026)
por: Liu, Lei, et al.
Publicado: (2026)
Ejemplares similares
-
OpenRSD: Towards Open-prompts for Object Detection in Remote Sensing Images
por: Huang, Ziyue, et al.
Publicado: (2025) -
MutDet: Mutually Optimizing Pre-training for Remote Sensing Object Detection
por: Huang, Ziyue, et al.
Publicado: (2024) -
PACF: Prototype Augmented Compact Features for Improving Domain Adaptive Object Detection
por: Liu, Chenguang, et al.
Publicado: (2025) -
Uni-MDTrack: Learning Decoupled Memory and Dynamic States for Parameter-Efficient Visual Tracking in All Modality
por: Cai, Wenrui, et al.
Publicado: (2026) -
HIPTrack: Visual Tracking with Historical Prompts
por: Cai, Wenrui, et al.
Publicado: (2023)