YOLO-World: Real-Time Open-Vocabulary Object Detection
Fuente:
arXiv
Saved in:
| Main Authors: | Cheng, Tianheng, Song, Lin, Ge, Yixiao, Liu, Wenyu, Wang, Xinggang, Shan, Ying |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Mask-Adapter: The Devil is in the Masks for Open-Vocabulary Segmentation
by: Li, Yongkang, et al.
Published: (2024)
by: Li, Yongkang, et al.
Published: (2024)
Cross-Layer Attentive Feature Upsampling for Low-latency Semantic Segmentation
by: Cheng, Tianheng, et al.
Published: (2026)
by: Cheng, Tianheng, et al.
Published: (2026)
Mamba-YOLO-World: Marrying YOLO-World with Mamba for Open-Vocabulary Detection
by: Wang, Haoxuan, et al.
Published: (2024)
by: Wang, Haoxuan, et al.
Published: (2024)
YOLO-UniOW: Efficient Universal Open-World Object Detection
by: Liu, Lihao, et al.
Published: (2024)
by: Liu, Lihao, et al.
Published: (2024)
Polar Parametrization for Vision-based Surround-View 3D Detection
by: Chen, Shaoyu, et al.
Published: (2022)
by: Chen, Shaoyu, et al.
Published: (2022)
Occupancy as Set of Points
by: Shi, Yiang, et al.
Published: (2024)
by: Shi, Yiang, et al.
Published: (2024)
GaussTR: Foundation Model-Aligned Gaussian Transformer for Self-Supervised 3D Spatial Understanding
by: Jiang, Haoyi, et al.
Published: (2024)
by: Jiang, Haoyi, et al.
Published: (2024)
YOLOE-26: Integrating YOLO26 with YOLOE for Real-Time Open-Vocabulary Instance Segmentation
by: Sapkota, Ranjan, et al.
Published: (2026)
by: Sapkota, Ranjan, et al.
Published: (2026)
YOLO-IOD: Towards Real Time Incremental Object Detection
by: Zhang, Shizhou, et al.
Published: (2025)
by: Zhang, Shizhou, et al.
Published: (2025)
Learning Background Prompts to Discover Implicit Knowledge for Open Vocabulary Object Detection
by: Li, Jiaming, et al.
Published: (2024)
by: Li, Jiaming, et al.
Published: (2024)
XS-VID: An Extremely Small Video Object Detection Dataset
by: Guo, Jiahao, et al.
Published: (2024)
by: Guo, Jiahao, et al.
Published: (2024)
RT-OVAD: Real-Time Open-Vocabulary Aerial Object Detection via Image-Text Collaboration
by: Wei, Guoting, et al.
Published: (2024)
by: Wei, Guoting, et al.
Published: (2024)
Open Vocabulary Monocular 3D Object Detection
by: Yao, Jin, et al.
Published: (2024)
by: Yao, Jin, et al.
Published: (2024)
OV-DEIM: Real-time DETR-Style Open-Vocabulary Object Detection with GridSynthetic Augmentation
by: Wang, Leilei, et al.
Published: (2026)
by: Wang, Leilei, et al.
Published: (2026)
FACTOR: Counterfactual Training-Free Test-Time Adaptation for Open-Vocabulary Object Detection
by: Zhao, Kaixiang, et al.
Published: (2026)
by: Zhao, Kaixiang, et al.
Published: (2026)
Multimodal Mamba: Decoder-only Multimodal State Space Model via Quadratic to Linear Distillation
by: Liao, Bencheng, et al.
Published: (2025)
by: Liao, Bencheng, et al.
Published: (2025)
Learning to Detect and Segment for Open Vocabulary Object Detection
by: Wang, Tao, et al.
Published: (2022)
by: Wang, Tao, et al.
Published: (2022)
EgoPlan-Bench2: A Benchmark for Multimodal Large Language Model Planning in Real-World Scenarios
by: Qiu, Lu, et al.
Published: (2024)
by: Qiu, Lu, et al.
Published: (2024)
FasterDiT: Towards Faster Diffusion Transformers Training without Architecture Modification
by: Yao, Jingfeng, et al.
Published: (2024)
by: Yao, Jingfeng, et al.
Published: (2024)
Scaling Open-Vocabulary Object Detection
by: Minderer, Matthias, et al.
Published: (2023)
by: Minderer, Matthias, et al.
Published: (2023)
MarvelOVD: Marrying Object Recognition and Vision-Language Models for Robust Open-Vocabulary Object Detection
by: Wang, Kuo, et al.
Published: (2024)
by: Wang, Kuo, et al.
Published: (2024)
PersonViT: Large-scale Self-supervised Vision Transformer for Person Re-Identification
by: Hu, Bin, et al.
Published: (2024)
by: Hu, Bin, et al.
Published: (2024)
Lane Graph as Path: Continuity-preserving Path-wise Modeling for Online Lane Graph Construction
by: Liao, Bencheng, et al.
Published: (2023)
by: Liao, Bencheng, et al.
Published: (2023)
ZS-VCOS: Zero-Shot Video Camouflaged Object Segmentation By Optical Flow and Open Vocabulary Object Detection
by: Guo, Wenqi, et al.
Published: (2025)
by: Guo, Wenqi, et al.
Published: (2025)
LW-DETR: A Transformer Replacement to YOLO for Real-Time Detection
by: Chen, Qiang, et al.
Published: (2024)
by: Chen, Qiang, et al.
Published: (2024)
YOLO26: Key Architectural Enhancements and Performance Benchmarking for Real-Time Object Detection
by: Sapkota, Ranjan, et al.
Published: (2025)
by: Sapkota, Ranjan, et al.
Published: (2025)
Meta-Adapter: An Online Few-shot Learner for Vision-Language Model
by: Cheng, Cheng, et al.
Published: (2023)
by: Cheng, Cheng, et al.
Published: (2023)
YOLO-MS: Rethinking Multi-Scale Representation Learning for Real-time Object Detection
by: Chen, Yuming, et al.
Published: (2023)
by: Chen, Yuming, et al.
Published: (2023)
MRS-YOLO Railroad Transmission Line Foreign Object Detection Based on Improved YOLO11 and Channel Pruning
by: Liu, Siyuan, et al.
Published: (2025)
by: Liu, Siyuan, et al.
Published: (2025)
ControlAR: Controllable Image Generation with Autoregressive Models
by: Li, Zongming, et al.
Published: (2024)
by: Li, Zongming, et al.
Published: (2024)
Retrieval-Augmented Open-Vocabulary Object Detection
by: Kim, Jooyeon, et al.
Published: (2024)
by: Kim, Jooyeon, et al.
Published: (2024)
Real-Time Object Detection and Classification using YOLO for Edge FPGAs
by: Amin, Rashed Al, et al.
Published: (2025)
by: Amin, Rashed Al, et al.
Published: (2025)
Open-Vocabulary Spatio-Temporal Action Detection
by: Wu, Tao, et al.
Published: (2024)
by: Wu, Tao, et al.
Published: (2024)
ODOV: Benchmark the Open-Domain Open-Vocabulary Object Detection
by: Zhang, Yupeng, et al.
Published: (2025)
by: Zhang, Yupeng, et al.
Published: (2025)
Boosting Open-Vocabulary Object Detection by Handling Background Samples
by: Zeng, Ruizhe, et al.
Published: (2024)
by: Zeng, Ruizhe, et al.
Published: (2024)
AnimeGamer: Infinite Anime Life Simulation with Next Game State Prediction
by: Cheng, Junhao, et al.
Published: (2025)
by: Cheng, Junhao, et al.
Published: (2025)
EVF-SAM: Early Vision-Language Fusion for Text-Prompted Segment Anything Model
by: Zhang, Yuxuan, et al.
Published: (2024)
by: Zhang, Yuxuan, et al.
Published: (2024)
GroundingSuite: Measuring Complex Multi-Granular Pixel Grounding
by: Hu, Rui, et al.
Published: (2025)
by: Hu, Rui, et al.
Published: (2025)
Mamba Capsule Routing Towards Part-Whole Relational Camouflaged Object Detection
by: Zhang, Dingwen, et al.
Published: (2024)
by: Zhang, Dingwen, et al.
Published: (2024)
CLDA-YOLO: Visual Contrastive Learning Based Domain Adaptive YOLO Detector
by: Qiu, Tianheng, et al.
Published: (2024)
by: Qiu, Tianheng, et al.
Published: (2024)
Similar Items
-
Mask-Adapter: The Devil is in the Masks for Open-Vocabulary Segmentation
by: Li, Yongkang, et al.
Published: (2024) -
Cross-Layer Attentive Feature Upsampling for Low-latency Semantic Segmentation
by: Cheng, Tianheng, et al.
Published: (2026) -
Mamba-YOLO-World: Marrying YOLO-World with Mamba for Open-Vocabulary Detection
by: Wang, Haoxuan, et al.
Published: (2024) -
YOLO-UniOW: Efficient Universal Open-World Object Detection
by: Liu, Lihao, et al.
Published: (2024) -
Polar Parametrization for Vision-based Surround-View 3D Detection
by: Chen, Shaoyu, et al.
Published: (2022)