YOLO-MS: Rethinking Multi-Scale Representation Learning for Real-time Object Detection
Fuente:
arXiv
Saved in:
| Main Authors: | Chen, Yuming, Yuan, Xinbin, Wang, Jiabao, Wu, Ruiqi, Li, Xiang, Hou, Qibin, Cheng, Ming-Ming |
|---|---|
| Format: | Preprint |
| Published: |
2023
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
CrossKD: Cross-Head Knowledge Distillation for Object Detection
by: Wang, Jiabao, et al.
Published: (2023)
by: Wang, Jiabao, et al.
Published: (2023)
Strip R-CNN: Large Strip Convolution for Remote Sensing Object Detection
by: Yuan, Xinbin, et al.
Published: (2025)
by: Yuan, Xinbin, et al.
Published: (2025)
Zone Evaluation: Revealing Spatial Bias in Object Detection
by: Zheng, Zhaohui, et al.
Published: (2023)
by: Zheng, Zhaohui, et al.
Published: (2023)
Towards Stable 3D Object Detection
by: Wang, Jiabao, et al.
Published: (2024)
by: Wang, Jiabao, et al.
Published: (2024)
DFormer: Rethinking RGBD Representation Learning for Semantic Segmentation
by: Yin, Bowen, et al.
Published: (2023)
by: Yin, Bowen, et al.
Published: (2023)
Rethinking RGB-D Salient Object Detection: Models, Data Sets, and Large-Scale Benchmarks
by: Fan, Deng-Ping, et al.
Published: (2019)
by: Fan, Deng-Ping, et al.
Published: (2019)
Referring Camouflaged Object Detection
by: Zhang, Xuying, et al.
Published: (2023)
by: Zhang, Xuying, et al.
Published: (2023)
OmniSegmentor: A Flexible Multi-Modal Learning Framework for Semantic Segmentation
by: Yin, Bo-Wen, et al.
Published: (2025)
by: Yin, Bo-Wen, et al.
Published: (2025)
SM3Det: A Unified Model for Multi-Modal Remote Sensing Object Detection
by: Li, Yuxuan, et al.
Published: (2024)
by: Li, Yuxuan, et al.
Published: (2024)
Re-Aligning Language to Visual Objects with an Agentic Workflow
by: Chen, Yuming, et al.
Published: (2025)
by: Chen, Yuming, et al.
Published: (2025)
MS-YOLO: A Multi-Scale Model for Accurate and Efficient Blood Cell Detection
by: Wu, Guohua, et al.
Published: (2025)
by: Wu, Guohua, et al.
Published: (2025)
Rethinking Token-Level Policy Optimization for Multimodal Chain-of-Thought
by: Li, Yunheng, et al.
Published: (2026)
by: Li, Yunheng, et al.
Published: (2026)
YOLO-World: Real-Time Open-Vocabulary Object Detection
by: Cheng, Tianheng, et al.
Published: (2024)
by: Cheng, Tianheng, et al.
Published: (2024)
SARDet-100K: Towards Open-Source Benchmark and ToolKit for Large-Scale SAR Object Detection
by: Li, Yuxuan, et al.
Published: (2024)
by: Li, Yuxuan, et al.
Published: (2024)
YOLO-IOD: Towards Real Time Incremental Object Detection
by: Zhang, Shizhou, et al.
Published: (2025)
by: Zhang, Shizhou, et al.
Published: (2025)
Revisiting Efficient Semantic Segmentation: Learning Offsets for Better Spatial and Class Feature Alignment
by: Zhang, Shi-Chen, et al.
Published: (2025)
by: Zhang, Shi-Chen, et al.
Published: (2025)
TACR-YOLO: A Real-time Detection Framework for Abnormal Human Behaviors Enhanced with Coordinate and Task-Aware Representations
by: Yin, Xinyi, et al.
Published: (2025)
by: Yin, Xinyi, et al.
Published: (2025)
Unifying Heterogeneous Multi-Modal Remote Sensing Detection Via Language-Pivoted Pretraining
by: Li, Yuxuan, et al.
Published: (2026)
by: Li, Yuxuan, et al.
Published: (2026)
Assessing the Capability of YOLO- and Transformer-based Object Detectors for Real-time Weed Detection
by: Allmendinger, Alicia, et al.
Published: (2025)
by: Allmendinger, Alicia, et al.
Published: (2025)
Predictive Regularization Against Visual Representation Degradation in Multimodal Large Language Models
by: Wang, Enguang, et al.
Published: (2026)
by: Wang, Enguang, et al.
Published: (2026)
OPUS: Occupancy Prediction Using a Sparse Set
by: Wang, Jiabao, et al.
Published: (2024)
by: Wang, Jiabao, et al.
Published: (2024)
MambaNeXt-YOLO: A Hybrid State Space Model for Real-time Object Detection
by: Lei, Xiaochun, et al.
Published: (2025)
by: Lei, Xiaochun, et al.
Published: (2025)
ControlSR: Taming Diffusion Models for Consistent Real-World Image Super Resolution
by: Wan, Yuhao, et al.
Published: (2024)
by: Wan, Yuhao, et al.
Published: (2024)
StoryDiffusion: Consistent Self-Attention for Long-Range Image and Video Generation
by: Zhou, Yupeng, et al.
Published: (2024)
by: Zhou, Yupeng, et al.
Published: (2024)
DFormerv2: Geometry Self-Attention for RGBD Semantic Segmentation
by: Yin, Bo-Wen, et al.
Published: (2025)
by: Yin, Bo-Wen, et al.
Published: (2025)
3A-YOLO: New Real-Time Object Detectors with Triple Discriminative Awareness and Coordinated Representations
by: Wu, Xuecheng, et al.
Published: (2024)
by: Wu, Xuecheng, et al.
Published: (2024)
Cascade-CLIP: Cascaded Vision-Language Embeddings Alignment for Zero-Shot Semantic Segmentation
by: Li, Yunheng, et al.
Published: (2024)
by: Li, Yunheng, et al.
Published: (2024)
Mutual Forcing: Dual-Mode Self-Evolution for Fast Autoregressive Audio-Video Character Generation
by: Zhou, Yupeng, et al.
Published: (2026)
by: Zhou, Yupeng, et al.
Published: (2026)
Multi-Scale Representations by Varying Window Attention for Semantic Segmentation
by: Yan, Haotian, et al.
Published: (2024)
by: Yan, Haotian, et al.
Published: (2024)
KAC: Kolmogorov-Arnold Classifier for Continual Learning
by: Hu, Yusong, et al.
Published: (2025)
by: Hu, Yusong, et al.
Published: (2025)
Concealed Object Detection
by: Fan, Deng-Ping, et al.
Published: (2021)
by: Fan, Deng-Ping, et al.
Published: (2021)
Multi-Token Enhancing for Vision Representation Learning
by: Li, Zhong-Yu, et al.
Published: (2024)
by: Li, Zhong-Yu, et al.
Published: (2024)
AR-1-to-3: Single Image to Consistent 3D Object Generation via Next-View Prediction
by: Zhang, Xuying, et al.
Published: (2025)
by: Zhang, Xuying, et al.
Published: (2025)
Unbiased Region-Language Alignment for Open-Vocabulary Dense Prediction
by: Li, Yunheng, et al.
Published: (2024)
by: Li, Yunheng, et al.
Published: (2024)
Sora Generates Videos with Stunning Geometrical Consistency
by: Li, Xuanyi, et al.
Published: (2024)
by: Li, Xuanyi, et al.
Published: (2024)
GSO-YOLO: Global Stability Optimization YOLO for Construction Site Detection
by: Zhang, Yuming, et al.
Published: (2024)
by: Zhang, Yuming, et al.
Published: (2024)
HopTrack: A Real-time Multi-Object Tracking System for Embedded Devices
by: Li, Xiang, et al.
Published: (2024)
by: Li, Xiang, et al.
Published: (2024)
YOLO26: Key Architectural Enhancements and Performance Benchmarking for Real-Time Object Detection
by: Sapkota, Ranjan, et al.
Published: (2025)
by: Sapkota, Ranjan, et al.
Published: (2025)
TempSamp-R1: Effective Temporal Sampling with Reinforcement Fine-Tuning for Video LLMs
by: Li, Yunheng, et al.
Published: (2025)
by: Li, Yunheng, et al.
Published: (2025)
MS-YOLO: Infrared Object Detection for Edge Deployment via MobileNetV4 and SlideLoss
by: Zhang, Jiali, et al.
Published: (2025)
by: Zhang, Jiali, et al.
Published: (2025)
Similar Items
-
CrossKD: Cross-Head Knowledge Distillation for Object Detection
by: Wang, Jiabao, et al.
Published: (2023) -
Strip R-CNN: Large Strip Convolution for Remote Sensing Object Detection
by: Yuan, Xinbin, et al.
Published: (2025) -
Zone Evaluation: Revealing Spatial Bias in Object Detection
by: Zheng, Zhaohui, et al.
Published: (2023) -
Towards Stable 3D Object Detection
by: Wang, Jiabao, et al.
Published: (2024) -
DFormer: Rethinking RGBD Representation Learning for Semantic Segmentation
by: Yin, Bowen, et al.
Published: (2023)