Benchmarking Object Detectors with COCO: A New Path Forward
Fuente:
arXiv
Saved in:
| Main Authors: | Singh, Shweta, Yadav, Aayan, Jain, Jitesh, Shi, Humphrey, Johnson, Justin, Desai, Karan |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Elevating Visual Perception in Multimodal LLMs with Visual Embedding Distillation
by: Jain, Jitesh, et al.
Published: (2024)
by: Jain, Jitesh, et al.
Published: (2024)
StegaVision: Enhancing Steganography with Attention Mechanism
by: Kumar, Abhinav, et al.
Published: (2024)
by: Kumar, Abhinav, et al.
Published: (2024)
From COCO to COCO-FP: A Deep Dive into Background False Positives for COCO Detectors
by: Liu, Longfei, et al.
Published: (2024)
by: Liu, Longfei, et al.
Published: (2024)
Impact of Language Guidance: A Reproducibility Study
by: Puniani, Cherish, et al.
Published: (2025)
by: Puniani, Cherish, et al.
Published: (2025)
Provenance Detection for AI-Generated Images: Combining Perceptual Hashing, Homomorphic Encryption, and AI Detection Models
by: Singhi, Shree, et al.
Published: (2025)
by: Singhi, Shree, et al.
Published: (2025)
Slow-Fast Architecture for Video Multi-Modal Large Language Models
by: Shi, Min, et al.
Published: (2025)
by: Shi, Min, et al.
Published: (2025)
CuMo: Scaling Multimodal LLM with Co-Upcycled Mixture-of-Experts
by: Li, Jiachen, et al.
Published: (2024)
by: Li, Jiachen, et al.
Published: (2024)
Hyperbolic Image-Text Representations
by: Desai, Karan, et al.
Published: (2023)
by: Desai, Karan, et al.
Published: (2023)
SAGE: Training Smart Any-Horizon Agents for Long Video Reasoning with Reinforcement Learning
by: Jain, Jitesh, et al.
Published: (2025)
by: Jain, Jitesh, et al.
Published: (2025)
MCUBench: A Benchmark of Tiny Object Detectors on MCUs
by: Sah, Sudhakar, et al.
Published: (2024)
by: Sah, Sudhakar, et al.
Published: (2024)
COCO-OLAC: A Benchmark for Occluded Panoptic Segmentation and Image Understanding
by: Wei, Wenbo, et al.
Published: (2024)
by: Wei, Wenbo, et al.
Published: (2024)
RoCOCO: Robustness Benchmark of MS-COCO to Stress-test Image-Text Matching Models
by: Park, Seulki, et al.
Published: (2023)
by: Park, Seulki, et al.
Published: (2023)
PAI-Bench: A Comprehensive Benchmark For Physical AI
by: Zhou, Fengzhe, et al.
Published: (2025)
by: Zhou, Fengzhe, et al.
Published: (2025)
Balancing Stability and Plasticity in Pretrained Detector: A Dual-Path Framework for Incremental Object Detection
by: Li, Songze, et al.
Published: (2025)
by: Li, Songze, et al.
Published: (2025)
Plain-Det: A Plain Multi-Dataset Object Detector
by: Shi, Cheng, et al.
Published: (2024)
by: Shi, Cheng, et al.
Published: (2024)
COCONut: Modernizing COCO Segmentation
by: Deng, Xueqing, et al.
Published: (2024)
by: Deng, Xueqing, et al.
Published: (2024)
Comparison Of Deep Object Detectors On A New Vulnerable Pedestrian Dataset
by: Sharma, Devansh, et al.
Published: (2022)
by: Sharma, Devansh, et al.
Published: (2022)
3D-COCO: extension of MS-COCO dataset for image detection and 3D reconstruction modules
by: Bideaux, Maxence, et al.
Published: (2024)
by: Bideaux, Maxence, et al.
Published: (2024)
Benchmarking Object Detectors under Real-World Distribution Shifts in Satellite Imagery
by: Al-Emadi, Sara, et al.
Published: (2025)
by: Al-Emadi, Sara, et al.
Published: (2025)
Benchmarking CNN and Transformer-Based Object Detectors for UAV Solar Panel Inspection
by: Rodrigo, Ashen, et al.
Published: (2025)
by: Rodrigo, Ashen, et al.
Published: (2025)
FactorizePhys: Matrix Factorization for Multidimensional Attention in Remote Physiological Sensing
by: Joshi, Jitesh, et al.
Published: (2024)
by: Joshi, Jitesh, et al.
Published: (2024)
COCO-Inpaint: A Benchmark for Detecting and Localizing Inpainting-Based Image Manipulations
by: Yan, Haozhen, et al.
Published: (2025)
by: Yan, Haozhen, et al.
Published: (2025)
Learning Feature Inversion for Multi-class Anomaly Detection under General-purpose COCO-AD Benchmark
by: Zhang, Jiangning, et al.
Published: (2024)
by: Zhang, Jiangning, et al.
Published: (2024)
Efficient and Robust Multidimensional Attention in Remote Physiological Sensing through Target Signal Constrained Factorization
by: Joshi, Jitesh, et al.
Published: (2025)
by: Joshi, Jitesh, et al.
Published: (2025)
PathVG: A New Benchmark and Dataset for Pathology Visual Grounding
by: Zhong, Chunlin, et al.
Published: (2025)
by: Zhong, Chunlin, et al.
Published: (2025)
Stable Diffusion for Data Augmentation in COCO and Weed Datasets
by: Deng, Boyang
Published: (2023)
by: Deng, Boyang
Published: (2023)
Object Detection using Event Camera: A MoE Heat Conduction based Detector and A New Benchmark Dataset
by: Wang, Xiao, et al.
Published: (2024)
by: Wang, Xiao, et al.
Published: (2024)
Beyond Realism: Learning the Art of Expressive Composition with StickerNet
by: Lu, Haoming, et al.
Published: (2025)
by: Lu, Haoming, et al.
Published: (2025)
Task Integration Distillation for Object Detectors
by: Su, Hai, et al.
Published: (2024)
by: Su, Hai, et al.
Published: (2024)
Diagnosing Visual Reasoning: Challenges, Insights, and a Path Forward
by: Bi, Jing, et al.
Published: (2025)
by: Bi, Jing, et al.
Published: (2025)
A New People-Object Interaction Dataset and NVS Benchmarks
by: Guo, Shuai, et al.
Published: (2024)
by: Guo, Shuai, et al.
Published: (2024)
Motion State: A New Benchmark Multiple Object Tracking
by: Feng, Yang, et al.
Published: (2023)
by: Feng, Yang, et al.
Published: (2023)
VASE: Object-Centric Appearance and Shape Manipulation of Real Videos
by: Peruzzo, Elia, et al.
Published: (2024)
by: Peruzzo, Elia, et al.
Published: (2024)
3A-YOLO: New Real-Time Object Detectors with Triple Discriminative Awareness and Coordinated Representations
by: Wu, Xuecheng, et al.
Published: (2024)
by: Wu, Xuecheng, et al.
Published: (2024)
Bringing together invertible UNets with invertible attention modules for memory-efficient diffusion models
by: Jain, Karan, et al.
Published: (2025)
by: Jain, Karan, et al.
Published: (2025)
Arcee: Differentiable Recurrent State Chain for Generative Vision Modeling with Mamba SSMs
by: Chavan, Jitesh, et al.
Published: (2025)
by: Chavan, Jitesh, et al.
Published: (2025)
COCO is "ALL'' You Need for Visual Instruction Fine-tuning
by: Han, Xiaotian, et al.
Published: (2024)
by: Han, Xiaotian, et al.
Published: (2024)
WavShadow: Wavelet Based Shadow Segmentation and Removal
by: Jain, Shreyans, et al.
Published: (2024)
by: Jain, Shreyans, et al.
Published: (2024)
Benchmarking Scientific Image Forgery Detectors
by: Cardenuto, João P., et al.
Published: (2021)
by: Cardenuto, João P., et al.
Published: (2021)
Combating Label Noise With A General Surrogate Model For Sample Selection
by: Liang, Chao, et al.
Published: (2023)
by: Liang, Chao, et al.
Published: (2023)
Similar Items
-
Elevating Visual Perception in Multimodal LLMs with Visual Embedding Distillation
by: Jain, Jitesh, et al.
Published: (2024) -
StegaVision: Enhancing Steganography with Attention Mechanism
by: Kumar, Abhinav, et al.
Published: (2024) -
From COCO to COCO-FP: A Deep Dive into Background False Positives for COCO Detectors
by: Liu, Longfei, et al.
Published: (2024) -
Impact of Language Guidance: A Reproducibility Study
by: Puniani, Cherish, et al.
Published: (2025) -
Provenance Detection for AI-Generated Images: Combining Perceptual Hashing, Homomorphic Encryption, and AI Detection Models
by: Singhi, Shree, et al.
Published: (2025)