RoCOCO: Robustness Benchmark of MS-COCO to Stress-test Image-Text Matching Models
Fuente:
arXiv
Saved in:
| Main Authors: | Park, Seulki, Um, Daeho, Yoon, Hajung, Chun, Sanghyuk, Yun, Sangdoo, Choi, Jin Young |
|---|---|
| Format: | Preprint |
| Published: |
2023
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
ECCV Caption: Correcting False Negatives by Collecting Machine-and-Human-verified Image-Caption Associations for MS-COCO
by: Chun, Sanghyuk, et al.
Published: (2022)
by: Chun, Sanghyuk, et al.
Published: (2022)
LongProLIP: A Probabilistic Vision-Language Model with Long Context Text
by: Chun, Sanghyuk, et al.
Published: (2025)
by: Chun, Sanghyuk, et al.
Published: (2025)
Probabilistic Language-Image Pre-Training
by: Chun, Sanghyuk, et al.
Published: (2024)
by: Chun, Sanghyuk, et al.
Published: (2024)
Emergence of Text Readability in Vision Language Models
by: Park, Jaeyoo, et al.
Published: (2025)
by: Park, Jaeyoo, et al.
Published: (2025)
HYPE: Hyperbolic Entailment Filtering for Underspecified Images and Texts
by: Kim, Wonjae, et al.
Published: (2024)
by: Kim, Wonjae, et al.
Published: (2024)
3D-COCO: extension of MS-COCO dataset for image detection and 3D reconstruction modules
by: Bideaux, Maxence, et al.
Published: (2024)
by: Bideaux, Maxence, et al.
Published: (2024)
From COCO to COCO-FP: A Deep Dive into Background False Positives for COCO Detectors
by: Liu, Longfei, et al.
Published: (2024)
by: Liu, Longfei, et al.
Published: (2024)
Toward Interactive Regional Understanding in Vision-Large Language Models
by: Lee, Jungbeom, et al.
Published: (2024)
by: Lee, Jungbeom, et al.
Published: (2024)
COCO-OLAC: A Benchmark for Occluded Panoptic Segmentation and Image Understanding
by: Wei, Wenbo, et al.
Published: (2024)
by: Wei, Wenbo, et al.
Published: (2024)
Improved Probabilistic Image-Text Representations
by: Chun, Sanghyuk
Published: (2023)
by: Chun, Sanghyuk
Published: (2023)
Finding Optimal Video Moment without Training: Gaussian Boundary Optimization for Weakly Supervised Video Grounding
by: Kim, Sunoh, et al.
Published: (2026)
by: Kim, Sunoh, et al.
Published: (2026)
COCONut: Modernizing COCO Segmentation
by: Deng, Xueqing, et al.
Published: (2024)
by: Deng, Xueqing, et al.
Published: (2024)
Language-only Efficient Training of Zero-shot Composed Image Retrieval
by: Gu, Geonmo, et al.
Published: (2023)
by: Gu, Geonmo, et al.
Published: (2023)
UOD: Unseen Object Detection in 3D Point Cloud
by: Choi, Hyunjun, et al.
Published: (2024)
by: Choi, Hyunjun, et al.
Published: (2024)
Enhancing Weakly Supervised Video Grounding via Diverse Inference Strategies for Boundary and Prediction Selection
by: Kim, Sunoh, et al.
Published: (2025)
by: Kim, Sunoh, et al.
Published: (2025)
Benchmarking Object Detectors with COCO: A New Path Forward
by: Singh, Shweta, et al.
Published: (2024)
by: Singh, Shweta, et al.
Published: (2024)
COCO-Inpaint: A Benchmark for Detecting and Localizing Inpainting-Based Image Manipulations
by: Yan, Haozhen, et al.
Published: (2025)
by: Yan, Haozhen, et al.
Published: (2025)
Leveraging Generative AI Models to Explore Human Identity
by: Yeo, Yunha, et al.
Published: (2025)
by: Yeo, Yunha, et al.
Published: (2025)
CompoDiff: Versatile Composed Image Retrieval With Latent Diffusion
by: Gu, Geonmo, et al.
Published: (2023)
by: Gu, Geonmo, et al.
Published: (2023)
Stable Diffusion for Data Augmentation in COCO and Weed Datasets
by: Deng, Boyang
Published: (2023)
by: Deng, Boyang
Published: (2023)
COCO is "ALL'' You Need for Visual Instruction Fine-tuning
by: Han, Xiaotian, et al.
Published: (2024)
by: Han, Xiaotian, et al.
Published: (2024)
Direct Unlearning Optimization for Robust and Safe Text-to-Image Models
by: Park, Yong-Hyun, et al.
Published: (2024)
by: Park, Yong-Hyun, et al.
Published: (2024)
Learning Feature Inversion for Multi-class Anomaly Detection under General-purpose COCO-AD Benchmark
by: Zhang, Jiangning, et al.
Published: (2024)
by: Zhang, Jiangning, et al.
Published: (2024)
Fine-Tuning Without Forgetting: Adaptation of YOLOv8 Preserves COCO Performance
by: Gandhi, Vishal, et al.
Published: (2025)
by: Gandhi, Vishal, et al.
Published: (2025)
COCO-Tree: Compositional Hierarchical Concept Trees for Enhanced Reasoning in Vision Language Models
by: Sinha, Sanchit, et al.
Published: (2025)
by: Sinha, Sanchit, et al.
Published: (2025)
COCO-Urdu: A Large-Scale Urdu Image-Caption Dataset with Multimodal Quality Estimation
by: Hassan, Umair
Published: (2025)
by: Hassan, Umair
Published: (2025)
Multiplicity is an Inevitable and Inherent Challenge in Multimodal Learning
by: Chun, Sanghyuk
Published: (2025)
by: Chun, Sanghyuk
Published: (2025)
Mitigating Cross-Image Information Leakage in LVLMs for Multi-Image Tasks
by: Park, Yeji, et al.
Published: (2025)
by: Park, Yeji, et al.
Published: (2025)
Pix2Cap-COCO: Advancing Visual Comprehension via Pixel-Level Captioning
by: You, Zuyao, et al.
Published: (2025)
by: You, Zuyao, et al.
Published: (2025)
COCO Pose val2017
by: Moghimi, Sina
Published: (2025)
by: Moghimi, Sina
Published: (2025)
A COCO-Formatted Instance-Level Dataset for Plasmodium Falciparum Detection in Giemsa-Stained Blood Smears
by: Wilm, Frauke, et al.
Published: (2025)
by: Wilm, Frauke, et al.
Published: (2025)
Can AI Recognize the Style of Art? Analyzing Aesthetics through the Lens of Style Transfer
by: Yeo, Yunha, et al.
Published: (2025)
by: Yeo, Yunha, et al.
Published: (2025)
Match me if you can: Semi-Supervised Semantic Correspondence Learning with Unpaired Images
by: Kim, Jiwon, et al.
Published: (2023)
by: Kim, Jiwon, et al.
Published: (2023)
Pose Matters: Evaluating Vision Transformers and CNNs for Human Action Recognition on Small COCO Subsets
by: Tang, MingZe, et al.
Published: (2025)
by: Tang, MingZe, et al.
Published: (2025)
RoMa: Robust Dense Feature Matching
by: Edstedt, Johan, et al.
Published: (2023)
by: Edstedt, Johan, et al.
Published: (2023)
ALEGORÍAS CON COCO Y ANÍS
by: María Clara Escobar Gaitán
Published: (2009)
by: María Clara Escobar Gaitán
Published: (2009)
An Efficient Post-hoc Framework for Reducing Task Discrepancy of Text Encoders for Composed Image Retrieval
by: Byun, Jaeseok, et al.
Published: (2024)
by: Byun, Jaeseok, et al.
Published: (2024)
Seeing What You Say: Expressive Image Generation from Speech
by: Lee, Jiyoung, et al.
Published: (2025)
by: Lee, Jiyoung, et al.
Published: (2025)
Steering Guidance for Personalized Text-to-Image Diffusion Models
by: Park, Sunghyun, et al.
Published: (2025)
by: Park, Sunghyun, et al.
Published: (2025)
RoSe: Robust Self-supervised Stereo Matching under Adverse Weather Conditions
by: Wang, Yun, et al.
Published: (2025)
by: Wang, Yun, et al.
Published: (2025)
Similar Items
-
ECCV Caption: Correcting False Negatives by Collecting Machine-and-Human-verified Image-Caption Associations for MS-COCO
by: Chun, Sanghyuk, et al.
Published: (2022) -
LongProLIP: A Probabilistic Vision-Language Model with Long Context Text
by: Chun, Sanghyuk, et al.
Published: (2025) -
Probabilistic Language-Image Pre-Training
by: Chun, Sanghyuk, et al.
Published: (2024) -
Emergence of Text Readability in Vision Language Models
by: Park, Jaeyoo, et al.
Published: (2025) -
HYPE: Hyperbolic Entailment Filtering for Underspecified Images and Texts
by: Kim, Wonjae, et al.
Published: (2024)