SliceLens: Fine-Grained and Grounded Error Slice Discovery for Multi-Instance Vision Tasks

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhang, Wei, Wang, Chaoqun, Guan, Zixuan, Kao, Sam, Zhao, Pengfei, Wu, Peng, He, Sifeng
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911347576406016
author Zhang, Wei
Wang, Chaoqun
Guan, Zixuan
Kao, Sam
Zhao, Pengfei
Wu, Peng
He, Sifeng
author_facet Zhang, Wei
Wang, Chaoqun
Guan, Zixuan
Kao, Sam
Zhao, Pengfei
Wu, Peng
He, Sifeng
contents Systematic failures of computer vision models on subsets with coherent visual patterns, known as error slices, pose a critical challenge for robust model evaluation. Existing slice discovery methods are primarily developed for image classification, limiting their applicability to multi-instance tasks such as detection, segmentation, and pose estimation. In real-world scenarios, error slices often arise from corner cases involving complex visual relationships, where existing instance-level approaches lacking fine-grained reasoning struggle to yield meaningful insights. Moreover, current benchmarks are typically tailored to specific algorithms or biased toward image classification, with artificial ground truth that fails to reflect real model failures. To address these limitations, we propose SliceLens, a hypothesis-driven framework that leverages LLMs and VLMs to generate and verify diverse failure hypotheses through grounded visual reasoning, enabling reliable identification of fine-grained and interpretable error slices. We further introduce FeSD (Fine-grained Slice Discovery), the first benchmark specifically designed for evaluating fine-grained error slice discovery across instance-level vision tasks, featuring expert-annotated and carefully refined ground-truth slices with precise grounding to local error regions. Extensive experiments on both existing benchmarks and FeSD demonstrate that SliceLens achieves state-of-the-art performance, improving Precision@10 by 0.42 (0.73 vs. 0.31) on FeSD, and identifies interpretable slices that facilitate actionable model improvements, as validated through model repair experiments.
format Preprint
id arxiv_https___arxiv_org_abs_2512_24592
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle SliceLens: Fine-Grained and Grounded Error Slice Discovery for Multi-Instance Vision Tasks
Zhang, Wei
Wang, Chaoqun
Guan, Zixuan
Kao, Sam
Zhao, Pengfei
Wu, Peng
He, Sifeng
Computer Vision and Pattern Recognition
Systematic failures of computer vision models on subsets with coherent visual patterns, known as error slices, pose a critical challenge for robust model evaluation. Existing slice discovery methods are primarily developed for image classification, limiting their applicability to multi-instance tasks such as detection, segmentation, and pose estimation. In real-world scenarios, error slices often arise from corner cases involving complex visual relationships, where existing instance-level approaches lacking fine-grained reasoning struggle to yield meaningful insights. Moreover, current benchmarks are typically tailored to specific algorithms or biased toward image classification, with artificial ground truth that fails to reflect real model failures. To address these limitations, we propose SliceLens, a hypothesis-driven framework that leverages LLMs and VLMs to generate and verify diverse failure hypotheses through grounded visual reasoning, enabling reliable identification of fine-grained and interpretable error slices. We further introduce FeSD (Fine-grained Slice Discovery), the first benchmark specifically designed for evaluating fine-grained error slice discovery across instance-level vision tasks, featuring expert-annotated and carefully refined ground-truth slices with precise grounding to local error regions. Extensive experiments on both existing benchmarks and FeSD demonstrate that SliceLens achieves state-of-the-art performance, improving Precision@10 by 0.42 (0.73 vs. 0.31) on FeSD, and identifies interpretable slices that facilitate actionable model improvements, as validated through model repair experiments.
title SliceLens: Fine-Grained and Grounded Error Slice Discovery for Multi-Instance Vision Tasks
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2512.24592