Visually Consistent Hierarchical Image Classification

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Park, Seulki, Zhang, Youren, Yu, Stella X., Beery, Sara, Huang, Jonathan
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916693375188992
author Park, Seulki
Zhang, Youren
Yu, Stella X.
Beery, Sara
Huang, Jonathan
author_facet Park, Seulki
Zhang, Youren
Yu, Stella X.
Beery, Sara
Huang, Jonathan
contents Hierarchical classification predicts labels across multiple levels of a taxonomy, e.g., from coarse-level 'Bird' to mid-level 'Hummingbird' to fine-level 'Green hermit', allowing flexible recognition under varying visual conditions. It is commonly framed as multiple single-level tasks, but each level may rely on different visual cues: Distinguishing 'Bird' from 'Plant' relies on global features like feathers or leaves, while separating 'Anna's hummingbird' from 'Green hermit' requires local details such as head coloration. Prior methods improve accuracy using external semantic supervision, but such statistical learning criteria fail to ensure consistent visual grounding at test time, resulting in incorrect hierarchical classification. We propose, for the first time, to enforce internal visual consistency by aligning fine-to-coarse predictions through intra-image segmentation. Our method outperforms zero-shot CLIP and state-of-the-art baselines on hierarchical classification benchmarks, achieving both higher accuracy and more consistent predictions. It also improves internal image segmentation without requiring pixel-level annotations.
format Preprint
id arxiv_https___arxiv_org_abs_2406_11608
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Visually Consistent Hierarchical Image Classification
Park, Seulki
Zhang, Youren
Yu, Stella X.
Beery, Sara
Huang, Jonathan
Computer Vision and Pattern Recognition
Hierarchical classification predicts labels across multiple levels of a taxonomy, e.g., from coarse-level 'Bird' to mid-level 'Hummingbird' to fine-level 'Green hermit', allowing flexible recognition under varying visual conditions. It is commonly framed as multiple single-level tasks, but each level may rely on different visual cues: Distinguishing 'Bird' from 'Plant' relies on global features like feathers or leaves, while separating 'Anna's hummingbird' from 'Green hermit' requires local details such as head coloration. Prior methods improve accuracy using external semantic supervision, but such statistical learning criteria fail to ensure consistent visual grounding at test time, resulting in incorrect hierarchical classification. We propose, for the first time, to enforce internal visual consistency by aligning fine-to-coarse predictions through intra-image segmentation. Our method outperforms zero-shot CLIP and state-of-the-art baselines on hierarchical classification benchmarks, achieving both higher accuracy and more consistent predictions. It also improves internal image segmentation without requiring pixel-level annotations.
title Visually Consistent Hierarchical Image Classification
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2406.11608