What You See is (Usually) What You Get: Multimodal Prototype Networks that Abstain from Expensive Modalities

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Bahng, Muchang, Berens, Charlie, Donnelly, Jon, Chen, Eric, Chen, Chaofan, Rudin, Cynthia
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917102071316480
author Bahng, Muchang
Berens, Charlie
Donnelly, Jon
Chen, Eric
Chen, Chaofan
Rudin, Cynthia
author_facet Bahng, Muchang
Berens, Charlie
Donnelly, Jon
Chen, Eric
Chen, Chaofan
Rudin, Cynthia
contents Species detection is important for monitoring the health of ecosystems and identifying invasive species, serving a crucial role in guiding conservation efforts. Multimodal neural networks have seen increasing use for identifying species to help automate this task, but they have two major drawbacks. First, their black-box nature prevents the interpretability of their decision making process. Second, collecting genetic data is often expensive and requires invasive procedures, often necessitating researchers to capture or kill the target specimen. We address both of these problems by extending prototype networks (ProtoPNets), which are a popular and interpretable alternative to traditional neural networks, to the multimodal, cost-aware setting. We ensemble prototypes from each modality, using an associated weight to determine how much a given prediction relies on each modality. We further introduce methods to identify cases for which we do not need the expensive genetic information to make confident predictions. We demonstrate that our approach can intelligently allocate expensive genetic data for fine-grained distinctions while using abundant image data for clearer visual classifications and achieving comparable accuracy to models that consistently use both modalities.
format Preprint
id arxiv_https___arxiv_org_abs_2511_19752
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle What You See is (Usually) What You Get: Multimodal Prototype Networks that Abstain from Expensive Modalities
Bahng, Muchang
Berens, Charlie
Donnelly, Jon
Chen, Eric
Chen, Chaofan
Rudin, Cynthia
Computer Vision and Pattern Recognition
Species detection is important for monitoring the health of ecosystems and identifying invasive species, serving a crucial role in guiding conservation efforts. Multimodal neural networks have seen increasing use for identifying species to help automate this task, but they have two major drawbacks. First, their black-box nature prevents the interpretability of their decision making process. Second, collecting genetic data is often expensive and requires invasive procedures, often necessitating researchers to capture or kill the target specimen. We address both of these problems by extending prototype networks (ProtoPNets), which are a popular and interpretable alternative to traditional neural networks, to the multimodal, cost-aware setting. We ensemble prototypes from each modality, using an associated weight to determine how much a given prediction relies on each modality. We further introduce methods to identify cases for which we do not need the expensive genetic information to make confident predictions. We demonstrate that our approach can intelligently allocate expensive genetic data for fine-grained distinctions while using abundant image data for clearer visual classifications and achieving comparable accuracy to models that consistently use both modalities.
title What You See is (Usually) What You Get: Multimodal Prototype Networks that Abstain from Expensive Modalities
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2511.19752