Tile-Based ViT Inference with Visual-Cluster Priors for Zero-Shot Multi-Species Plant Identification

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Gustineli, Murilo, Miyaguchi, Anthony, Cheung, Adrian, Khattak, Divyansh
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913932366577664
author Gustineli, Murilo
Miyaguchi, Anthony
Cheung, Adrian
Khattak, Divyansh
author_facet Gustineli, Murilo
Miyaguchi, Anthony
Cheung, Adrian
Khattak, Divyansh
contents We describe DS@GT's second-place solution to the PlantCLEF 2025 challenge on multi-species plant identification in vegetation quadrat images. Our pipeline combines (i) a fine-tuned Vision Transformer ViTD2PC24All for patch-level inference, (ii) a 4x4 tiling strategy that aligns patch size with the network's 518x518 receptive field, and (iii) domain-prior adaptation through PaCMAP + K-Means visual clustering and geolocation filtering. Tile predictions are aggregated by majority vote and re-weighted with cluster-specific Bayesian priors, yielding a macro-averaged F1 of 0.348 (private leaderboard) while requiring no additional training. All code, configuration files, and reproducibility scripts are publicly available at https://github.com/dsgt-arc/plantclef-2025.
format Preprint
id arxiv_https___arxiv_org_abs_2507_06093
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Tile-Based ViT Inference with Visual-Cluster Priors for Zero-Shot Multi-Species Plant Identification
Gustineli, Murilo
Miyaguchi, Anthony
Cheung, Adrian
Khattak, Divyansh
Computer Vision and Pattern Recognition
Information Retrieval
Machine Learning
We describe DS@GT's second-place solution to the PlantCLEF 2025 challenge on multi-species plant identification in vegetation quadrat images. Our pipeline combines (i) a fine-tuned Vision Transformer ViTD2PC24All for patch-level inference, (ii) a 4x4 tiling strategy that aligns patch size with the network's 518x518 receptive field, and (iii) domain-prior adaptation through PaCMAP + K-Means visual clustering and geolocation filtering. Tile predictions are aggregated by majority vote and re-weighted with cluster-specific Bayesian priors, yielding a macro-averaged F1 of 0.348 (private leaderboard) while requiring no additional training. All code, configuration files, and reproducibility scripts are publicly available at https://github.com/dsgt-arc/plantclef-2025.
title Tile-Based ViT Inference with Visual-Cluster Priors for Zero-Shot Multi-Species Plant Identification
topic Computer Vision and Pattern Recognition
Information Retrieval
Machine Learning
url https://arxiv.org/abs/2507.06093