Tile-Based ViT Inference with Visual-Cluster Priors for Zero-Shot Multi-Species Plant Identification
Fuente:
arXiv
Saved in:
| Main Authors: | Gustineli, Murilo, Miyaguchi, Anthony, Cheung, Adrian, Khattak, Divyansh |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Multi-Label Plant Species Classification with Self-Supervised Vision Transformers
by: Gustineli, Murilo, et al.
Published: (2024)
by: Gustineli, Murilo, et al.
Published: (2024)
Transfer Learning and Mixup for Fine-Grained Few-Shot Fungi Classification
by: Tam, Jason Kahei, et al.
Published: (2025)
by: Tam, Jason Kahei, et al.
Published: (2025)
Transfer Learning with Self-Supervised Vision Transformers for Snake Identification
by: Miyaguchi, Anthony, et al.
Published: (2024)
by: Miyaguchi, Anthony, et al.
Published: (2024)
Transfer Learning with Semi-Supervised Dataset Annotation for Birdcall Classification
by: Miyaguchi, Anthony, et al.
Published: (2023)
by: Miyaguchi, Anthony, et al.
Published: (2023)
Distilling Spectrograms into Tokens: Fast and Lightweight Bioacoustic Classification for BirdCLEF+ 2025
by: Miyaguchi, Anthony, et al.
Published: (2025)
by: Miyaguchi, Anthony, et al.
Published: (2025)
Transfer Learning with Pseudo Multi-Label Birdcall Classification for DS@GT BirdCLEF 2024
by: Miyaguchi, Anthony, et al.
Published: (2024)
by: Miyaguchi, Anthony, et al.
Published: (2024)
Visual Zero-Shot E-Commerce Product Attribute Value Extraction
by: Gong, Jiaying, et al.
Published: (2025)
by: Gong, Jiaying, et al.
Published: (2025)
DS@GT at LongEval: Evaluating Temporal Performance in Web Search Systems and Topics with Two-Stage Retrieval
by: Miyaguchi, Anthony, et al.
Published: (2025)
by: Miyaguchi, Anthony, et al.
Published: (2025)
Zero-Shot Hashing Based on Reconstruction With Part Alignment
by: Jiang, Yan, et al.
Published: (2025)
by: Jiang, Yan, et al.
Published: (2025)
MLLM-Guided VLM Fine-Tuning with Joint Inference for Zero-Shot Composed Image Retrieval
by: Tu, Rong-Cheng, et al.
Published: (2025)
by: Tu, Rong-Cheng, et al.
Published: (2025)
DS@GT at Touché: Large Language Models for Retrieval-Augmented Debate
by: Miyaguchi, Anthony, et al.
Published: (2025)
by: Miyaguchi, Anthony, et al.
Published: (2025)
DS@GT eRisk 2024: Sentence Transformers for Social Media Risk Assessment
by: Guecha, David, et al.
Published: (2024)
by: Guecha, David, et al.
Published: (2024)
Multi-Label Plant Species Prediction with Metadata-Enhanced Multi-Head Vision Transformers
by: Herasimchyk, Hanna, et al.
Published: (2025)
by: Herasimchyk, Hanna, et al.
Published: (2025)
ViDR: Grounding Multimodal Deep Research Reports in Source Visual Evidence
by: Shi, Zhuofan, et al.
Published: (2026)
by: Shi, Zhuofan, et al.
Published: (2026)
CoTMR: Chain-of-Thought Multi-Scale Reasoning for Training-Free Zero-Shot Composed Image Retrieval
by: Sun, Zelong, et al.
Published: (2025)
by: Sun, Zelong, et al.
Published: (2025)
Multimodal Reasoning Agent for Zero-Shot Composed Image Retrieval
by: Tu, Rong-Cheng, et al.
Published: (2025)
by: Tu, Rong-Cheng, et al.
Published: (2025)
Vote-in-Context: Turning VLMs into Zero-Shot Rank Fusers
by: Eltahir, Mohamed, et al.
Published: (2025)
by: Eltahir, Mohamed, et al.
Published: (2025)
iSEARLE: Improving Textual Inversion for Zero-Shot Composed Image Retrieval
by: Agnolucci, Lorenzo, et al.
Published: (2024)
by: Agnolucci, Lorenzo, et al.
Published: (2024)
DS@GT at TREC TOT 2025: Bridging Vague Recollection with Fusion Retrieval and Learned Reranking
by: Zhou, Wenxin, et al.
Published: (2026)
by: Zhou, Wenxin, et al.
Published: (2026)
Multi-Label Zero-Shot Product Attribute-Value Extraction
by: Gong, Jiaying, et al.
Published: (2024)
by: Gong, Jiaying, et al.
Published: (2024)
UNION: A Lightweight Target Representation for Efficient Zero-Shot Image-Guided Retrieval with Optional Textual Queries
by: Le, Hoang-Bao, et al.
Published: (2025)
by: Le, Hoang-Bao, et al.
Published: (2025)
SERVAL: Surprisingly Effective Zero-Shot Visual Document Retrieval Powered by Large Vision and Language Models
by: Nguyen, Thong, et al.
Published: (2025)
by: Nguyen, Thong, et al.
Published: (2025)
Retrieval Augmented Zero-Shot Text Classification
by: Abdullahi, Tassallah, et al.
Published: (2024)
by: Abdullahi, Tassallah, et al.
Published: (2024)
ZeroGR: A Generalizable and Scalable Framework for Zero-Shot Generative Retrieval
by: Sun, Weiwei, et al.
Published: (2025)
by: Sun, Weiwei, et al.
Published: (2025)
ViDoRe Benchmark V2: Raising the Bar for Visual Retrieval
by: Macé, Quentin, et al.
Published: (2025)
by: Macé, Quentin, et al.
Published: (2025)
Zero-Shot Retrieval for Scalable Visual Search in a Two-Sided Marketplace
by: Rusli, Andre, et al.
Published: (2025)
by: Rusli, Andre, et al.
Published: (2025)
Zero-Shot Complex Question-Answering on Long Scientific Documents
by: Wang, Wanting
Published: (2025)
by: Wang, Wanting
Published: (2025)
A Multi-Granularity Retrieval Framework for Visually-Rich Documents
by: Xu, Mingjun, et al.
Published: (2025)
by: Xu, Mingjun, et al.
Published: (2025)
GMM-Based Comprehensive Feature Extraction and Relative Distance Preservation For Few-Shot Cross-Modal Retrieval
by: Sun, Chengsong, et al.
Published: (2025)
by: Sun, Chengsong, et al.
Published: (2025)
ViFi-ReID: A Two-Stream Vision-WiFi Multimodal Approach for Person Re-identification
by: Mao, Chen, et al.
Published: (2024)
by: Mao, Chen, et al.
Published: (2024)
Zero-Shot Image Moderation in Google Ads with LLM-Assisted Textual Descriptions and Cross-modal Co-embeddings
by: Luo, Enming, et al.
Published: (2024)
by: Luo, Enming, et al.
Published: (2024)
RecBase: Generative Foundation Model Pretraining for Zero-Shot Recommendation
by: Zhou, Sashuai, et al.
Published: (2025)
by: Zhou, Sashuai, et al.
Published: (2025)
Identity-Decoupled Anonymization for Visual Evidence in Multi-modal Retrieval-Augmented Generation
by: Cheng, Zehua, et al.
Published: (2026)
by: Cheng, Zehua, et al.
Published: (2026)
A Cooperative Multi-Agent Framework for Zero-Shot Named Entity Recognition
by: Wang, Zihan, et al.
Published: (2025)
by: Wang, Zihan, et al.
Published: (2025)
Leveraging Reference Documents for Zero-Shot Ranking via Large Language Models
by: Li, Jieran, et al.
Published: (2025)
by: Li, Jieran, et al.
Published: (2025)
Col-Bandit: Zero-Shot Query-Time Pruning for Late-Interaction Retrieval
by: Pony, Roi, et al.
Published: (2026)
by: Pony, Roi, et al.
Published: (2026)
Magic Tokens: Select Diverse Tokens for Multi-modal Object Re-Identification
by: Zhang, Pingping, et al.
Published: (2024)
by: Zhang, Pingping, et al.
Published: (2024)
Towards Human-Like Machine Comprehension: Few-Shot Relational Learning in Visually-Rich Documents
by: Wang, Hao, et al.
Published: (2024)
by: Wang, Hao, et al.
Published: (2024)
Rethinking Detection Based Table Structure Recognition for Visually Rich Document Images
by: Xiao, Bin, et al.
Published: (2023)
by: Xiao, Bin, et al.
Published: (2023)
ViDoRAG: Visual Document Retrieval-Augmented Generation via Dynamic Iterative Reasoning Agents
by: Wang, Qiuchen, et al.
Published: (2025)
by: Wang, Qiuchen, et al.
Published: (2025)
Similar Items
-
Multi-Label Plant Species Classification with Self-Supervised Vision Transformers
by: Gustineli, Murilo, et al.
Published: (2024) -
Transfer Learning and Mixup for Fine-Grained Few-Shot Fungi Classification
by: Tam, Jason Kahei, et al.
Published: (2025) -
Transfer Learning with Self-Supervised Vision Transformers for Snake Identification
by: Miyaguchi, Anthony, et al.
Published: (2024) -
Transfer Learning with Semi-Supervised Dataset Annotation for Birdcall Classification
by: Miyaguchi, Anthony, et al.
Published: (2023) -
Distilling Spectrograms into Tokens: Fast and Lightweight Bioacoustic Classification for BirdCLEF+ 2025
by: Miyaguchi, Anthony, et al.
Published: (2025)