Vision Transformers for Zero-Shot Clustering of Animal Images: A Comparative Benchmarking Study
Fuente:
arXiv
Guardado en:
| Autores principales: | Markoff, Hugo, Bengtson, Stefan Hein, Ørsted, Michael |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Zero-Shot Wildlife Sorting Using Vision Transformers: Evaluating Clustering and Continuous Similarity Ordering
por: Markoff, Hugo, et al.
Publicado: (2025)
por: Markoff, Hugo, et al.
Publicado: (2025)
Hierarchical Re-Classification: Combining Animal Classification Models with Vision Transformers
por: Markoff, Hugo, et al.
Publicado: (2025)
por: Markoff, Hugo, et al.
Publicado: (2025)
Benchmarking Zero-Shot Recognition with Vision-Language Models: Challenges on Granularity and Specificity
por: Xu, Zhenlin, et al.
Publicado: (2023)
por: Xu, Zhenlin, et al.
Publicado: (2023)
Binary Verification for Zero-Shot Vision
por: Hu, Rongbin, et al.
Publicado: (2025)
por: Hu, Rongbin, et al.
Publicado: (2025)
Sea-ing Through Scattered Rays: Revisiting the Image Formation Model for Realistic Underwater Image Generation
por: Ismiroglou, Vasiliki, et al.
Publicado: (2025)
por: Ismiroglou, Vasiliki, et al.
Publicado: (2025)
VLAgeBench: Benchmarking Large Vision-Language Models for Zero-Shot Human Age Estimation
por: Sajib, Rakib Hossain, et al.
Publicado: (2026)
por: Sajib, Rakib Hossain, et al.
Publicado: (2026)
Benchmarking Foundation Models for Zero-Shot Biometric Tasks
por: Sony, Redwan, et al.
Publicado: (2025)
por: Sony, Redwan, et al.
Publicado: (2025)
Rethinking Plant Disease Diagnosis: Bridging the Academic-Practical Gap with Vision Transformers and Zero-Shot Learning
por: Benabbas, Wassim, et al.
Publicado: (2025)
por: Benabbas, Wassim, et al.
Publicado: (2025)
LesionLocator: Zero-Shot Universal Tumor Segmentation and Tracking in 3D Whole-Body Imaging
por: Rokuss, Maximilian, et al.
Publicado: (2025)
por: Rokuss, Maximilian, et al.
Publicado: (2025)
HeatPrompt: Zero-Shot Vision-Language Modeling of Urban Heat Demand from Satellite Images
por: Thota, Kundan, et al.
Publicado: (2026)
por: Thota, Kundan, et al.
Publicado: (2026)
Benchmarking Unlearning for Vision Transformers
por: Zhao, Kairan, et al.
Publicado: (2026)
por: Zhao, Kairan, et al.
Publicado: (2026)
Efficient Zero-Shot AI-Generated Image Detection
por: Sonoda, Ryosuke, et al.
Publicado: (2026)
por: Sonoda, Ryosuke, et al.
Publicado: (2026)
Zero-Shot Vision-and-Language Navigation with Collision Mitigation in Continuous Environment
por: Jeong, Seongjun, et al.
Publicado: (2024)
por: Jeong, Seongjun, et al.
Publicado: (2024)
Zero-Shot Visual Reasoning by Vision-Language Models: Benchmarking and Analysis
por: Nagar, Aishik, et al.
Publicado: (2024)
por: Nagar, Aishik, et al.
Publicado: (2024)
Unconstrained Open Vocabulary Image Classification: Zero-Shot Transfer from Text to Image via CLIP Inversion
por: Allgeuer, Philipp, et al.
Publicado: (2024)
por: Allgeuer, Philipp, et al.
Publicado: (2024)
LightZeroNav: Zero-Shot Vision Language Navigation in Continuous Environments Based on Lightweight VLMs
por: Luo, Kun, et al.
Publicado: (2026)
por: Luo, Kun, et al.
Publicado: (2026)
SAM 3D Animal: Promptable Animal 3D Reconstruction from Images in the Wild
por: Hu, Xuyi, et al.
Publicado: (2026)
por: Hu, Xuyi, et al.
Publicado: (2026)
ZSPAPrune: Zero-Shot Prompt-Aware Token Pruning for Vision-Language Models
por: Zhang, Pu, et al.
Publicado: (2025)
por: Zhang, Pu, et al.
Publicado: (2025)
TINA: Think, Interaction, and Action Framework for Zero-Shot Vision Language Navigation
por: Li, Dingbang, et al.
Publicado: (2024)
por: Li, Dingbang, et al.
Publicado: (2024)
LLM meets Vision-Language Models for Zero-Shot One-Class Classification
por: Bendou, Yassir, et al.
Publicado: (2024)
por: Bendou, Yassir, et al.
Publicado: (2024)
Intriguing Differences Between Zero-Shot and Systematic Evaluations of Vision-Language Transformer Models
por: Salman, Shaeke, et al.
Publicado: (2024)
por: Salman, Shaeke, et al.
Publicado: (2024)
A Framework for Evaluating Zero-Shot Image Generation in Concept-based Explainability
por: Astolfi, Giacomo, et al.
Publicado: (2026)
por: Astolfi, Giacomo, et al.
Publicado: (2026)
Are Video Models Ready as Zero-Shot Reasoners? An Empirical Study with the MME-CoF Benchmark
por: Guo, Ziyu, et al.
Publicado: (2025)
por: Guo, Ziyu, et al.
Publicado: (2025)
Transductive Zero-Shot and Few-Shot CLIP
por: Martin, Ségolène, et al.
Publicado: (2024)
por: Martin, Ségolène, et al.
Publicado: (2024)
PostEdit: Posterior Sampling for Efficient Zero-Shot Image Editing
por: Tian, Feng, et al.
Publicado: (2024)
por: Tian, Feng, et al.
Publicado: (2024)
Zero-Shot Industrial Anomaly Segmentation with Image-Aware Prompt Generation
por: Park, SoYoung, et al.
Publicado: (2025)
por: Park, SoYoung, et al.
Publicado: (2025)
Are Vision Transformer Representations Semantically Meaningful? A Case Study in Medical Imaging
por: Shams, Montasir, et al.
Publicado: (2025)
por: Shams, Montasir, et al.
Publicado: (2025)
TransVLM: A Vision-Language Framework and Benchmark for Detecting Any Shot Transitions
por: Chen, Ce, et al.
Publicado: (2026)
por: Chen, Ce, et al.
Publicado: (2026)
Navigating Data Scarcity using Foundation Models: A Benchmark of Few-Shot and Zero-Shot Learning Approaches in Medical Imaging
por: Woerner, Stefano, et al.
Publicado: (2024)
por: Woerner, Stefano, et al.
Publicado: (2024)
Benchmarking Zero-Shot Robustness of Multimodal Foundation Models: A Pilot Study
por: Wang, Chenguang, et al.
Publicado: (2024)
por: Wang, Chenguang, et al.
Publicado: (2024)
Text-Guided Attention is All You Need for Zero-Shot Robustness in Vision-Language Models
por: Yu, Lu, et al.
Publicado: (2024)
por: Yu, Lu, et al.
Publicado: (2024)
Evaluating Vision-Language Models for Zero-Shot Detection, Classification, and Association of Motorcycles, Passengers, and Helmets
por: Choi, Lucas, et al.
Publicado: (2024)
por: Choi, Lucas, et al.
Publicado: (2024)
Decoupling Endpoint and Semantic Transition Learning for Zero-Shot Composed Image Retrieval
por: Liu, Mingyu, et al.
Publicado: (2026)
por: Liu, Mingyu, et al.
Publicado: (2026)
Buffer Anytime: Zero-Shot Video Depth and Normal from Image Priors
por: Kuang, Zhengfei, et al.
Publicado: (2024)
por: Kuang, Zhengfei, et al.
Publicado: (2024)
Toward Ethical Facial Age Estimation: A Generalized Zero-Shot Benchmark Without Training on Children's Data
por: Petrucci, Caio, et al.
Publicado: (2026)
por: Petrucci, Caio, et al.
Publicado: (2026)
Efficient Masked Attention Transformer for Few-Shot Classification and Segmentation
por: Carrión-Ojeda, Dustin, et al.
Publicado: (2025)
por: Carrión-Ojeda, Dustin, et al.
Publicado: (2025)
Navigating the Trade-off: A Synthesis of Defensive Strategies for Zero-Shot Adversarial Robustness in Vision-Language Models
por: Xu, Zane, et al.
Publicado: (2025)
por: Xu, Zane, et al.
Publicado: (2025)
SparseFormer: Detecting Objects in HRW Shots via Sparse Vision Transformer
por: Li, Wenxi, et al.
Publicado: (2025)
por: Li, Wenxi, et al.
Publicado: (2025)
Temporal Object-Aware Vision Transformer for Few-Shot Video Object Detection
por: Kumar, Yogesh, et al.
Publicado: (2025)
por: Kumar, Yogesh, et al.
Publicado: (2025)
CompareBench: A Benchmark for Visual Comparison Reasoning in Vision-Language Models
por: Cai, Jie, et al.
Publicado: (2025)
por: Cai, Jie, et al.
Publicado: (2025)
Ejemplares similares
-
Zero-Shot Wildlife Sorting Using Vision Transformers: Evaluating Clustering and Continuous Similarity Ordering
por: Markoff, Hugo, et al.
Publicado: (2025) -
Hierarchical Re-Classification: Combining Animal Classification Models with Vision Transformers
por: Markoff, Hugo, et al.
Publicado: (2025) -
Benchmarking Zero-Shot Recognition with Vision-Language Models: Challenges on Granularity and Specificity
por: Xu, Zhenlin, et al.
Publicado: (2023) -
Binary Verification for Zero-Shot Vision
por: Hu, Rongbin, et al.
Publicado: (2025) -
Sea-ing Through Scattered Rays: Revisiting the Image Formation Model for Realistic Underwater Image Generation
por: Ismiroglou, Vasiliki, et al.
Publicado: (2025)