BioCLIP: A Vision Foundation Model for the Tree of Life
Fuente:
arXiv
Saved in:
| Main Authors: | Stevens, Samuel, Wu, Jiaman, Thompson, Matthew J, Campolongo, Elizabeth G, Song, Chan Hee, Carlyn, David Edward, Dong, Li, Dahdul, Wasila M, Stewart, Charles, Berger-Wolf, Tanya, Chao, Wei-Lun, Su, Yu |
|---|---|
| Format: | Preprint |
| Published: |
2023
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
BioCLIP 2: Emergent Properties from Scaling Hierarchical Contrastive Learning
by: Gu, Jianyang, et al.
Published: (2025)
by: Gu, Jianyang, et al.
Published: (2025)
BioCAP: Exploiting Synthetic Captions Beyond Labels in Biological Foundation Models
by: Zhang, Ziheng, et al.
Published: (2025)
by: Zhang, Ziheng, et al.
Published: (2025)
Interpretable and Testable Vision Features via Sparse Autoencoders
by: Stevens, Samuel, et al.
Published: (2025)
by: Stevens, Samuel, et al.
Published: (2025)
What Do You See in Common? Learning Hierarchical Prototypes over Tree-of-Life to Discover Evolutionary Traits
by: Manogaran, Harish Babu, et al.
Published: (2024)
by: Manogaran, Harish Babu, et al.
Published: (2024)
VLM4Bio: A Benchmark Dataset to Evaluate Pretrained Vision-Language Models for Trait Discovery from Biological Images
by: Maruf, M., et al.
Published: (2024)
by: Maruf, M., et al.
Published: (2024)
Hierarchical Conditioning of Diffusion Models Using Tree-of-Life for Studying Species Evolution
by: Khurana, Mridul, et al.
Published: (2024)
by: Khurana, Mridul, et al.
Published: (2024)
Fish-Vista: A Multi-Purpose Dataset for Understanding & Identification of Traits from Images
by: Mehrab, Kazi Sajeed, et al.
Published: (2024)
by: Mehrab, Kazi Sajeed, et al.
Published: (2024)
Towards Open-Ended Visual Scientific Discovery with Sparse Autoencoders
by: Stevens, Samuel, et al.
Published: (2025)
by: Stevens, Samuel, et al.
Published: (2025)
Finer-CAM: Spotting the Difference Reveals Finer Details for Visual Explanation
by: Zhang, Ziheng, et al.
Published: (2025)
by: Zhang, Ziheng, et al.
Published: (2025)
Prompt-CAM: Making Vision Transformers Interpretable for Fine-Grained Analysis
by: Chowdhury, Arpita, et al.
Published: (2025)
by: Chowdhury, Arpita, et al.
Published: (2025)
A Simple Interpretable Transformer for Fine-Grained Image Classification and Analysis
by: Paul, Dipanjyoti, et al.
Published: (2023)
by: Paul, Dipanjyoti, et al.
Published: (2023)
kabr-tools: Automated Framework for Multi-Species Behavioral Monitoring
by: Kline, Jenna, et al.
Published: (2025)
by: Kline, Jenna, et al.
Published: (2025)
BioBench: A Blueprint to Move Beyond ImageNet for Scientific ML Benchmarks
by: Stevens, Samuel
Published: (2025)
by: Stevens, Samuel
Published: (2025)
TaxaAdapter: Vision Taxonomy Models are Key to Fine-grained Image Generation over the Tree of Life
by: Khurana, Mridul, et al.
Published: (2026)
by: Khurana, Mridul, et al.
Published: (2026)
DOFA-CLIP: Multimodal Vision-Language Foundation Models for Earth Observation
by: Xiong, Zhitong, et al.
Published: (2025)
by: Xiong, Zhitong, et al.
Published: (2025)
A continental-scale dataset of ground beetles with high-resolution images and validated morphological trait measurements
by: Rayeed, S M, et al.
Published: (2026)
by: Rayeed, S M, et al.
Published: (2026)
Finer-Personalization Rank: Fine-Grained Retrieval Examines Identity Preservation for Personalized Generation
by: Kilrain, Connor, et al.
Published: (2025)
by: Kilrain, Connor, et al.
Published: (2025)
RemoteCLIP: A Vision Language Foundation Model for Remote Sensing
by: Liu, Fan, et al.
Published: (2023)
by: Liu, Fan, et al.
Published: (2023)
Reviving the Context: Camera Trap Species Classification as Link Prediction on Multimodal Knowledge Graphs
by: Pahuja, Vardaan, et al.
Published: (2023)
by: Pahuja, Vardaan, et al.
Published: (2023)
Optimizing Image Capture for Computer Vision-Powered Taxonomic Identification and Trait Recognition of Biodiversity Specimens
by: East, Alyson, et al.
Published: (2025)
by: East, Alyson, et al.
Published: (2025)
Dual-View Visual Contextualization for Web Navigation
by: Kil, Jihyung, et al.
Published: (2024)
by: Kil, Jihyung, et al.
Published: (2024)
MoralCLIP: Contrastive Alignment of Vision-and-Language Representations with Moral Foundations Theory
by: Condez, Ana Carolina, et al.
Published: (2025)
by: Condez, Ana Carolina, et al.
Published: (2025)
Leveraging Latent Visual Reasoning in Silence
by: Zhu, Dongyao, et al.
Published: (2026)
by: Zhu, Dongyao, et al.
Published: (2026)
HotSpotter - Patterned Species Instance Recognition
by: Crall, Jonathan P., et al.
Published: (2025)
by: Crall, Jonathan P., et al.
Published: (2025)
Tracking Phenological Status and Ecological Interactions in a Hawaiian Cloud Forest Understory using Low-Cost Camera Traps and Visual Foundation Models
by: Meyers, Luke, et al.
Published: (2026)
by: Meyers, Luke, et al.
Published: (2026)
AgentZero++: Modeling Fear-Based Behavior
by: Malhotra, Vrinda, et al.
Published: (2025)
by: Malhotra, Vrinda, et al.
Published: (2025)
Mind the (Data) Gap: Evaluating Vision Systems in Small Data Applications
by: Stevens, Samuel, et al.
Published: (2025)
by: Stevens, Samuel, et al.
Published: (2025)
TripletCLIP: Improving Compositional Reasoning of CLIP via Synthetic Vision-Language Negatives
by: Patel, Maitreya, et al.
Published: (2024)
by: Patel, Maitreya, et al.
Published: (2024)
SAM-CLIP: Merging Vision Foundation Models towards Semantic and Spatial Understanding
by: Wang, Haoxiang, et al.
Published: (2023)
by: Wang, Haoxiang, et al.
Published: (2023)
AquaticCLIP: A Vision-Language Foundation Model for Underwater Scene Analysis
by: Alawode, Basit, et al.
Published: (2025)
by: Alawode, Basit, et al.
Published: (2025)
Fine-Tuning is Fine, if Calibrated
by: Mai, Zheda, et al.
Published: (2024)
by: Mai, Zheda, et al.
Published: (2024)
Static Segmentation by Tracking: A Label-Efficient Approach for Fine-Grained Specimen Image Segmentation
by: Feng, Zhenyang, et al.
Published: (2025)
by: Feng, Zhenyang, et al.
Published: (2025)
ClearCLIP: Decomposing CLIP Representations for Dense Vision-Language Inference
by: Lan, Mengcheng, et al.
Published: (2024)
by: Lan, Mengcheng, et al.
Published: (2024)
UAVD-Mamba: Deformable Token Fusion Vision Mamba for Multimodal UAV Detection
by: Li, Wei, et al.
Published: (2025)
by: Li, Wei, et al.
Published: (2025)
Proto-CLIP: Vision-Language Prototypical Network for Few-Shot Learning
by: P, Jishnu Jaykumar, et al.
Published: (2023)
by: P, Jishnu Jaykumar, et al.
Published: (2023)
AVA-Bench: Atomic Visual Ability Benchmark for Vision Foundation Models
by: Mai, Zheda, et al.
Published: (2025)
by: Mai, Zheda, et al.
Published: (2025)
CLIP-Adapter: Better Vision-Language Models with Feature Adapters
by: Gao, Peng, et al.
Published: (2021)
by: Gao, Peng, et al.
Published: (2021)
Cardiac-CLIP: A Vision-Language Foundation Model for 3D Cardiac CT Images
by: Hu, Yutao, et al.
Published: (2025)
by: Hu, Yutao, et al.
Published: (2025)
Mammo-CLIP: A Vision Language Foundation Model to Enhance Data Efficiency and Robustness in Mammography
by: Ghosh, Shantanu, et al.
Published: (2024)
by: Ghosh, Shantanu, et al.
Published: (2024)
MedProbCLIP: Probabilistic Adaptation of Vision-Language Foundation Model for Reliable Radiograph-Report Retrieval
by: Elallaf, Ahmad, et al.
Published: (2026)
by: Elallaf, Ahmad, et al.
Published: (2026)
Similar Items
-
BioCLIP 2: Emergent Properties from Scaling Hierarchical Contrastive Learning
by: Gu, Jianyang, et al.
Published: (2025) -
BioCAP: Exploiting Synthetic Captions Beyond Labels in Biological Foundation Models
by: Zhang, Ziheng, et al.
Published: (2025) -
Interpretable and Testable Vision Features via Sparse Autoencoders
by: Stevens, Samuel, et al.
Published: (2025) -
What Do You See in Common? Learning Hierarchical Prototypes over Tree-of-Life to Discover Evolutionary Traits
by: Manogaran, Harish Babu, et al.
Published: (2024) -
VLM4Bio: A Benchmark Dataset to Evaluate Pretrained Vision-Language Models for Trait Discovery from Biological Images
by: Maruf, M., et al.
Published: (2024)