TaxaBind: A Unified Embedding Space for Ecological Applications
Fuente:
arXiv
Saved in:
| Main Authors: | Sastry, Srikumar, Khanal, Subash, Dhakal, Aayush, Ahmad, Adeel, Jacobs, Nathan |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Sat2Sound: A Unified Framework for Zero-Shot Soundscape Mapping
by: Khanal, Subash, et al.
Published: (2025)
by: Khanal, Subash, et al.
Published: (2025)
RANGE: Retrieval Augmented Neural Fields for Multi-Resolution Geo-Embeddings
by: Dhakal, Aayush, et al.
Published: (2025)
by: Dhakal, Aayush, et al.
Published: (2025)
GeoSynth: Contextually-Aware High-Resolution Satellite Image Synthesis
by: Sastry, Srikumar, et al.
Published: (2024)
by: Sastry, Srikumar, et al.
Published: (2024)
LD-SDM: Language-Driven Hierarchical Species Distribution Modeling
by: Sastry, Srikumar, et al.
Published: (2023)
by: Sastry, Srikumar, et al.
Published: (2023)
Sat2Cap: Mapping Fine-Grained Textual Descriptions from Satellite Images
by: Dhakal, Aayush, et al.
Published: (2023)
by: Dhakal, Aayush, et al.
Published: (2023)
ProM3E: Probabilistic Masked MultiModal Embedding Model for Ecology
by: Sastry, Srikumar, et al.
Published: (2025)
by: Sastry, Srikumar, et al.
Published: (2025)
Global and Local Entailment Learning for Natural World Imagery
by: Sastry, Srikumar, et al.
Published: (2025)
by: Sastry, Srikumar, et al.
Published: (2025)
PSM: Learning Probabilistic Embeddings for Multi-scale Zero-Shot Soundscape Mapping
by: Khanal, Subash, et al.
Published: (2024)
by: Khanal, Subash, et al.
Published: (2024)
GEOBIND: Binding Text, Image, and Audio through Satellite Images
by: Dhakal, Aayush, et al.
Published: (2024)
by: Dhakal, Aayush, et al.
Published: (2024)
GeoDiT: Point-Conditioned Diffusion Transformer for Satellite Image Synthesis
by: Sastry, Srikumar, et al.
Published: (2026)
by: Sastry, Srikumar, et al.
Published: (2026)
SimLBR: Learning to Detect Fake Images by Learning to Detect Real Images
by: Dhakal, Aayush, et al.
Published: (2026)
by: Dhakal, Aayush, et al.
Published: (2026)
DiffVAS: Diffusion-Guided Visual Active Search in Partially Observable Environments
by: Sarkar, Anindya, et al.
Published: (2026)
by: Sarkar, Anindya, et al.
Published: (2026)
GOMAA-Geo: GOal Modality Agnostic Active Geo-localization
by: Sarkar, Anindya, et al.
Published: (2024)
by: Sarkar, Anindya, et al.
Published: (2024)
VLSM-Adapter: Finetuning Vision-Language Segmentation Efficiently with Lightweight Blocks
by: Dhakal, Manish, et al.
Published: (2024)
by: Dhakal, Manish, et al.
Published: (2024)
Exploring Transfer Learning in Medical Image Segmentation using Vision-Language Models
by: Poudel, Kanchan, et al.
Published: (2023)
by: Poudel, Kanchan, et al.
Published: (2023)
CodeBind: Decoupled Representation Learning for Multimodal Alignment with Unified Compositional Codebook
by: Chen, Zeyu, et al.
Published: (2026)
by: Chen, Zeyu, et al.
Published: (2026)
VectorSynth: Fine-Grained Satellite Image Synthesis with Structured Semantics
by: Cher, Daniel, et al.
Published: (2025)
by: Cher, Daniel, et al.
Published: (2025)
ChartGemma: Visual Instruction-tuning for Chart Reasoning in the Wild
by: Masry, Ahmed, et al.
Published: (2024)
by: Masry, Ahmed, et al.
Published: (2024)
Analyzing Diffusion and Autoregressive Vision Language Models in Multimodal Embedding Space
by: Wang, Zihang, et al.
Published: (2026)
by: Wang, Zihang, et al.
Published: (2026)
TuneVLSeg: Prompt Tuning Benchmark for Vision-Language Segmentation Models
by: Adhikari, Rabin, et al.
Published: (2024)
by: Adhikari, Rabin, et al.
Published: (2024)
Does CLIP Bind Concepts? Probing Compositionality in Large Image Models
by: Lewis, Martha, et al.
Published: (2022)
by: Lewis, Martha, et al.
Published: (2022)
RichSpace: Enriching Text-to-Video Prompt Space via Text Embedding Interpolation
by: Cao, Yuefan, et al.
Published: (2025)
by: Cao, Yuefan, et al.
Published: (2025)
Interfacing Foundation Models' Embeddings
by: Zou, Xueyan, et al.
Published: (2023)
by: Zou, Xueyan, et al.
Published: (2023)
Mixed-View Panorama Synthesis using Geospatially Guided Diffusion
by: Xiong, Zhexiao, et al.
Published: (2024)
by: Xiong, Zhexiao, et al.
Published: (2024)
Unified Vision-Language Modeling via Concept Space Alignment
by: Qiu, Yifu, et al.
Published: (2026)
by: Qiu, Yifu, et al.
Published: (2026)
UMBRAE: Unified Multimodal Brain Decoding
by: Xia, Weihao, et al.
Published: (2024)
by: Xia, Weihao, et al.
Published: (2024)
Beta Distribution Learning for Reliable Roadway Crash Risk Assessment
by: Elallaf, Ahmad, et al.
Published: (2025)
by: Elallaf, Ahmad, et al.
Published: (2025)
Unsupervised Document and Template Clustering using Multimodal Embeddings
by: Sampaio, Phillipe R., et al.
Published: (2025)
by: Sampaio, Phillipe R., et al.
Published: (2025)
MATE: Meet At The Embedding -- Connecting Images with Long Texts
by: Jang, Young Kyun, et al.
Published: (2024)
by: Jang, Young Kyun, et al.
Published: (2024)
CAFe: Unifying Representation and Generation with Contrastive-Autoregressive Finetuning
by: Yu, Hao, et al.
Published: (2025)
by: Yu, Hao, et al.
Published: (2025)
SCOPE: Sign Language Contextual Processing with Embedding from LLMs
by: Liu, Yuqi, et al.
Published: (2024)
by: Liu, Yuqi, et al.
Published: (2024)
On the Limits of Token Reduction for Efficient Unified Vision Language Training
by: Chen, Siyi, et al.
Published: (2026)
by: Chen, Siyi, et al.
Published: (2026)
Janus: Decoupling Visual Encoding for Unified Multimodal Understanding and Generation
by: Wu, Chengyue, et al.
Published: (2024)
by: Wu, Chengyue, et al.
Published: (2024)
e5-omni: Explicit Cross-modal Alignment for Omni-modal Embeddings
by: Chen, Haonan, et al.
Published: (2026)
by: Chen, Haonan, et al.
Published: (2026)
Janus-Pro: Unified Multimodal Understanding and Generation with Data and Model Scaling
by: Chen, Xiaokang, et al.
Published: (2025)
by: Chen, Xiaokang, et al.
Published: (2025)
Bootstrapping Action-Grounded Visual Dynamics in Unified Vision-Language Models
by: Qiu, Yifu, et al.
Published: (2025)
by: Qiu, Yifu, et al.
Published: (2025)
VLM2Vec: Training Vision-Language Models for Massive Multimodal Embedding Tasks
by: Jiang, Ziyan, et al.
Published: (2024)
by: Jiang, Ziyan, et al.
Published: (2024)
Lingshu: A Generalist Foundation Model for Unified Multimodal Medical Understanding and Reasoning
by: LASA Team, et al.
Published: (2025)
by: LASA Team, et al.
Published: (2025)
UniCode: Learning a Unified Codebook for Multimodal Large Language Models
by: Zheng, Sipeng, et al.
Published: (2024)
by: Zheng, Sipeng, et al.
Published: (2024)
JanusFlow: Harmonizing Autoregression and Rectified Flow for Unified Multimodal Understanding and Generation
by: Ma, Yiyang, et al.
Published: (2024)
by: Ma, Yiyang, et al.
Published: (2024)
Similar Items
-
Sat2Sound: A Unified Framework for Zero-Shot Soundscape Mapping
by: Khanal, Subash, et al.
Published: (2025) -
RANGE: Retrieval Augmented Neural Fields for Multi-Resolution Geo-Embeddings
by: Dhakal, Aayush, et al.
Published: (2025) -
GeoSynth: Contextually-Aware High-Resolution Satellite Image Synthesis
by: Sastry, Srikumar, et al.
Published: (2024) -
LD-SDM: Language-Driven Hierarchical Species Distribution Modeling
by: Sastry, Srikumar, et al.
Published: (2023) -
Sat2Cap: Mapping Fine-Grained Textual Descriptions from Satellite Images
by: Dhakal, Aayush, et al.
Published: (2023)