Composing Open-domain Vision with RAG for Ocean Monitoring and Conservation
Fuente:
arXiv
Saved in:
| Main Authors: | Dyanatkar, Sepand, Li, Angran, Dungate, Alexander |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Efficient Adaptation of Large Vision Transformer via Adapter Re-Composing
by: Dong, Wei, et al.
Published: (2023)
by: Dong, Wei, et al.
Published: (2023)
MMed-RAG: Versatile Multimodal RAG System for Medical Vision Language Models
by: Xia, Peng, et al.
Published: (2024)
by: Xia, Peng, et al.
Published: (2024)
ViewFusion: Learning Composable Diffusion Models for Novel View Synthesis
by: Spiegl, Bernard, et al.
Published: (2024)
by: Spiegl, Bernard, et al.
Published: (2024)
Open-Insect: Benchmarking Open-Set Recognition of Novel Species in Biodiversity Monitoring
by: Chen, Yuyan, et al.
Published: (2025)
by: Chen, Yuyan, et al.
Published: (2025)
Enhancing Pollinator Conservation towards Agriculture 4.0: Monitoring of Bees through Object Recognition
by: Alex, Ajay John, et al.
Published: (2024)
by: Alex, Ajay John, et al.
Published: (2024)
Decomposing and Composing: Towards Efficient Vision-Language Continual Learning via Rank-1 Expert Pool in a Single LoRA
by: Fa, Zhan, et al.
Published: (2026)
by: Fa, Zhan, et al.
Published: (2026)
Garbage Vulnerable Point Monitoring using IoT and Computer Vision
by: Kumar, R., et al.
Published: (2025)
by: Kumar, R., et al.
Published: (2025)
Controllable Image Generation with Composed Parallel Token Prediction
by: Stirling, Jamie, et al.
Published: (2024)
by: Stirling, Jamie, et al.
Published: (2024)
Interaction Asymmetry: A General Principle for Learning Composable Abstractions
by: Brady, Jack, et al.
Published: (2024)
by: Brady, Jack, et al.
Published: (2024)
Decompose-and-Compose: A Compositional Approach to Mitigating Spurious Correlation
by: Noohdani, Fahimeh Hosseini, et al.
Published: (2024)
by: Noohdani, Fahimeh Hosseini, et al.
Published: (2024)
Learning to Compose: Improving Object Centric Learning by Injecting Compositionality
by: Jung, Whie, et al.
Published: (2024)
by: Jung, Whie, et al.
Published: (2024)
FIRE-CIR: Fine-grained Reasoning for Composed Fashion Image Retrieval
by: Gardères, François, et al.
Published: (2026)
by: Gardères, François, et al.
Published: (2026)
GeoFormer: A Vision and Sequence Transformer-based Approach for Greenhouse Gas Monitoring
by: Khirwar, Madhav, et al.
Published: (2024)
by: Khirwar, Madhav, et al.
Published: (2024)
Toward Autonomous Laboratory Safety Monitoring with Vision Language Models: Learning to See Hazards Through Scene Structure
by: Chakraborty, Trishna, et al.
Published: (2026)
by: Chakraborty, Trishna, et al.
Published: (2026)
ConText-CIR: Learning from Concepts in Text for Composed Image Retrieval
by: Xing, Eric, et al.
Published: (2025)
by: Xing, Eric, et al.
Published: (2025)
Drifting Fields are not Conservative
by: Franz, Leonard T., et al.
Published: (2026)
by: Franz, Leonard T., et al.
Published: (2026)
Seeing and Knowing in the Wild: Open-domain Visual Entity Recognition with Large-scale Knowledge Graphs via Contrastive Learning
by: Zhou, Hongkuan, et al.
Published: (2025)
by: Zhou, Hongkuan, et al.
Published: (2025)
Benchmarking Vision, Language, & Action Models in Procedurally Generated, Open Ended Action Environments
by: Guruprasad, Pranav, et al.
Published: (2025)
by: Guruprasad, Pranav, et al.
Published: (2025)
Open Set Domain Adaptation with Vision-language models via Gradient-aware Separation
by: Chen, Haoyang
Published: (2025)
by: Chen, Haoyang
Published: (2025)
Align Your Gaussians: Text-to-4D with Dynamic 3D Gaussians and Composed Diffusion Models
by: Ling, Huan, et al.
Published: (2023)
by: Ling, Huan, et al.
Published: (2023)
Reason, Retrieve, Re-rank: A Zero-Shot Reasoning-Aware Framework for Composed Video Retrieval
by: Alavi, Ali
Published: (2026)
by: Alavi, Ali
Published: (2026)
Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models
by: Deitke, Matt, et al.
Published: (2024)
by: Deitke, Matt, et al.
Published: (2024)
Composing Parts for Expressive Object Generation
by: Rangwani, Harsh, et al.
Published: (2024)
by: Rangwani, Harsh, et al.
Published: (2024)
Beyond Spatial Compression: Interface-Centric Generative States for Open-World 3D Structure
by: Chen, Xiang, et al.
Published: (2026)
by: Chen, Xiang, et al.
Published: (2026)
T-QPM: Enabling Temporal Out-Of-Distribution Detection and Domain Generalization for Vision-Language Models in Open-World
by: Naiknaware, Aditi, et al.
Published: (2026)
by: Naiknaware, Aditi, et al.
Published: (2026)
Closing the Visual Sim-to-Real Gap with Object-Composable NeRFs
by: Mishra, Nikhil, et al.
Published: (2024)
by: Mishra, Nikhil, et al.
Published: (2024)
ComposableNav: Instruction-Following Navigation in Dynamic Environments via Composable Diffusion
by: Hu, Zichao, et al.
Published: (2025)
by: Hu, Zichao, et al.
Published: (2025)
How Diffusion Models Learn to Factorize and Compose
by: Liang, Qiyao, et al.
Published: (2024)
by: Liang, Qiyao, et al.
Published: (2024)
RULE: Reliable Multimodal RAG for Factuality in Medical Vision Language Models
by: Xia, Peng, et al.
Published: (2024)
by: Xia, Peng, et al.
Published: (2024)
Transferable Adversarial Attacks on Black-Box Vision-Language Models
by: Hu, Kai, et al.
Published: (2025)
by: Hu, Kai, et al.
Published: (2025)
Navigating the Emotion Tree: Hierarchical Hyperbolic RAG for Multimodal Emotion Recognition
by: Wang, Zeheng, et al.
Published: (2026)
by: Wang, Zeheng, et al.
Published: (2026)
iRAG: Advancing RAG for Videos with an Incremental Approach
by: Arefeen, Md Adnan, et al.
Published: (2024)
by: Arefeen, Md Adnan, et al.
Published: (2024)
Composable Part-Based Manipulation
by: Liu, Weiyu, et al.
Published: (2024)
by: Liu, Weiyu, et al.
Published: (2024)
Pytorch-Wildlife: A Collaborative Deep Learning Framework for Conservation
by: Hernandez, Andres, et al.
Published: (2024)
by: Hernandez, Andres, et al.
Published: (2024)
Source-Free Domain Adaptation Guided by Vision and Vision-Language Pre-Training
by: Zhang, Wenyu, et al.
Published: (2024)
by: Zhang, Wenyu, et al.
Published: (2024)
FairRAG: Fair Human Generation via Fair Retrieval Augmentation
by: Shrestha, Robik, et al.
Published: (2024)
by: Shrestha, Robik, et al.
Published: (2024)
Vision-based Manipulation from Single Human Video with Open-World Object Graphs
by: Zhu, Yifeng, et al.
Published: (2024)
by: Zhu, Yifeng, et al.
Published: (2024)
Unlocking the Power of Open Set : A New Perspective for Open-Set Noisy Label Learning
by: Wan, Wenhai, et al.
Published: (2023)
by: Wan, Wenhai, et al.
Published: (2023)
Vision Transformer-based Adversarial Domain Adaptation
by: Li, Yahan, et al.
Published: (2024)
by: Li, Yahan, et al.
Published: (2024)
Cancer-Net SCa-Synth: An Open Access Synthetically Generated 2D Skin Lesion Dataset for Skin Cancer Classification
by: Tai, Chi-en Amy, et al.
Published: (2024)
by: Tai, Chi-en Amy, et al.
Published: (2024)
Similar Items
-
Efficient Adaptation of Large Vision Transformer via Adapter Re-Composing
by: Dong, Wei, et al.
Published: (2023) -
MMed-RAG: Versatile Multimodal RAG System for Medical Vision Language Models
by: Xia, Peng, et al.
Published: (2024) -
ViewFusion: Learning Composable Diffusion Models for Novel View Synthesis
by: Spiegl, Bernard, et al.
Published: (2024) -
Open-Insect: Benchmarking Open-Set Recognition of Novel Species in Biodiversity Monitoring
by: Chen, Yuyan, et al.
Published: (2025) -
Enhancing Pollinator Conservation towards Agriculture 4.0: Monitoring of Bees through Object Recognition
by: Alex, Ajay John, et al.
Published: (2024)