AgriCLIP: Adapting CLIP for Agriculture and Livestock via Domain-Specialized Cross-Model Alignment
Fuente:
arXiv
Guardado en:
| Autores principales: | Nawaz, Umair, Awais, Muhammad, Gani, Hanan, Naseer, Muzammal, Khan, Fahad, Khan, Salman, Anwer, Rao Muhammad |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
AI in Agriculture: A Survey of Deep Learning Techniques for Crops, Fisheries and Livestock
por: Nawaz, Umair, et al.
Publicado: (2025)
por: Nawaz, Umair, et al.
Publicado: (2025)
MedContext: Learning Contextual Cues for Efficient Volumetric Medical Segmentation
por: Gani, Hanan, et al.
Publicado: (2024)
por: Gani, Hanan, et al.
Publicado: (2024)
UniMed-CLIP: Towards a Unified Image-Text Pretraining Paradigm for Diverse Medical Imaging Modalities
por: Khattak, Muhammad Uzair, et al.
Publicado: (2024)
por: Khattak, Muhammad Uzair, et al.
Publicado: (2024)
BAPLe: Backdoor Attacks on Medical Foundational Models using Prompt Learning
por: Hanif, Asif, et al.
Publicado: (2024)
por: Hanif, Asif, et al.
Publicado: (2024)
VANE-Bench: Video Anomaly Evaluation Benchmark for Conversational LMMs
por: Bharadwaj, Rohit, et al.
Publicado: (2024)
por: Bharadwaj, Rohit, et al.
Publicado: (2024)
Align Your Prompts: Test-Time Prompting with Distribution Alignment for Zero-Shot Generalization
por: Hassan, Jameel, et al.
Publicado: (2023)
por: Hassan, Jameel, et al.
Publicado: (2023)
Multi-modal Generation via Cross-Modal In-Context Learning
por: Kumar, Amandeep, et al.
Publicado: (2024)
por: Kumar, Amandeep, et al.
Publicado: (2024)
Composed Video Retrieval via Enriched Context and Discriminative Embeddings
por: Thawakar, Omkar, et al.
Publicado: (2024)
por: Thawakar, Omkar, et al.
Publicado: (2024)
Language Guided Domain Generalized Medical Image Segmentation
por: Kunhimon, Shahina, et al.
Publicado: (2024)
por: Kunhimon, Shahina, et al.
Publicado: (2024)
CDChat: A Large Multimodal Model for Remote Sensing Change Description
por: Noman, Mubashir, et al.
Publicado: (2024)
por: Noman, Mubashir, et al.
Publicado: (2024)
Cross-Modal Self-Training: Aligning Images and Pointclouds to Learn Classification without Labels
por: Dharmasiri, Amaya, et al.
Publicado: (2024)
por: Dharmasiri, Amaya, et al.
Publicado: (2024)
CLIP-Decoder : ZeroShot Multilabel Classification using Multimodal CLIP Aligned Representation
por: Ali, Muhammad, et al.
Publicado: (2024)
por: Ali, Muhammad, et al.
Publicado: (2024)
LLM Blueprint: Enabling Text-to-Image Generation with Complex and Detailed Prompts
por: Gani, Hanan, et al.
Publicado: (2023)
por: Gani, Hanan, et al.
Publicado: (2023)
Enhancing Novel Object Detection via Cooperative Foundational Models
por: Bharadwaj, Rohit, et al.
Publicado: (2023)
por: Bharadwaj, Rohit, et al.
Publicado: (2023)
Rethinking Transformers Pre-training for Multi-Spectral Satellite Imagery
por: Noman, Mubashir, et al.
Publicado: (2024)
por: Noman, Mubashir, et al.
Publicado: (2024)
ObjectCompose: Evaluating Resilience of Vision-Based Models on Object-to-Background Compositional Changes
por: Malik, Hashmat Shadab, et al.
Publicado: (2024)
por: Malik, Hashmat Shadab, et al.
Publicado: (2024)
Hierarchical Text-to-Vision Self Supervised Alignment for Improved Histopathology Representation Learning
por: Watawana, Hasindri, et al.
Publicado: (2024)
por: Watawana, Hasindri, et al.
Publicado: (2024)
Efficient 3D-Aware Facial Image Editing via Attribute-Specific Prompt Learning
por: Kumar, Amandeep, et al.
Publicado: (2024)
por: Kumar, Amandeep, et al.
Publicado: (2024)
VURF: A General-purpose Reasoning and Self-refinement Framework for Video Understanding
por: Mahmood, Ahmad, et al.
Publicado: (2024)
por: Mahmood, Ahmad, et al.
Publicado: (2024)
Robust-LLaVA: On the Effectiveness of Large-Scale Robust Image Encoders for Multi-modal Large Language Models
por: Malik, Hashmat Shadab, et al.
Publicado: (2025)
por: Malik, Hashmat Shadab, et al.
Publicado: (2025)
MATRIX: Multimodal Agent Tuning for Robust Tool-Use Reasoning
por: Ashraf, Tajamul, et al.
Publicado: (2025)
por: Ashraf, Tajamul, et al.
Publicado: (2025)
How Good is my Video LMM? Complex Video Reasoning and Robustness Evaluation Suite for Video-LMMs
por: Khattak, Muhammad Uzair, et al.
Publicado: (2024)
por: Khattak, Muhammad Uzair, et al.
Publicado: (2024)
Learnable Weight Initialization for Volumetric Medical Image Segmentation
por: Kunhimon, Shahina, et al.
Publicado: (2023)
por: Kunhimon, Shahina, et al.
Publicado: (2023)
Hierarchical Self-Supervised Adversarial Training for Robust Vision Models in Histopathology
por: Malik, Hashmat Shadab, et al.
Publicado: (2025)
por: Malik, Hashmat Shadab, et al.
Publicado: (2025)
Towards Evaluating the Robustness of Visual State Space Models
por: Malik, Hashmat Shadab, et al.
Publicado: (2024)
por: Malik, Hashmat Shadab, et al.
Publicado: (2024)
Video-GroundingDINO: Towards Open-Vocabulary Spatio-Temporal Video Grounding
por: Wasim, Syed Talal, et al.
Publicado: (2023)
por: Wasim, Syed Talal, et al.
Publicado: (2023)
DEFT: Decompositional Efficient Fine-Tuning for Text-to-Image Models
por: Kumar, Komal, et al.
Publicado: (2025)
por: Kumar, Komal, et al.
Publicado: (2025)
Tracking Meets Large Multimodal Models for Driving Scenario Understanding
por: Ishaq, Ayesha, et al.
Publicado: (2025)
por: Ishaq, Ayesha, et al.
Publicado: (2025)
Towards Fine-Grained Adaptation of CLIP via a Self-Trained Alignment Score
por: Ali, Eman, et al.
Publicado: (2025)
por: Ali, Eman, et al.
Publicado: (2025)
MedROV: Towards Real-Time Open-Vocabulary Detection Across Diverse Medical Imaging Modalities
por: Sheikh, Tooba Tehreem, et al.
Publicado: (2025)
por: Sheikh, Tooba Tehreem, et al.
Publicado: (2025)
TerraFM: A Scalable Foundation Model for Unified Multisensor Earth Observation
por: Danish, Muhammad Sohail, et al.
Publicado: (2025)
por: Danish, Muhammad Sohail, et al.
Publicado: (2025)
microCLIP: Unsupervised CLIP Adaptation via Coarse-Fine Token Fusion for Fine-Grained Image Classification
por: Silva, Sathira, et al.
Publicado: (2025)
por: Silva, Sathira, et al.
Publicado: (2025)
Beyond Simple Edits: Composed Video Retrieval with Dense Modifications
por: Thawakar, Omkar, et al.
Publicado: (2025)
por: Thawakar, Omkar, et al.
Publicado: (2025)
WorldCache: Content-Aware Caching for Accelerated Video World Models
por: Nawaz, Umair, et al.
Publicado: (2026)
por: Nawaz, Umair, et al.
Publicado: (2026)
Makeup-Guided Facial Privacy Protection via Untrained Neural Network Priors
por: Shamshad, Fahad, et al.
Publicado: (2024)
por: Shamshad, Fahad, et al.
Publicado: (2024)
AD-CLIP: Adapting Domains in Prompt Space Using CLIP
por: Singha, Mainak, et al.
Publicado: (2023)
por: Singha, Mainak, et al.
Publicado: (2023)
Agent-X: Evaluating Deep Multimodal Reasoning in Vision-Centric Agentic Tasks
por: Ashraf, Tajamul, et al.
Publicado: (2025)
por: Ashraf, Tajamul, et al.
Publicado: (2025)
EvoLMM: Self-Evolving Large Multimodal Models with Continuous Rewards
por: Thawakar, Omkar, et al.
Publicado: (2025)
por: Thawakar, Omkar, et al.
Publicado: (2025)
VideoGLaMM: A Large Multimodal Model for Pixel-Level Visual Grounding in Videos
por: Munasinghe, Shehan, et al.
Publicado: (2024)
por: Munasinghe, Shehan, et al.
Publicado: (2024)
Video-ChatGPT: Towards Detailed Video Understanding via Large Vision and Language Models
por: Maaz, Muhammad, et al.
Publicado: (2023)
por: Maaz, Muhammad, et al.
Publicado: (2023)
Ejemplares similares
-
AI in Agriculture: A Survey of Deep Learning Techniques for Crops, Fisheries and Livestock
por: Nawaz, Umair, et al.
Publicado: (2025) -
MedContext: Learning Contextual Cues for Efficient Volumetric Medical Segmentation
por: Gani, Hanan, et al.
Publicado: (2024) -
UniMed-CLIP: Towards a Unified Image-Text Pretraining Paradigm for Diverse Medical Imaging Modalities
por: Khattak, Muhammad Uzair, et al.
Publicado: (2024) -
BAPLe: Backdoor Attacks on Medical Foundational Models using Prompt Learning
por: Hanif, Asif, et al.
Publicado: (2024) -
VANE-Bench: Video Anomaly Evaluation Benchmark for Conversational LMMs
por: Bharadwaj, Rohit, et al.
Publicado: (2024)