Gespeichert in:
| Hauptverfasser: | Dudhane, Akshay, Thawakar, Omkar, Zamir, Syed Waqas, Khan, Salman, Khan, Fahad Shahbaz, Yang, Ming-Hsuan |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | https://arxiv.org/abs/2404.02154 |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
AdaIR: Adaptive All-in-One Image Restoration via Frequency Mining and Modulation
von: Cui, Yuning, et al.
Veröffentlicht: (2024)
von: Cui, Yuning, et al.
Veröffentlicht: (2024)
Vocabulary-free Fine-grained Visual Recognition via Enriched Contextually Grounded Vision-Language Model
von: Demidov, Dmitry, et al.
Veröffentlicht: (2025)
von: Demidov, Dmitry, et al.
Veröffentlicht: (2025)
Video-GroundingDINO: Towards Open-Vocabulary Spatio-Temporal Video Grounding
von: Wasim, Syed Talal, et al.
Veröffentlicht: (2023)
von: Wasim, Syed Talal, et al.
Veröffentlicht: (2023)
Composed Video Retrieval via Enriched Context and Discriminative Embeddings
von: Thawakar, Omkar, et al.
Veröffentlicht: (2024)
von: Thawakar, Omkar, et al.
Veröffentlicht: (2024)
Beyond Simple Edits: Composed Video Retrieval with Dense Modifications
von: Thawakar, Omkar, et al.
Veröffentlicht: (2025)
von: Thawakar, Omkar, et al.
Veröffentlicht: (2025)
Efficient Video Object Segmentation via Modulated Cross-Attention Memory
von: Shaker, Abdelrahman, et al.
Veröffentlicht: (2024)
von: Shaker, Abdelrahman, et al.
Veröffentlicht: (2024)
UNETR++: Delving into Efficient and Accurate 3D Medical Image Segmentation
von: Shaker, Abdelrahman, et al.
Veröffentlicht: (2022)
von: Shaker, Abdelrahman, et al.
Veröffentlicht: (2022)
LLM Post-Training: A Deep Dive into Reasoning Large Language Models
von: Kumar, Komal, et al.
Veröffentlicht: (2025)
von: Kumar, Komal, et al.
Veröffentlicht: (2025)
XrayGPT: Chest Radiographs Summarization using Medical Vision-Language Models
von: Thawakar, Omkar, et al.
Veröffentlicht: (2023)
von: Thawakar, Omkar, et al.
Veröffentlicht: (2023)
Rethinking Transformers Pre-training for Multi-Spectral Satellite Imagery
von: Noman, Mubashir, et al.
Veröffentlicht: (2024)
von: Noman, Mubashir, et al.
Veröffentlicht: (2024)
Time Travel: A Comprehensive Benchmark to Evaluate LMMs on Historical and Cultural Artifacts
von: Ghaboura, Sara, et al.
Veröffentlicht: (2025)
von: Ghaboura, Sara, et al.
Veröffentlicht: (2025)
GroupMamba: Efficient Group-Based Visual State Space Model
von: Shaker, Abdelrahman, et al.
Veröffentlicht: (2024)
von: Shaker, Abdelrahman, et al.
Veröffentlicht: (2024)
Language Guided Domain Generalized Medical Image Segmentation
von: Kunhimon, Shahina, et al.
Veröffentlicht: (2024)
von: Kunhimon, Shahina, et al.
Veröffentlicht: (2024)
Video-ChatGPT: Towards Detailed Video Understanding via Large Vision and Language Models
von: Maaz, Muhammad, et al.
Veröffentlicht: (2023)
von: Maaz, Muhammad, et al.
Veröffentlicht: (2023)
Video-CoM: Interactive Video Reasoning via Chain of Manipulations
von: Rasheed, Hanoona, et al.
Veröffentlicht: (2025)
von: Rasheed, Hanoona, et al.
Veröffentlicht: (2025)
ThinkGeo: Evaluating Tool-Augmented Agents for Remote Sensing Tasks
von: Shabbir, Akashah, et al.
Veröffentlicht: (2025)
von: Shabbir, Akashah, et al.
Veröffentlicht: (2025)
Cross-Modal Self-Training: Aligning Images and Pointclouds to Learn Classification without Labels
von: Dharmasiri, Amaya, et al.
Veröffentlicht: (2024)
von: Dharmasiri, Amaya, et al.
Veröffentlicht: (2024)
UniMed-CLIP: Towards a Unified Image-Text Pretraining Paradigm for Diverse Medical Imaging Modalities
von: Khattak, Muhammad Uzair, et al.
Veröffentlicht: (2024)
von: Khattak, Muhammad Uzair, et al.
Veröffentlicht: (2024)
EvoLMM: Self-Evolving Large Multimodal Models with Continuous Rewards
von: Thawakar, Omkar, et al.
Veröffentlicht: (2025)
von: Thawakar, Omkar, et al.
Veröffentlicht: (2025)
Mobile-O: Unified Multimodal Understanding and Generation on Mobile Device
von: Shaker, Abdelrahman, et al.
Veröffentlicht: (2026)
von: Shaker, Abdelrahman, et al.
Veröffentlicht: (2026)
Towards Evaluating the Robustness of Visual State Space Models
von: Malik, Hashmat Shadab, et al.
Veröffentlicht: (2024)
von: Malik, Hashmat Shadab, et al.
Veröffentlicht: (2024)
VideoMathQA: Benchmarking Mathematical Reasoning via Multimodal Understanding in Videos
von: Rasheed, Hanoona, et al.
Veröffentlicht: (2025)
von: Rasheed, Hanoona, et al.
Veröffentlicht: (2025)
DEFT: Decompositional Efficient Fine-Tuning for Text-to-Image Models
von: Kumar, Komal, et al.
Veröffentlicht: (2025)
von: Kumar, Komal, et al.
Veröffentlicht: (2025)
ELGC-Net: Efficient Local-Global Context Aggregation for Remote Sensing Change Detection
von: Noman, Mubashir, et al.
Veröffentlicht: (2024)
von: Noman, Mubashir, et al.
Veröffentlicht: (2024)
Interpretable Zero-Shot Learning with Locally-Aligned Vision-Language Model
von: Chen, Shiming, et al.
Veröffentlicht: (2025)
von: Chen, Shiming, et al.
Veröffentlicht: (2025)
Video-R2: Reinforcing Consistent and Grounded Reasoning in Multimodal Language Models
von: Maaz, Muhammad, et al.
Veröffentlicht: (2025)
von: Maaz, Muhammad, et al.
Veröffentlicht: (2025)
Learnable Weight Initialization for Volumetric Medical Image Segmentation
von: Kunhimon, Shahina, et al.
Veröffentlicht: (2023)
von: Kunhimon, Shahina, et al.
Veröffentlicht: (2023)
How Good are Foundation Models in Step-by-Step Embodied Reasoning?
von: Dissanayake, Dinura, et al.
Veröffentlicht: (2025)
von: Dissanayake, Dinura, et al.
Veröffentlicht: (2025)
GenZSL: Generative Zero-Shot Learning Via Inductive Variational Autoencoder
von: Chen, Shiming, et al.
Veröffentlicht: (2025)
von: Chen, Shiming, et al.
Veröffentlicht: (2025)
Progressive Semantic-Guided Vision Transformer for Zero-Shot Learning
von: Chen, Shiming, et al.
Veröffentlicht: (2024)
von: Chen, Shiming, et al.
Veröffentlicht: (2024)
Enhancing Novel Object Detection via Cooperative Foundational Models
von: Bharadwaj, Rohit, et al.
Veröffentlicht: (2023)
von: Bharadwaj, Rohit, et al.
Veröffentlicht: (2023)
Dual Hyperspectral Mamba for Efficient Spectral Compressive Imaging
von: Dong, Jiahua, et al.
Veröffentlicht: (2024)
von: Dong, Jiahua, et al.
Veröffentlicht: (2024)
EvoIR: Towards All-in-One Image Restoration via Evolutionary Frequency Modulation
von: Ma, Jiaqi, et al.
Veröffentlicht: (2025)
von: Ma, Jiaqi, et al.
Veröffentlicht: (2025)
MobiLlama: Towards Accurate and Lightweight Fully Transparent GPT
von: Thawakar, Omkar, et al.
Veröffentlicht: (2024)
von: Thawakar, Omkar, et al.
Veröffentlicht: (2024)
MedROV: Towards Real-Time Open-Vocabulary Detection Across Diverse Medical Imaging Modalities
von: Sheikh, Tooba Tehreem, et al.
Veröffentlicht: (2025)
von: Sheikh, Tooba Tehreem, et al.
Veröffentlicht: (2025)
DuwatBench: Bridging Language and Visual Heritage through an Arabic Calligraphy Benchmark for Multimodal Understanding
von: Patle, Shubham, et al.
Veröffentlicht: (2026)
von: Patle, Shubham, et al.
Veröffentlicht: (2026)
EarthDial: Turning Multi-sensory Earth Observations to Interactive Dialogues
von: Soni, Sagar, et al.
Veröffentlicht: (2024)
von: Soni, Sagar, et al.
Veröffentlicht: (2024)
TerraFM: A Scalable Foundation Model for Unified Multisensor Earth Observation
von: Danish, Muhammad Sohail, et al.
Veröffentlicht: (2025)
von: Danish, Muhammad Sohail, et al.
Veröffentlicht: (2025)
VURF: A General-purpose Reasoning and Self-refinement Framework for Video Understanding
von: Mahmood, Ahmad, et al.
Veröffentlicht: (2024)
von: Mahmood, Ahmad, et al.
Veröffentlicht: (2024)
VANE-Bench: Video Anomaly Evaluation Benchmark for Conversational LMMs
von: Bharadwaj, Rohit, et al.
Veröffentlicht: (2024)
von: Bharadwaj, Rohit, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
AdaIR: Adaptive All-in-One Image Restoration via Frequency Mining and Modulation
von: Cui, Yuning, et al.
Veröffentlicht: (2024) -
Vocabulary-free Fine-grained Visual Recognition via Enriched Contextually Grounded Vision-Language Model
von: Demidov, Dmitry, et al.
Veröffentlicht: (2025) -
Video-GroundingDINO: Towards Open-Vocabulary Spatio-Temporal Video Grounding
von: Wasim, Syed Talal, et al.
Veröffentlicht: (2023) -
Composed Video Retrieval via Enriched Context and Discriminative Embeddings
von: Thawakar, Omkar, et al.
Veröffentlicht: (2024) -
Beyond Simple Edits: Composed Video Retrieval with Dense Modifications
von: Thawakar, Omkar, et al.
Veröffentlicht: (2025)