Gespeichert in:
| Hauptverfasser: | Kumar, Raja, Singhal, Raghav, Kulkarni, Pranamya, Mehta, Deval, Jadhav, Kshitij |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | https://arxiv.org/abs/2409.17777 |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Interpretable Few-Shot Retinal Disease Diagnosis with Concept-Guided Prompting of Vision-Language Models
von: Mehta, Deval, et al.
Veröffentlicht: (2025)
von: Mehta, Deval, et al.
Veröffentlicht: (2025)
MIRAGE: Multimodal Identification and Recognition of Annotations in Indian General Prescriptions
von: Mankash, Tavish, et al.
Veröffentlicht: (2024)
von: Mankash, Tavish, et al.
Veröffentlicht: (2024)
One Shot GANs for Long Tail Problem in Skin Lesion Dataset using novel content space assessment metric
von: Deo, Kunal, et al.
Veröffentlicht: (2024)
von: Deo, Kunal, et al.
Veröffentlicht: (2024)
Neurosymbolic Framework for Concept-Driven Logical Reasoning in Skeleton-Based Human Action Recognition
von: Ilyas, Talha, et al.
Veröffentlicht: (2026)
von: Ilyas, Talha, et al.
Veröffentlicht: (2026)
Efficient Learning for Product Attributes with Compact Multimodal Models
von: Kulkarni, Mandar
Veröffentlicht: (2025)
von: Kulkarni, Mandar
Veröffentlicht: (2025)
XBusNet: Text-Guided Breast Ultrasound Segmentation via Multimodal Vision-Language Learning
von: Mallina, Raja, et al.
Veröffentlicht: (2025)
von: Mallina, Raja, et al.
Veröffentlicht: (2025)
SONIC-O1: A Real-World Benchmark for Evaluating Multimodal Large Language Models on Audio-Video Understanding
von: Radwan, Ahmed Y., et al.
Veröffentlicht: (2026)
von: Radwan, Ahmed Y., et al.
Veröffentlicht: (2026)
FESS Loss: Feature-Enhanced Spatial Segmentation Loss for Optimizing Medical Image Analysis
von: Chodvadiya, Charulkumar, et al.
Veröffentlicht: (2024)
von: Chodvadiya, Charulkumar, et al.
Veröffentlicht: (2024)
MCN-CL: Multimodal Cross-Attention Network and Contrastive Learning for Multimodal Emotion Recognition
von: Li, Feng, et al.
Veröffentlicht: (2025)
von: Li, Feng, et al.
Veröffentlicht: (2025)
Understanding and Harnessing Sparsity in Unified Multimodal Models
von: He, Shwai, et al.
Veröffentlicht: (2025)
von: He, Shwai, et al.
Veröffentlicht: (2025)
CtrlSynth: Controllable Image Text Synthesis for Data-Efficient Multimodal Learning
von: Cao, Qingqing, et al.
Veröffentlicht: (2024)
von: Cao, Qingqing, et al.
Veröffentlicht: (2024)
Pseudo Contrastive Learning for Diagram Comprehension in Multimodal Models
von: Sasaki, Hiroshi
Veröffentlicht: (2026)
von: Sasaki, Hiroshi
Veröffentlicht: (2026)
Multimodal Medical Image Classification via Synergistic Learning Pre-training
von: Lin, Qinghua, et al.
Veröffentlicht: (2025)
von: Lin, Qinghua, et al.
Veröffentlicht: (2025)
Robust Defense Strategies for Multimodal Contrastive Learning: Efficient Fine-tuning Against Backdoor Attacks
von: Hossain, Md. Iqbal, et al.
Veröffentlicht: (2025)
von: Hossain, Md. Iqbal, et al.
Veröffentlicht: (2025)
SupReMix: Supervised Contrastive Learning for Medical Imaging Regression with Mixup
von: Wu, Yilei, et al.
Veröffentlicht: (2023)
von: Wu, Yilei, et al.
Veröffentlicht: (2023)
IMPROVE: Improving Medical Plausibility without Reliance on HumanValidation -- An Enhanced Prototype-Guided Diffusion Framework
von: Shandilya, Anurag, et al.
Veröffentlicht: (2024)
von: Shandilya, Anurag, et al.
Veröffentlicht: (2024)
HumaniBench: A Human-Centric Framework for Large Multimodal Models Evaluation
von: Raza, Shaina, et al.
Veröffentlicht: (2025)
von: Raza, Shaina, et al.
Veröffentlicht: (2025)
MedSAGa: Few-shot Memory Efficient Medical Image Segmentation using Gradient Low-Rank Projection in SAM
von: Mahla, Navyansh, et al.
Veröffentlicht: (2024)
von: Mahla, Navyansh, et al.
Veröffentlicht: (2024)
Automatic dataset shift identification to support safe deployment of medical imaging AI
von: Roschewitz, Mélanie, et al.
Veröffentlicht: (2024)
von: Roschewitz, Mélanie, et al.
Veröffentlicht: (2024)
Test-Time Mixup Augmentation for Data and Class-Specific Uncertainty Estimation in Deep Learning Image Classification
von: Lee, Hansang, et al.
Veröffentlicht: (2022)
von: Lee, Hansang, et al.
Veröffentlicht: (2022)
NULLBUS: Multimodal Mixed-Supervision for Breast Ultrasound Segmentation via Nullable Global-Local Prompts
von: Mallina, Raja, et al.
Veröffentlicht: (2025)
von: Mallina, Raja, et al.
Veröffentlicht: (2025)
Text-VQA Aug: Pipelined Harnessing of Large Multimodal Models for Automated Synthesis
von: Joshi, Soham, et al.
Veröffentlicht: (2025)
von: Joshi, Soham, et al.
Veröffentlicht: (2025)
Harnessing PDF Data for Improving Japanese Large Multimodal Models
von: Baek, Jeonghun, et al.
Veröffentlicht: (2025)
von: Baek, Jeonghun, et al.
Veröffentlicht: (2025)
Explaining How Visual, Textual and Multimodal Encoders Share Concepts
von: Cornet, Clément, et al.
Veröffentlicht: (2025)
von: Cornet, Clément, et al.
Veröffentlicht: (2025)
Explaining and Mitigating the Modality Gap in Contrastive Multimodal Learning
von: Yaras, Can, et al.
Veröffentlicht: (2024)
von: Yaras, Can, et al.
Veröffentlicht: (2024)
Where are we with calibration under dataset shift in image classification?
von: Roschewitz, Mélanie, et al.
Veröffentlicht: (2025)
von: Roschewitz, Mélanie, et al.
Veröffentlicht: (2025)
Enhancing Multimodal In-Context Learning for Image Classification through Coreset Optimization
von: Chen, Huiyi, et al.
Veröffentlicht: (2025)
von: Chen, Huiyi, et al.
Veröffentlicht: (2025)
DeepInsert: Early Layer Bypass for Efficient and Performant Multimodal Understanding
von: Choraria, Moulik, et al.
Veröffentlicht: (2025)
von: Choraria, Moulik, et al.
Veröffentlicht: (2025)
Multimodal Connectome Fusion via Cross-Attention for Autism Spectrum Disorder Classification Using Graph Learning
von: Rahman, Ansar, et al.
Veröffentlicht: (2026)
von: Rahman, Ansar, et al.
Veröffentlicht: (2026)
PQV-Mobile: A Combined Pruning and Quantization Toolkit to Optimize Vision Transformers for Mobile Applications
von: Bhardwaj, Kshitij
Veröffentlicht: (2024)
von: Bhardwaj, Kshitij
Veröffentlicht: (2024)
R1-ShareVL: Incentivizing Reasoning Capability of Multimodal Large Language Models via Share-GRPO
von: Yao, Huanjin, et al.
Veröffentlicht: (2025)
von: Yao, Huanjin, et al.
Veröffentlicht: (2025)
Enhancing Generalization in Data-free Quantization via Mixup-class Prompting
von: Park, Jiwoong, et al.
Veröffentlicht: (2025)
von: Park, Jiwoong, et al.
Veröffentlicht: (2025)
ConVQG: Contrastive Visual Question Generation with Multimodal Guidance
von: Mi, Li, et al.
Veröffentlicht: (2024)
von: Mi, Li, et al.
Veröffentlicht: (2024)
Multimodal Contrastive Pretraining of CBCT and IOS for Enhanced Tooth Segmentation
von: Son, Moo Hyun, et al.
Veröffentlicht: (2025)
von: Son, Moo Hyun, et al.
Veröffentlicht: (2025)
The More, the Merrier: Contrastive Fusion for Higher-Order Multimodal Alignment
von: Koutoupis, Stefanos, et al.
Veröffentlicht: (2025)
von: Koutoupis, Stefanos, et al.
Veröffentlicht: (2025)
Robust Multimodal Learning via Representation Decoupling
von: Wei, Shicai, et al.
Veröffentlicht: (2024)
von: Wei, Shicai, et al.
Veröffentlicht: (2024)
Designing Production-Scale OCR for India: Multilingual and Domain-Specific Systems
von: Faraz, Ali, et al.
Veröffentlicht: (2026)
von: Faraz, Ali, et al.
Veröffentlicht: (2026)
PatrolVision: Automated License Plate Recognition in the wild
von: Singhal, Anmol Singhal Navya
Veröffentlicht: (2025)
von: Singhal, Anmol Singhal Navya
Veröffentlicht: (2025)
Cross-Domain Few-Shot Learning for Hyperspectral Image Classification Based on Mixup Foundation Model
von: Paeedeh, Naeem, et al.
Veröffentlicht: (2026)
von: Paeedeh, Naeem, et al.
Veröffentlicht: (2026)
Multimodal Deep Learning for Phyllodes Tumor Classification from Ultrasound and Clinical Data
von: Abir, Farhan Fuad, et al.
Veröffentlicht: (2025)
von: Abir, Farhan Fuad, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Interpretable Few-Shot Retinal Disease Diagnosis with Concept-Guided Prompting of Vision-Language Models
von: Mehta, Deval, et al.
Veröffentlicht: (2025) -
MIRAGE: Multimodal Identification and Recognition of Annotations in Indian General Prescriptions
von: Mankash, Tavish, et al.
Veröffentlicht: (2024) -
One Shot GANs for Long Tail Problem in Skin Lesion Dataset using novel content space assessment metric
von: Deo, Kunal, et al.
Veröffentlicht: (2024) -
Neurosymbolic Framework for Concept-Driven Logical Reasoning in Skeleton-Based Human Action Recognition
von: Ilyas, Talha, et al.
Veröffentlicht: (2026) -
Efficient Learning for Product Attributes with Compact Multimodal Models
von: Kulkarni, Mandar
Veröffentlicht: (2025)