Harnessing Shared Relations via Multimodal Mixup Contrastive Learning for Multimodal Classification
Fuente:
arXiv
Guardado en:
| Autores principales: | Kumar, Raja, Singhal, Raghav, Kulkarni, Pranamya, Mehta, Deval, Jadhav, Kshitij |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
MIRAGE: Multimodal Identification and Recognition of Annotations in Indian General Prescriptions
por: Mankash, Tavish, et al.
Publicado: (2024)
por: Mankash, Tavish, et al.
Publicado: (2024)
Interpretable Few-Shot Retinal Disease Diagnosis with Concept-Guided Prompting of Vision-Language Models
por: Mehta, Deval, et al.
Publicado: (2025)
por: Mehta, Deval, et al.
Publicado: (2025)
One Shot GANs for Long Tail Problem in Skin Lesion Dataset using novel content space assessment metric
por: Deo, Kunal, et al.
Publicado: (2024)
por: Deo, Kunal, et al.
Publicado: (2024)
Efficient Learning for Product Attributes with Compact Multimodal Models
por: Kulkarni, Mandar
Publicado: (2025)
por: Kulkarni, Mandar
Publicado: (2025)
Neurosymbolic Framework for Concept-Driven Logical Reasoning in Skeleton-Based Human Action Recognition
por: Ilyas, Talha, et al.
Publicado: (2026)
por: Ilyas, Talha, et al.
Publicado: (2026)
XBusNet: Text-Guided Breast Ultrasound Segmentation via Multimodal Vision-Language Learning
por: Mallina, Raja, et al.
Publicado: (2025)
por: Mallina, Raja, et al.
Publicado: (2025)
MCN-CL: Multimodal Cross-Attention Network and Contrastive Learning for Multimodal Emotion Recognition
por: Li, Feng, et al.
Publicado: (2025)
por: Li, Feng, et al.
Publicado: (2025)
Understanding and Harnessing Sparsity in Unified Multimodal Models
por: He, Shwai, et al.
Publicado: (2025)
por: He, Shwai, et al.
Publicado: (2025)
Pseudo Contrastive Learning for Diagram Comprehension in Multimodal Models
por: Sasaki, Hiroshi
Publicado: (2026)
por: Sasaki, Hiroshi
Publicado: (2026)
SONIC-O1: A Real-World Benchmark for Evaluating Multimodal Large Language Models on Audio-Video Understanding
por: Radwan, Ahmed Y., et al.
Publicado: (2026)
por: Radwan, Ahmed Y., et al.
Publicado: (2026)
CtrlSynth: Controllable Image Text Synthesis for Data-Efficient Multimodal Learning
por: Cao, Qingqing, et al.
Publicado: (2024)
por: Cao, Qingqing, et al.
Publicado: (2024)
Multimodal Medical Image Classification via Synergistic Learning Pre-training
por: Lin, Qinghua, et al.
Publicado: (2025)
por: Lin, Qinghua, et al.
Publicado: (2025)
Robust Defense Strategies for Multimodal Contrastive Learning: Efficient Fine-tuning Against Backdoor Attacks
por: Hossain, Md. Iqbal, et al.
Publicado: (2025)
por: Hossain, Md. Iqbal, et al.
Publicado: (2025)
FESS Loss: Feature-Enhanced Spatial Segmentation Loss for Optimizing Medical Image Analysis
por: Chodvadiya, Charulkumar, et al.
Publicado: (2024)
por: Chodvadiya, Charulkumar, et al.
Publicado: (2024)
Test-Time Mixup Augmentation for Data and Class-Specific Uncertainty Estimation in Deep Learning Image Classification
por: Lee, Hansang, et al.
Publicado: (2022)
por: Lee, Hansang, et al.
Publicado: (2022)
SupReMix: Supervised Contrastive Learning for Medical Imaging Regression with Mixup
por: Wu, Yilei, et al.
Publicado: (2023)
por: Wu, Yilei, et al.
Publicado: (2023)
HumaniBench: A Human-Centric Framework for Large Multimodal Models Evaluation
por: Raza, Shaina, et al.
Publicado: (2025)
por: Raza, Shaina, et al.
Publicado: (2025)
NULLBUS: Multimodal Mixed-Supervision for Breast Ultrasound Segmentation via Nullable Global-Local Prompts
por: Mallina, Raja, et al.
Publicado: (2025)
por: Mallina, Raja, et al.
Publicado: (2025)
Automatic dataset shift identification to support safe deployment of medical imaging AI
por: Roschewitz, Mélanie, et al.
Publicado: (2024)
por: Roschewitz, Mélanie, et al.
Publicado: (2024)
Text-VQA Aug: Pipelined Harnessing of Large Multimodal Models for Automated Synthesis
por: Joshi, Soham, et al.
Publicado: (2025)
por: Joshi, Soham, et al.
Publicado: (2025)
Enhancing Multimodal In-Context Learning for Image Classification through Coreset Optimization
por: Chen, Huiyi, et al.
Publicado: (2025)
por: Chen, Huiyi, et al.
Publicado: (2025)
Explaining How Visual, Textual and Multimodal Encoders Share Concepts
por: Cornet, Clément, et al.
Publicado: (2025)
por: Cornet, Clément, et al.
Publicado: (2025)
Multimodal Connectome Fusion via Cross-Attention for Autism Spectrum Disorder Classification Using Graph Learning
por: Rahman, Ansar, et al.
Publicado: (2026)
por: Rahman, Ansar, et al.
Publicado: (2026)
MedSAGa: Few-shot Memory Efficient Medical Image Segmentation using Gradient Low-Rank Projection in SAM
por: Mahla, Navyansh, et al.
Publicado: (2024)
por: Mahla, Navyansh, et al.
Publicado: (2024)
Robust Multimodal Learning via Representation Decoupling
por: Wei, Shicai, et al.
Publicado: (2024)
por: Wei, Shicai, et al.
Publicado: (2024)
Harnessing PDF Data for Improving Japanese Large Multimodal Models
por: Baek, Jeonghun, et al.
Publicado: (2025)
por: Baek, Jeonghun, et al.
Publicado: (2025)
ConVQG: Contrastive Visual Question Generation with Multimodal Guidance
por: Mi, Li, et al.
Publicado: (2024)
por: Mi, Li, et al.
Publicado: (2024)
Multimodal Contrastive Pretraining of CBCT and IOS for Enhanced Tooth Segmentation
por: Son, Moo Hyun, et al.
Publicado: (2025)
por: Son, Moo Hyun, et al.
Publicado: (2025)
The More, the Merrier: Contrastive Fusion for Higher-Order Multimodal Alignment
por: Koutoupis, Stefanos, et al.
Publicado: (2025)
por: Koutoupis, Stefanos, et al.
Publicado: (2025)
Enhancing Generalization in Data-free Quantization via Mixup-class Prompting
por: Park, Jiwoong, et al.
Publicado: (2025)
por: Park, Jiwoong, et al.
Publicado: (2025)
Explaining and Mitigating the Modality Gap in Contrastive Multimodal Learning
por: Yaras, Can, et al.
Publicado: (2024)
por: Yaras, Can, et al.
Publicado: (2024)
Multimodal Deep Learning for Phyllodes Tumor Classification from Ultrasound and Clinical Data
por: Abir, Farhan Fuad, et al.
Publicado: (2025)
por: Abir, Farhan Fuad, et al.
Publicado: (2025)
R1-ShareVL: Incentivizing Reasoning Capability of Multimodal Large Language Models via Share-GRPO
por: Yao, Huanjin, et al.
Publicado: (2025)
por: Yao, Huanjin, et al.
Publicado: (2025)
IMPROVE: Improving Medical Plausibility without Reliance on HumanValidation -- An Enhanced Prototype-Guided Diffusion Framework
por: Shandilya, Anurag, et al.
Publicado: (2024)
por: Shandilya, Anurag, et al.
Publicado: (2024)
Reconstruction-Driven Multimodal Representation Learning for Automated Media Understanding
por: Benhammou, Yassir, et al.
Publicado: (2025)
por: Benhammou, Yassir, et al.
Publicado: (2025)
DeepInsert: Early Layer Bypass for Efficient and Performant Multimodal Understanding
por: Choraria, Moulik, et al.
Publicado: (2025)
por: Choraria, Moulik, et al.
Publicado: (2025)
PQV-Mobile: A Combined Pruning and Quantization Toolkit to Optimize Vision Transformers for Mobile Applications
por: Bhardwaj, Kshitij
Publicado: (2024)
por: Bhardwaj, Kshitij
Publicado: (2024)
CL3DOR: Contrastive Learning for 3D Large Multimodal Models via Odds Ratio on High-Resolution Point Clouds
por: Kim, Keonwoo, et al.
Publicado: (2025)
por: Kim, Keonwoo, et al.
Publicado: (2025)
Img-Diff: Contrastive Data Synthesis for Multimodal Large Language Models
por: Jiao, Qirui, et al.
Publicado: (2024)
por: Jiao, Qirui, et al.
Publicado: (2024)
Boosting the Transferability of Adversarial Examples via Local Mixup and Adaptive Step Size
por: Liu, Junlin, et al.
Publicado: (2024)
por: Liu, Junlin, et al.
Publicado: (2024)
Ejemplares similares
-
MIRAGE: Multimodal Identification and Recognition of Annotations in Indian General Prescriptions
por: Mankash, Tavish, et al.
Publicado: (2024) -
Interpretable Few-Shot Retinal Disease Diagnosis with Concept-Guided Prompting of Vision-Language Models
por: Mehta, Deval, et al.
Publicado: (2025) -
One Shot GANs for Long Tail Problem in Skin Lesion Dataset using novel content space assessment metric
por: Deo, Kunal, et al.
Publicado: (2024) -
Efficient Learning for Product Attributes with Compact Multimodal Models
por: Kulkarni, Mandar
Publicado: (2025) -
Neurosymbolic Framework for Concept-Driven Logical Reasoning in Skeleton-Based Human Action Recognition
por: Ilyas, Talha, et al.
Publicado: (2026)