Simultaneous Long-tailed Recognition and Multi-modal Fusion for Highly Imbalanced Multi-modal Data
Fuente:
arXiv
Saved in:
| Main Authors: | Yoon, Heegeon, Kim, Heeyoung |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Multimodal Deep Generative Model for Semi-Supervised Learning under Class Imbalance
by: Yoon, Heegeon, et al.
Published: (2026)
by: Yoon, Heegeon, et al.
Published: (2026)
Multi-modal Co-learning for Earth Observation: Enhancing single-modality models via modality collaboration
by: Mena, Francisco, et al.
Published: (2025)
by: Mena, Francisco, et al.
Published: (2025)
Actor-agnostic Multi-label Action Recognition with Multi-modal Query
by: Mondal, Anindya, et al.
Published: (2023)
by: Mondal, Anindya, et al.
Published: (2023)
Visual Hallucinations of Multi-modal Large Language Models
by: Huang, Wen, et al.
Published: (2024)
by: Huang, Wen, et al.
Published: (2024)
Rationale-Enhanced Decoding for Multi-modal Chain-of-Thought
by: Yamaguchi, Shin'ya, et al.
Published: (2025)
by: Yamaguchi, Shin'ya, et al.
Published: (2025)
Difficulty-aware Balancing Margin Loss for Long-tailed Recognition
by: Son, Minseok, et al.
Published: (2024)
by: Son, Minseok, et al.
Published: (2024)
Mitigating Visual Forgetting via Take-along Visual Conditioning for Multi-modal Long CoT Reasoning
by: Sun, Hai-Long, et al.
Published: (2025)
by: Sun, Hai-Long, et al.
Published: (2025)
Generative Multi-modal Models are Good Class-Incremental Learners
by: Cao, Xusheng, et al.
Published: (2024)
by: Cao, Xusheng, et al.
Published: (2024)
Multi-modal Vision Pre-training for Medical Image Analysis
by: Rui, Shaohao, et al.
Published: (2024)
by: Rui, Shaohao, et al.
Published: (2024)
Multi-modal Machine Learning for Vehicle Rating Predictions Using Image, Text, and Parametric Data
by: Su, Hanqi, et al.
Published: (2023)
by: Su, Hanqi, et al.
Published: (2023)
Analyzing and Boosting the Power of Fine-Grained Visual Recognition for Multi-modal Large Language Models
by: He, Hulingxiao, et al.
Published: (2025)
by: He, Hulingxiao, et al.
Published: (2025)
Skin Lesion Phenotyping via Nested Multi-modal Contrastive Learning
by: Christopoulos, Dionysis, et al.
Published: (2025)
by: Christopoulos, Dionysis, et al.
Published: (2025)
MOCHA: Multi-modal Objects-aware Cross-arcHitecture Alignment
by: Camuffo, Elena, et al.
Published: (2025)
by: Camuffo, Elena, et al.
Published: (2025)
GenSim2: Scaling Robot Data Generation with Multi-modal and Reasoning LLMs
by: Hua, Pu, et al.
Published: (2024)
by: Hua, Pu, et al.
Published: (2024)
VIAssist: Adapting Multi-modal Large Language Models for Users with Visual Impairments
by: Yang, Bufang, et al.
Published: (2024)
by: Yang, Bufang, et al.
Published: (2024)
CNC: Cross-modal Normality Constraint for Unsupervised Multi-class Anomaly Detection
by: Wang, Xiaolei, et al.
Published: (2024)
by: Wang, Xiaolei, et al.
Published: (2024)
Knowledge Graph Enhanced Generative Multi-modal Models for Class-Incremental Learning
by: Cao, Xusheng, et al.
Published: (2025)
by: Cao, Xusheng, et al.
Published: (2025)
MuMA-ToM: Multi-modal Multi-Agent Theory of Mind
by: Shi, Haojun, et al.
Published: (2024)
by: Shi, Haojun, et al.
Published: (2024)
MAVEN: Multi-modal Attention for Valence-Arousal Emotion Network
by: Ahire, Vrushank, et al.
Published: (2025)
by: Ahire, Vrushank, et al.
Published: (2025)
The Labyrinth of Links: Navigating the Associative Maze of Multi-modal LLMs
by: Li, Hong, et al.
Published: (2024)
by: Li, Hong, et al.
Published: (2024)
Multi-modal Masked Siamese Network Improves Chest X-Ray Representation Learning
by: Shurrab, Saeed, et al.
Published: (2024)
by: Shurrab, Saeed, et al.
Published: (2024)
Preserving Pre-trained Representation Space: On Effectiveness of Prefix-tuning for Large Multi-modal Models
by: Kim, Donghoon, et al.
Published: (2024)
by: Kim, Donghoon, et al.
Published: (2024)
SPHINX-X: Scaling Data and Parameters for a Family of Multi-modal Large Language Models
by: Liu, Dongyang, et al.
Published: (2024)
by: Liu, Dongyang, et al.
Published: (2024)
GET: Unlocking the Multi-modal Potential of CLIP for Generalized Category Discovery
by: Wang, Enguang, et al.
Published: (2024)
by: Wang, Enguang, et al.
Published: (2024)
Hierarchical Multi-modal Transformer for Cross-modal Long Document Classification
by: Liu, Tengfei, et al.
Published: (2024)
by: Liu, Tengfei, et al.
Published: (2024)
Refusing Safe Prompts for Multi-modal Large Language Models
by: Shao, Zedian, et al.
Published: (2024)
by: Shao, Zedian, et al.
Published: (2024)
MMCTAgent: Multi-modal Critical Thinking Agent Framework for Complex Visual Reasoning
by: Kumar, Somnath, et al.
Published: (2024)
by: Kumar, Somnath, et al.
Published: (2024)
R2GenKG: Hierarchical Multi-modal Knowledge Graph for LLM-based Radiology Report Generation
by: Wang, Futian, et al.
Published: (2025)
by: Wang, Futian, et al.
Published: (2025)
Multi-modal Data Spectrum: Multi-modal Datasets are Multi-dimensional
by: Madaan, Divyam, et al.
Published: (2025)
by: Madaan, Divyam, et al.
Published: (2025)
Multi-modal Preference Alignment Remedies Degradation of Visual Instruction Tuning on Language Models
by: Li, Shengzhi, et al.
Published: (2024)
by: Li, Shengzhi, et al.
Published: (2024)
SurvMamba: State Space Model with Multi-grained Multi-modal Interaction for Survival Prediction
by: Chen, Ying, et al.
Published: (2024)
by: Chen, Ying, et al.
Published: (2024)
DirMixE: Harnessing Test Agnostic Long-tail Recognition with Hierarchical Label Vartiations
by: Yang, Zhiyong, et al.
Published: (2024)
by: Yang, Zhiyong, et al.
Published: (2024)
Robust Domain Generalization for Multi-modal Object Recognition
by: Qiao, Yuxin, et al.
Published: (2024)
by: Qiao, Yuxin, et al.
Published: (2024)
An Enhanced Classification Method Based on Adaptive Multi-Scale Fusion for Long-tailed Multispectral Point Clouds
by: Liu, TianZhu, et al.
Published: (2024)
by: Liu, TianZhu, et al.
Published: (2024)
From Consistency to Complementarity: Aligned and Disentangled Multi-modal Learning for Time Series Understanding and Reasoning
by: Ni, Hang, et al.
Published: (2026)
by: Ni, Hang, et al.
Published: (2026)
Countering Multi-modal Representation Collapse through Rank-targeted Fusion
by: Kim, Seulgi, et al.
Published: (2025)
by: Kim, Seulgi, et al.
Published: (2025)
MultiWay-Adapater: Adapting large-scale multi-modal models for scalable image-text retrieval
by: Long, Zijun, et al.
Published: (2023)
by: Long, Zijun, et al.
Published: (2023)
Multi-modal Generative AI: Multi-modal LLMs, Diffusions, and the Unification
by: Wang, Xin, et al.
Published: (2024)
by: Wang, Xin, et al.
Published: (2024)
MultiFloodSynth: Multi-Annotated Flood Synthetic Dataset Generation
by: Kang, YoonJe, et al.
Published: (2025)
by: Kang, YoonJe, et al.
Published: (2025)
MathVerse: Does Your Multi-modal LLM Truly See the Diagrams in Visual Math Problems?
by: Zhang, Renrui, et al.
Published: (2024)
by: Zhang, Renrui, et al.
Published: (2024)
Similar Items
-
Multimodal Deep Generative Model for Semi-Supervised Learning under Class Imbalance
by: Yoon, Heegeon, et al.
Published: (2026) -
Multi-modal Co-learning for Earth Observation: Enhancing single-modality models via modality collaboration
by: Mena, Francisco, et al.
Published: (2025) -
Actor-agnostic Multi-label Action Recognition with Multi-modal Query
by: Mondal, Anindya, et al.
Published: (2023) -
Visual Hallucinations of Multi-modal Large Language Models
by: Huang, Wen, et al.
Published: (2024) -
Rationale-Enhanced Decoding for Multi-modal Chain-of-Thought
by: Yamaguchi, Shin'ya, et al.
Published: (2025)