MDA: An Interpretable and Scalable Multi-Modal Fusion under Missing Modalities and Intrinsic Noise Conditions
Fuente:
arXiv
Saved in:
| Main Authors: | Fan, Lin, Ou, Yafei, Zheng, Cenyang, Dai, Pengyu, Kamishima, Tamotsu, Ikebe, Masayuki, Suzuki, Kenji, Gong, Xun |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
EvoCUA: Evolving Computer Use Agents via Learning from Scalable Synthetic Experience
by: Xue, Taofeng, et al.
Published: (2026)
by: Xue, Taofeng, et al.
Published: (2026)
Tri-VQA: Triangular Reasoning Medical Visual Question Answering for Multi-Attribute Analysis
by: Fan, Lin, et al.
Published: (2024)
by: Fan, Lin, et al.
Published: (2024)
Step-CoT: Stepwise Visual Chain-of-Thought for Medical Visual Question Answering
by: Fan, Lin, et al.
Published: (2026)
by: Fan, Lin, et al.
Published: (2026)
Human-Robot Dialogue Annotation for Multi-Modal Common Ground
by: Bonial, Claire, et al.
Published: (2024)
by: Bonial, Claire, et al.
Published: (2024)
RAM-H1200: A Unified Evaluation and Dataset on Hand Radiographs for Rheumatoid Arthritis
by: Yang, Songxiao, et al.
Published: (2026)
by: Yang, Songxiao, et al.
Published: (2026)
SCOUT: A Situated and Multi-Modal Human-Robot Dialogue Corpus
by: Lukin, Stephanie M., et al.
Published: (2024)
by: Lukin, Stephanie M., et al.
Published: (2024)
Context-Aware Network Based on Multi-scale Spatio-temporal Attention for Action Recognition in Videos
by: Li, Xiaoyang, et al.
Published: (2025)
by: Li, Xiaoyang, et al.
Published: (2025)
Distinguishing Visually Similar Actions: Prompt-Guided Semantic Prototype Modulation for Few-Shot Action Recognition
by: Li, Xiaoyang, et al.
Published: (2025)
by: Li, Xiaoyang, et al.
Published: (2025)
CLIP-Joint-Detect: End-to-End Joint Training of Object Detectors with Contrastive Vision-Language Supervision
by: Raoufi, Behnam, et al.
Published: (2025)
by: Raoufi, Behnam, et al.
Published: (2025)
MSPCaps: A Multi-Scale Patchify Capsule Network with Cross-Agreement Routing for Visual Recognition
by: Hu, Yudong, et al.
Published: (2025)
by: Hu, Yudong, et al.
Published: (2025)
Few-Shot Learning of a Graph-Based Neural Network Model Without Backpropagation
by: Lapin, Mykyta, et al.
Published: (2025)
by: Lapin, Mykyta, et al.
Published: (2025)
T-Rex: Task-Adaptive Spatial Representation Extraction for Robotic Manipulation with Vision-Language Models
by: Chen, Yiteng, et al.
Published: (2025)
by: Chen, Yiteng, et al.
Published: (2025)
AntiGrounding: Lifting Robotic Actions into VLM Representation Space for Decision Making
by: Li, Wenbo, et al.
Published: (2025)
by: Li, Wenbo, et al.
Published: (2025)
Lifelong Learning in Vision-Language Models: Enhanced EWC with Cross-Modal Knowledge Retention
by: Durrani, Hamza Ahmed, et al.
Published: (2026)
by: Durrani, Hamza Ahmed, et al.
Published: (2026)
Pan-Arctic Permafrost Landform and Human-built Infrastructure Feature Detection with Vision Transformers and Location Embeddings
by: Perera, Amal S., et al.
Published: (2025)
by: Perera, Amal S., et al.
Published: (2025)
Precision at Scale: Domain-Specific Datasets On-Demand
by: Rodríguez-de-Vera, Jesús M, et al.
Published: (2024)
by: Rodríguez-de-Vera, Jesús M, et al.
Published: (2024)
CCVA-FL: Cross-Client Variations Adaptive Federated Learning for Medical Imaging
by: Gupta, Sunny, et al.
Published: (2024)
by: Gupta, Sunny, et al.
Published: (2024)
Taming the Tail: Leveraging Asymmetric Loss and Pade Approximation to Overcome Medical Image Long-Tailed Class Imbalance
by: Kashyap, Pankhi, et al.
Published: (2024)
by: Kashyap, Pankhi, et al.
Published: (2024)
ViBED-Net: Video Based Engagement Detection Network Using Face-Aware and Scene-Aware Spatiotemporal Cues
by: Gothwal, Prateek, et al.
Published: (2025)
by: Gothwal, Prateek, et al.
Published: (2025)
OrganicHAR: Towards Activity Discovery in Organic Settings for Privacy Preserving Sensors Using Efficient Video Analysis
by: Patidar, Prasoon, et al.
Published: (2026)
by: Patidar, Prasoon, et al.
Published: (2026)
treeX: Unsupervised Tree Instance Segmentation in Dense Forest Point Clouds
by: Burmeister, Josafat-Mattias, et al.
Published: (2025)
by: Burmeister, Josafat-Mattias, et al.
Published: (2025)
FerretNet: Efficient Synthetic Image Detection via Local Pixel Dependencies
by: Liang, Shuqiao, et al.
Published: (2025)
by: Liang, Shuqiao, et al.
Published: (2025)
Interpreting Structured Perturbations in Image Protection Methods for Diffusion Models
by: Martin, Michael R., et al.
Published: (2025)
by: Martin, Michael R., et al.
Published: (2025)
Spiking Neural Networks for event-based action recognition: A new task to understand their advantage
by: Vicente-Sola, Alex, et al.
Published: (2022)
by: Vicente-Sola, Alex, et al.
Published: (2022)
IMUVIE: Pickup Timeline Action Localization via Motion Movies
by: Clapham, John, et al.
Published: (2024)
by: Clapham, John, et al.
Published: (2024)
Emotions in the Loop: A Survey of Affective Computing for Emotional Support
by: Hegde, Karishma, et al.
Published: (2025)
by: Hegde, Karishma, et al.
Published: (2025)
Interpretable Machine Learning-Derived Spectral Indices for Vegetation Monitoring
by: Lotfi, Ali, et al.
Published: (2025)
by: Lotfi, Ali, et al.
Published: (2025)
Look and Tell: A Dataset for Multimodal Grounding Across Egocentric and Exocentric Views
by: Deichler, Anna, et al.
Published: (2025)
by: Deichler, Anna, et al.
Published: (2025)
Yanyun-3: Enabling Cross-Platform Strategy Game Operation with Vision-Language Models
by: Wang, Guoyan, et al.
Published: (2025)
by: Wang, Guoyan, et al.
Published: (2025)
CourseTimeQA: A Lecture-Video Benchmark and a Latency-Constrained Cross-Modal Fusion Method for Timestamped QA
by: Kovalev, Vsevolod, et al.
Published: (2025)
by: Kovalev, Vsevolod, et al.
Published: (2025)
FlyMeThrough: Human-AI Collaborative 3D Indoor Mapping with Commodity Drones
by: Su, Xia, et al.
Published: (2025)
by: Su, Xia, et al.
Published: (2025)
MM-Conv: A Multimodal Dataset and Benchmark for Context-Aware Grounding in 3D Dialogue
by: Deichler, Anna, et al.
Published: (2026)
by: Deichler, Anna, et al.
Published: (2026)
ExpressNet-MoE: A Hybrid Deep Neural Network for Emotion Recognition
by: Banerjee, Deeptimaan, et al.
Published: (2025)
by: Banerjee, Deeptimaan, et al.
Published: (2025)
WaveMix: A Resource-efficient Neural Network for Image Analysis
by: Jeevan, Pranav, et al.
Published: (2022)
by: Jeevan, Pranav, et al.
Published: (2022)
Which Backbone to Use: A Resource-efficient Domain Specific Comparison for Computer Vision
by: Jeevan, Pranav, et al.
Published: (2024)
by: Jeevan, Pranav, et al.
Published: (2024)
Playing telephone with generative models: "verification disability," "compelled reliance," and accessibility in data visualization
by: Elavsky, Frank, et al.
Published: (2025)
by: Elavsky, Frank, et al.
Published: (2025)
Introspection in Learned Semantic Scene Graph Localisation
by: Bissessur, Manshika Charvi, et al.
Published: (2025)
by: Bissessur, Manshika Charvi, et al.
Published: (2025)
How Can One Choose the Best CAM-Based Explainability Method for a CNN Model?
by: Costa, Daniel da Silva, et al.
Published: (2026)
by: Costa, Daniel da Silva, et al.
Published: (2026)
Key-Scan-Based Mobile Robot Navigation: Integrated Mapping, Planning, and Control using Graphs of Scan Regions
by: Latha, Dharshan Bashkaran, et al.
Published: (2024)
by: Latha, Dharshan Bashkaran, et al.
Published: (2024)
OpenMap: Instruction Grounding via Open-Vocabulary Visual-Language Mapping
by: Li, Danyang, et al.
Published: (2025)
by: Li, Danyang, et al.
Published: (2025)
Similar Items
-
EvoCUA: Evolving Computer Use Agents via Learning from Scalable Synthetic Experience
by: Xue, Taofeng, et al.
Published: (2026) -
Tri-VQA: Triangular Reasoning Medical Visual Question Answering for Multi-Attribute Analysis
by: Fan, Lin, et al.
Published: (2024) -
Step-CoT: Stepwise Visual Chain-of-Thought for Medical Visual Question Answering
by: Fan, Lin, et al.
Published: (2026) -
Human-Robot Dialogue Annotation for Multi-Modal Common Ground
by: Bonial, Claire, et al.
Published: (2024) -
RAM-H1200: A Unified Evaluation and Dataset on Hand Radiographs for Rheumatoid Arthritis
by: Yang, Songxiao, et al.
Published: (2026)