Feature Alignment Determines Fusion Strategy: A Comparative Study of Cross-Attention and Concatenation in Multimodal Learning
Fuente:
arXiv
Saved in:
| Main Authors: | Zhou, Zhiqiang, Xie, Xuezhen |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Deep Learning-Based Approach for Identification of Potato Leaf Diseases Using Wrapper Feature Selection and Feature Concatenation
by: Naeem, Muhammad Ahtsam, et al.
Published: (2025)
by: Naeem, Muhammad Ahtsam, et al.
Published: (2025)
DashFusion: Dual-stream Alignment with Hierarchical Bottleneck Fusion for Multimodal Sentiment Analysis
by: Wen, Yuhua, et al.
Published: (2025)
by: Wen, Yuhua, et al.
Published: (2025)
Multimodal Fusion Strategies for Mapping Biophysical Landscape Features
by: Gordon, Lucia, et al.
Published: (2024)
by: Gordon, Lucia, et al.
Published: (2024)
Feature, Alignment, and Supervision in Category Learning: A Comparative Approach with Children and Neural Networks
by: Qiu, Fanxiao Wani, et al.
Published: (2026)
by: Qiu, Fanxiao Wani, et al.
Published: (2026)
CrossFuse: Learning Infrared and Visible Image Fusion by Cross-Sensor Top-K Vision Alignment and Beyond
by: Shi, Yukai, et al.
Published: (2025)
by: Shi, Yukai, et al.
Published: (2025)
Joint Attention-Guided Feature Fusion Network for Saliency Detection of Surface Defects
by: Jiang, Xiaoheng, et al.
Published: (2024)
by: Jiang, Xiaoheng, et al.
Published: (2024)
DAGLFNet: Deep Feature Attention Guided Global and Local Feature Fusion for Pseudo-Image Point Cloud Segmentation
by: Chen, Chuang, et al.
Published: (2025)
by: Chen, Chuang, et al.
Published: (2025)
DecoratingFusion: A LiDAR-Camera Fusion Network with the Combination of Point-level and Feature-level Fusion
by: Yin, Zixuan, et al.
Published: (2024)
by: Yin, Zixuan, et al.
Published: (2024)
Gated-Attention Feature-Fusion Based Framework for Poverty Prediction
by: Ramzan, Muhammad Umer, et al.
Published: (2024)
by: Ramzan, Muhammad Umer, et al.
Published: (2024)
MultiST: A Cross-Attention-Based Multimodal Model for Spatial Transcriptomic
by: Wang, Wei, et al.
Published: (2026)
by: Wang, Wei, et al.
Published: (2026)
ELMF4EggQ: Ensemble Learning with Multimodal Feature Fusion for Non-Destructive Egg Quality Assessment
by: Hassan, Md Zahim, et al.
Published: (2025)
by: Hassan, Md Zahim, et al.
Published: (2025)
Bayesian Cross-Modal Alignment Learning for Few-Shot Out-of-Distribution Generalization
by: Zhu, Lin, et al.
Published: (2025)
by: Zhu, Lin, et al.
Published: (2025)
Multi-Layer Feature Fusion with Cross-Channel Attention-Based U-Net for Kidney Tumor Segmentation
by: Neha, Fnu, et al.
Published: (2024)
by: Neha, Fnu, et al.
Published: (2024)
UrbanFusion: Stochastic Multimodal Fusion for Contrastive Learning of Robust Spatial Representations
by: Mühlematter, Dominik J., et al.
Published: (2025)
by: Mühlematter, Dominik J., et al.
Published: (2025)
Layer as Puzzle Pieces: Compressing Large Language Models through Layer Concatenation
by: Wang, Fei, et al.
Published: (2025)
by: Wang, Fei, et al.
Published: (2025)
Cross-Dataset Gaze Estimation by Evidential Inter-intra Fusion
by: Wang, Shijing, et al.
Published: (2024)
by: Wang, Shijing, et al.
Published: (2024)
Cross-Attention Multimodal Fusion for Breast Cancer Diagnosis: Integrating Mammography and Clinical Data with Explainability
by: Nantogmah, Muhaisin Tiyumba, et al.
Published: (2025)
by: Nantogmah, Muhaisin Tiyumba, et al.
Published: (2025)
Fusion Complexity Inversion: Why Simpler Cross View Modules Outperform SSMs and Cross View Attention Transformers for Pasture Biomass Regression
by: Mandal, Mridankan
Published: (2026)
by: Mandal, Mridankan
Published: (2026)
Cross-Modal Binary Attention: An Energy-Efficient Fusion Framework for Audio-Visual Learning
by: Saleh, Mohamed, et al.
Published: (2026)
by: Saleh, Mohamed, et al.
Published: (2026)
Robust Multi-View Learning via Representation Fusion of Sample-Level Attention and Alignment of Simulated Perturbation
by: Xu, Jie, et al.
Published: (2025)
by: Xu, Jie, et al.
Published: (2025)
ConceptAttention: Diffusion Transformers Learn Highly Interpretable Features
by: Helbling, Alec, et al.
Published: (2025)
by: Helbling, Alec, et al.
Published: (2025)
MolFM-Lite: Multi-Modal Molecular Property Prediction with Conformer Ensemble Attention and Cross-Modal Fusion
by: Shah, Syed Omer, et al.
Published: (2026)
by: Shah, Syed Omer, et al.
Published: (2026)
Alignment of Diffusion Models: Fundamentals, Challenges, and Future
by: Liu, Buhua, et al.
Published: (2024)
by: Liu, Buhua, et al.
Published: (2024)
Golden Noise for Diffusion Models: A Learning Framework
by: Zhou, Zikai, et al.
Published: (2024)
by: Zhou, Zikai, et al.
Published: (2024)
4D Multimodal Co-attention Fusion Network with Latent Contrastive Alignment for Alzheimer's Diagnosis
by: Wei, Yuxiang, et al.
Published: (2025)
by: Wei, Yuxiang, et al.
Published: (2025)
Feature Engineering vs. Deep Learning for Automated Coin Grading: A Comparative Study on Saint-Gaudens Double Eagles
by: Dogra, Tanmay, et al.
Published: (2025)
by: Dogra, Tanmay, et al.
Published: (2025)
Cross-Class Feature Augmentation for Class Incremental Learning
by: Kim, Taehoon, et al.
Published: (2023)
by: Kim, Taehoon, et al.
Published: (2023)
Personalized Federated Learning via Dual-Prompt Optimization and Cross Fusion
by: Zhang, Yuguang, et al.
Published: (2025)
by: Zhang, Yuguang, et al.
Published: (2025)
Feature Engineering is Not Dead: Reviving Classical Machine Learning with Entropy, HOG, and LBP Feature Fusion for Image Classification
by: Sen, Abhijit, et al.
Published: (2025)
by: Sen, Abhijit, et al.
Published: (2025)
Feature-based Graph Attention Networks Improve Online Continual Learning
by: Sim, Adjovi, et al.
Published: (2025)
by: Sim, Adjovi, et al.
Published: (2025)
GC-GAT: Multimodal Vehicular Trajectory Prediction using Graph Goal Conditioning and Cross-context Attention
by: Gulzar, Mahir, et al.
Published: (2025)
by: Gulzar, Mahir, et al.
Published: (2025)
On the Value of Cross-Modal Misalignment in Multimodal Representation Learning
by: Cai, Yichao, et al.
Published: (2025)
by: Cai, Yichao, et al.
Published: (2025)
Partial Multi-View Clustering via Meta-Learning and Contrastive Feature Alignment
by: Chen, BoHao
Published: (2024)
by: Chen, BoHao
Published: (2024)
A Multi-Scale Feature Extraction and Fusion Deep Learning Method for Classification of Wheat Diseases
by: Saleem, Sajjad, et al.
Published: (2025)
by: Saleem, Sajjad, et al.
Published: (2025)
Audio Deepfake Detection with Half-Truth Localisation Using Cross-Attentive Feature Fusion
by: Sutharya, S., et al.
Published: (2026)
by: Sutharya, S., et al.
Published: (2026)
Machine Unlearning on Pre-trained Models by Residual Feature Alignment Using LoRA
by: Qin, Laiqiao, et al.
Published: (2024)
by: Qin, Laiqiao, et al.
Published: (2024)
CrossLLM-Mamba: Multimodal State Space Fusion of LLMs for RNA Interaction Prediction
by: Sadia, Rabeya Tus, et al.
Published: (2026)
by: Sadia, Rabeya Tus, et al.
Published: (2026)
ViTaPEs: Visuotactile Position Encodings for Cross-Modal Alignment in Multimodal Transformers
by: Lygerakis, Fotios, et al.
Published: (2025)
by: Lygerakis, Fotios, et al.
Published: (2025)
Learning Temporal Saliency for Time Series Forecasting with Cross-Scale Attention
by: Delibasoglu, Ibrahim, et al.
Published: (2025)
by: Delibasoglu, Ibrahim, et al.
Published: (2025)
Beyond Simple Fusion: Adaptive Gated Fusion for Robust Multimodal Sentiment Analysis
by: Wu, Han, et al.
Published: (2025)
by: Wu, Han, et al.
Published: (2025)
Similar Items
-
Deep Learning-Based Approach for Identification of Potato Leaf Diseases Using Wrapper Feature Selection and Feature Concatenation
by: Naeem, Muhammad Ahtsam, et al.
Published: (2025) -
DashFusion: Dual-stream Alignment with Hierarchical Bottleneck Fusion for Multimodal Sentiment Analysis
by: Wen, Yuhua, et al.
Published: (2025) -
Multimodal Fusion Strategies for Mapping Biophysical Landscape Features
by: Gordon, Lucia, et al.
Published: (2024) -
Feature, Alignment, and Supervision in Category Learning: A Comparative Approach with Children and Neural Networks
by: Qiu, Fanxiao Wani, et al.
Published: (2026) -
CrossFuse: Learning Infrared and Visible Image Fusion by Cross-Sensor Top-K Vision Alignment and Beyond
by: Shi, Yukai, et al.
Published: (2025)