Gespeichert in:
| Hauptverfasser: | Wang, Rui, Yang, Shichun, Chen, Yuyi, Li, Zhuoyang, Tong, Zexiang, Xu, Jianyi, Lu, Jiayi, Feng, Xinjie, Cao, Yaoguang |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | https://arxiv.org/abs/2505.11066 |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Synesthesia of Vehicles: Tactile Data Synthesis from Visual Inputs
von: Wang, Rui, et al.
Veröffentlicht: (2026)
von: Wang, Rui, et al.
Veröffentlicht: (2026)
Emotional Cues Extraction and Fusion for Multi-modal Emotion Prediction and Recognition in Conversation
von: Shi, Haoxiang, et al.
Veröffentlicht: (2024)
von: Shi, Haoxiang, et al.
Veröffentlicht: (2024)
MMoFusion: Multi-modal Co-Speech Motion Generation with Diffusion Model
von: Wang, Sen, et al.
Veröffentlicht: (2024)
von: Wang, Sen, et al.
Veröffentlicht: (2024)
Terrain Diffusion Network: Climatic-Aware Terrain Generation with Geological Sketch Guidance
von: Hu, Zexin, et al.
Veröffentlicht: (2023)
von: Hu, Zexin, et al.
Veröffentlicht: (2023)
Clinical Multi-modal Fusion with Heterogeneous Graph and Disease Correlation Learning for Multi-Disease Prediction
von: Jiang, Yueheng, et al.
Veröffentlicht: (2025)
von: Jiang, Yueheng, et al.
Veröffentlicht: (2025)
MDF: A Dynamic Fusion Model for Multi-modal Fake News Detection
von: Lv, Hongzhen, et al.
Veröffentlicht: (2024)
von: Lv, Hongzhen, et al.
Veröffentlicht: (2024)
Tile Classification Based Viewport Prediction with Multi-modal Fusion Transformer
von: Zhang, Zhihao, et al.
Veröffentlicht: (2023)
von: Zhang, Zhihao, et al.
Veröffentlicht: (2023)
Multi-modal Segment Assemblage Network for Ad Video Editing with Importance-Coherence Reward
von: Tang, Yolo Yunlong, et al.
Veröffentlicht: (2022)
von: Tang, Yolo Yunlong, et al.
Veröffentlicht: (2022)
MCIHN: A Hybrid Network Model Based on Multi-path Cross-modal Interaction for Multimodal Emotion Recognition
von: Zhang, Haoyang, et al.
Veröffentlicht: (2025)
von: Zhang, Haoyang, et al.
Veröffentlicht: (2025)
Multi-modal and Metadata Capture Model for Micro Video Popularity Prediction
von: Lu, Jiacheng, et al.
Veröffentlicht: (2025)
von: Lu, Jiacheng, et al.
Veröffentlicht: (2025)
Deep Mamba Multi-modal Learning
von: Zhu, Jian, et al.
Veröffentlicht: (2024)
von: Zhu, Jian, et al.
Veröffentlicht: (2024)
An Emotion Recognition Framework via Cross-modal Alignment of EEG and Eye Movement Data
von: Wang, Jianlu, et al.
Veröffentlicht: (2025)
von: Wang, Jianlu, et al.
Veröffentlicht: (2025)
Structure-Aware Residual-Center Representation for Self-Supervised Open-Set 3D Cross-Modal Retrieval
von: Xu, Yang, et al.
Veröffentlicht: (2024)
von: Xu, Yang, et al.
Veröffentlicht: (2024)
Hierarchical Sub-action Tree for Continuous Sign Language Recognition
von: Yang, Dejie, et al.
Veröffentlicht: (2025)
von: Yang, Dejie, et al.
Veröffentlicht: (2025)
High-level Codes and Fine-grained Weights for Online Multi-modal Hashing Retrieval
von: Zhan, Yu-Wei, et al.
Veröffentlicht: (2024)
von: Zhan, Yu-Wei, et al.
Veröffentlicht: (2024)
EidetiCom: A Cross-modal Brain-Computer Semantic Communication Paradigm for Decoding Visual Perception
von: Zheng, Linfeng, et al.
Veröffentlicht: (2024)
von: Zheng, Linfeng, et al.
Veröffentlicht: (2024)
Interactive Spatial-Frequency Fusion Mamba for Multi-Modal Image Fusion
von: Zhu, Yixin, et al.
Veröffentlicht: (2026)
von: Zhu, Yixin, et al.
Veröffentlicht: (2026)
MMAPS: End-to-End Multi-Grained Multi-Modal Attribute-Aware Product Summarization
von: Chen, Tao, et al.
Veröffentlicht: (2023)
von: Chen, Tao, et al.
Veröffentlicht: (2023)
Continual Panoptic Perception: Towards Multi-modal Incremental Interpretation of Remote Sensing Images
von: Yuan, Bo, et al.
Veröffentlicht: (2024)
von: Yuan, Bo, et al.
Veröffentlicht: (2024)
A Survey of Multi-sensor Fusion Perception for Embodied AI: Background, Methods, Challenges and Prospects
von: Ruan, Shulan, et al.
Veröffentlicht: (2025)
von: Ruan, Shulan, et al.
Veröffentlicht: (2025)
LoginMEA: Local-to-Global Interaction Network for Multi-modal Entity Alignment
von: Su, Taoyu, et al.
Veröffentlicht: (2024)
von: Su, Taoyu, et al.
Veröffentlicht: (2024)
ViFusion: In-Network Tensor Fusion for Scalable Video Feature Indexing
von: Wang, Yisu, et al.
Veröffentlicht: (2025)
von: Wang, Yisu, et al.
Veröffentlicht: (2025)
AIM: Let Any Multi-modal Large Language Models Embrace Efficient In-Context Learning
von: Gao, Jun, et al.
Veröffentlicht: (2024)
von: Gao, Jun, et al.
Veröffentlicht: (2024)
LLM-based Fusion of Multi-modal Features for Commercial Memorability Prediction
von: Pramov, Aleksandar
Veröffentlicht: (2025)
von: Pramov, Aleksandar
Veröffentlicht: (2025)
CDI-DTI: A Strong Cross-domain Interpretable Drug-Target Interaction Prediction Framework Based on Multi-Strategy Fusion
von: Li, Xiangyu, et al.
Veröffentlicht: (2025)
von: Li, Xiangyu, et al.
Veröffentlicht: (2025)
Towards Structure-aware Model for Multi-modal Knowledge Graph Completion
von: Li, Linyu, et al.
Veröffentlicht: (2025)
von: Li, Linyu, et al.
Veröffentlicht: (2025)
Generative AI-enabled Mobile Tactical Multimedia Networks: Distribution, Generation, and Perception
von: Xu, Minrui, et al.
Veröffentlicht: (2024)
von: Xu, Minrui, et al.
Veröffentlicht: (2024)
IBMEA: Exploring Variational Information Bottleneck for Multi-modal Entity Alignment
von: Su, Taoyu, et al.
Veröffentlicht: (2024)
von: Su, Taoyu, et al.
Veröffentlicht: (2024)
Physics-Aware Novel-View Acoustic Synthesis with Vision-Language Priors and 3D Acoustic Environment Modeling
von: Fan, Congyi, et al.
Veröffentlicht: (2026)
von: Fan, Congyi, et al.
Veröffentlicht: (2026)
Voices, Faces, and Feelings: Multi-modal Emotion-Cognition Captioning for Mental Health Understanding
von: Zhou, Zhiyuan, et al.
Veröffentlicht: (2026)
von: Zhou, Zhiyuan, et al.
Veröffentlicht: (2026)
Routing Experts: Learning to Route Dynamic Experts in Multi-modal Large Language Models
von: Wu, Qiong, et al.
Veröffentlicht: (2024)
von: Wu, Qiong, et al.
Veröffentlicht: (2024)
Characterizing Multimedia Information Environment through Multi-modal Clustering of YouTube Videos
von: Yousefi, Niloofar, et al.
Veröffentlicht: (2024)
von: Yousefi, Niloofar, et al.
Veröffentlicht: (2024)
Multi-source Knowledge Enhanced Graph Attention Networks for Multimodal Fact Verification
von: Cao, Han, et al.
Veröffentlicht: (2024)
von: Cao, Han, et al.
Veröffentlicht: (2024)
Spatiotemporal Graph Guided Multi-modal Network for Livestreaming Product Retrieval
von: Hu, Xiaowan, et al.
Veröffentlicht: (2024)
von: Hu, Xiaowan, et al.
Veröffentlicht: (2024)
FinCall-Surprise: A Large Scale Multi-modal Benchmark for Earning Surprise Prediction
von: Shu, Dong, et al.
Veröffentlicht: (2025)
von: Shu, Dong, et al.
Veröffentlicht: (2025)
Quantifying and Enhancing Multi-modal Robustness with Modality Preference
von: Yang, Zequn, et al.
Veröffentlicht: (2024)
von: Yang, Zequn, et al.
Veröffentlicht: (2024)
Mixture-of-Prompt-Experts for Multi-modal Semantic Understanding
von: Wu, Zichen, et al.
Veröffentlicht: (2024)
von: Wu, Zichen, et al.
Veröffentlicht: (2024)
Interest-Aware Joint Caching, Computing, and Communication Optimization for Mobile VR Delivery in MEC Networks
von: Fu, Baojie, et al.
Veröffentlicht: (2024)
von: Fu, Baojie, et al.
Veröffentlicht: (2024)
Perception-Aware Video Semantic Communication
von: Huang, Yinhuan, et al.
Veröffentlicht: (2026)
von: Huang, Yinhuan, et al.
Veröffentlicht: (2026)
EEmo-Bench: A Benchmark for Multi-modal Large Language Models on Image Evoked Emotion Assessment
von: Gao, Lancheng, et al.
Veröffentlicht: (2025)
von: Gao, Lancheng, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Synesthesia of Vehicles: Tactile Data Synthesis from Visual Inputs
von: Wang, Rui, et al.
Veröffentlicht: (2026) -
Emotional Cues Extraction and Fusion for Multi-modal Emotion Prediction and Recognition in Conversation
von: Shi, Haoxiang, et al.
Veröffentlicht: (2024) -
MMoFusion: Multi-modal Co-Speech Motion Generation with Diffusion Model
von: Wang, Sen, et al.
Veröffentlicht: (2024) -
Terrain Diffusion Network: Climatic-Aware Terrain Generation with Geological Sketch Guidance
von: Hu, Zexin, et al.
Veröffentlicht: (2023) -
Clinical Multi-modal Fusion with Heterogeneous Graph and Disease Correlation Learning for Multi-Disease Prediction
von: Jiang, Yueheng, et al.
Veröffentlicht: (2025)