Bridging Modalities and Transferring Knowledge: Enhanced Multimodal Understanding and Recognition
Fuente:
arXiv
Gespeichert in:
| 1. Verfasser: | Radevski, Gorjan |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Eff-GRot: Efficient and Generalizable Rotation Estimation with Transformers
von: Mathioulakis, Fanis, et al.
Veröffentlicht: (2025)
von: Mathioulakis, Fanis, et al.
Veröffentlicht: (2025)
DAVE: Diagnostic benchmark for Audio Visual Evaluation
von: Radevski, Gorjan, et al.
Veröffentlicht: (2025)
von: Radevski, Gorjan, et al.
Veröffentlicht: (2025)
Classifying Novel 3D-Printed Objects without Retraining: Towards Post-Production Automation in Additive Manufacturing
von: Mathioulakis, Fanis, et al.
Veröffentlicht: (2026)
von: Mathioulakis, Fanis, et al.
Veröffentlicht: (2026)
Multimodal Knowledge Distillation for Egocentric Action Recognition Robust to Missing Modalities
von: Santos-Villafranca, Maria, et al.
Veröffentlicht: (2025)
von: Santos-Villafranca, Maria, et al.
Veröffentlicht: (2025)
Learning Modality Knowledge Alignment for Cross-Modality Transfer
von: Ma, Wenxuan, et al.
Veröffentlicht: (2024)
von: Ma, Wenxuan, et al.
Veröffentlicht: (2024)
Enhancing Zero-Shot Facial Expression Recognition by LLM Knowledge Transfer
von: Zhao, Zengqun, et al.
Veröffentlicht: (2024)
von: Zhao, Zengqun, et al.
Veröffentlicht: (2024)
SkeFi: Cross-Modal Knowledge Transfer for Wireless Skeleton-Based Action Recognition
von: Huang, Shunyu, et al.
Veröffentlicht: (2026)
von: Huang, Shunyu, et al.
Veröffentlicht: (2026)
Cross-Modality Gait Recognition: Bridging LiDAR and Camera Modalities for Human Identification
von: Wang, Rui, et al.
Veröffentlicht: (2024)
von: Wang, Rui, et al.
Veröffentlicht: (2024)
Reading, Not Thinking: Understanding and Bridging the Modality Gap When Text Becomes Pixels in Multimodal LLMs
von: Sun, Kaiser, et al.
Veröffentlicht: (2026)
von: Sun, Kaiser, et al.
Veröffentlicht: (2026)
Asymmetric Cross-Modal Knowledge Distillation: Bridging Modalities with Weak Semantic Consistency
von: Wei, Riling, et al.
Veröffentlicht: (2025)
von: Wei, Riling, et al.
Veröffentlicht: (2025)
Percept, Chat, and then Adapt: Multimodal Knowledge Transfer of Foundation Models for Open-World Video Recognition
von: Chen, Boyu, et al.
Veröffentlicht: (2024)
von: Chen, Boyu, et al.
Veröffentlicht: (2024)
Bridging the Skeleton-Text Modality Gap: Diffusion-Powered Modality Alignment for Zero-shot Skeleton-based Action Recognition
von: Do, Jeonghyeok, et al.
Veröffentlicht: (2024)
von: Do, Jeonghyeok, et al.
Veröffentlicht: (2024)
Multimodal Emotion Recognition via Causal-Diffusion Bridge (Affect-Diff)
von: Sanjyal, Ankit
Veröffentlicht: (2026)
von: Sanjyal, Ankit
Veröffentlicht: (2026)
Robust Brain Tumor Segmentation with Incomplete MRI Modalities Using Hölder Divergence and Mutual Information-Enhanced Knowledge Transfer
von: Cheng, Runze, et al.
Veröffentlicht: (2025)
von: Cheng, Runze, et al.
Veröffentlicht: (2025)
UniFork: Exploring Modality Alignment for Unified Multimodal Understanding and Generation
von: Li, Teng, et al.
Veröffentlicht: (2025)
von: Li, Teng, et al.
Veröffentlicht: (2025)
LLMs as Bridges: Reformulating Grounded Multimodal Named Entity Recognition
von: Li, Jinyuan, et al.
Veröffentlicht: (2024)
von: Li, Jinyuan, et al.
Veröffentlicht: (2024)
Towards Robust and Realible Multimodal Misinformation Recognition with Incomplete Modality
von: Zhou, Hengyang, et al.
Veröffentlicht: (2025)
von: Zhou, Hengyang, et al.
Veröffentlicht: (2025)
Correlation-Decoupled Knowledge Distillation for Multimodal Sentiment Analysis with Incomplete Modalities
von: Li, Mingcheng, et al.
Veröffentlicht: (2024)
von: Li, Mingcheng, et al.
Veröffentlicht: (2024)
Multi Teacher Privileged Knowledge Distillation for Multimodal Expression Recognition
von: Aslam, Muhammad Haseeb, et al.
Veröffentlicht: (2024)
von: Aslam, Muhammad Haseeb, et al.
Veröffentlicht: (2024)
Understanding Multimodal Procedural Knowledge by Sequencing Multimodal Instructional Manuals
von: Wu, Te-Lin, et al.
Veröffentlicht: (2021)
von: Wu, Te-Lin, et al.
Veröffentlicht: (2021)
Bridging the Intent Gap: Knowledge-Enhanced Visual Generation
von: Cheng, Yi, et al.
Veröffentlicht: (2024)
von: Cheng, Yi, et al.
Veröffentlicht: (2024)
Multimodal Prompt Learning with Missing Modalities for Sentiment Analysis and Emotion Recognition
von: Guo, Zirun, et al.
Veröffentlicht: (2024)
von: Guo, Zirun, et al.
Veröffentlicht: (2024)
ECMF: Enhanced Cross-Modal Fusion for Multimodal Emotion Recognition in MER-SEMI Challenge
von: Hu, Juewen, et al.
Veröffentlicht: (2025)
von: Hu, Juewen, et al.
Veröffentlicht: (2025)
Semantic-Enhanced Cross-Modal Place Recognition for Robust Robot Localization
von: Lin, Yujia, et al.
Veröffentlicht: (2025)
von: Lin, Yujia, et al.
Veröffentlicht: (2025)
HBridge: H-Shape Bridging of Heterogeneous Experts for Unified Multimodal Understanding and Generation
von: Wang, Xiang, et al.
Veröffentlicht: (2025)
von: Wang, Xiang, et al.
Veröffentlicht: (2025)
Category-Adaptive Cross-Modal Semantic Refinement and Transfer for Open-Vocabulary Multi-Label Recognition
von: Liu, Haijing, et al.
Veröffentlicht: (2024)
von: Liu, Haijing, et al.
Veröffentlicht: (2024)
Enhancing Modal Fusion by Alignment and Label Matching for Multimodal Emotion Recognition
von: Li, Qifei, et al.
Veröffentlicht: (2024)
von: Li, Qifei, et al.
Veröffentlicht: (2024)
Bridging the Gap in Missing Modalities: Leveraging Knowledge Distillation and Style Matching for Brain Tumor Segmentation
von: Zhu, Shenghao, et al.
Veröffentlicht: (2025)
von: Zhu, Shenghao, et al.
Veröffentlicht: (2025)
VKnowU: Evaluating Visual Knowledge Understanding in Multimodal LLMs
von: Jiang, Tianxiang, et al.
Veröffentlicht: (2025)
von: Jiang, Tianxiang, et al.
Veröffentlicht: (2025)
Contrasting Intra-Modal and Ranking Cross-Modal Hard Negatives to Enhance Visio-Linguistic Compositional Understanding
von: Zhang, Le, et al.
Veröffentlicht: (2023)
von: Zhang, Le, et al.
Veröffentlicht: (2023)
IMUGPT 2.0: Language-Based Cross Modality Transfer for Sensor-Based Human Activity Recognition
von: Leng, Zikang, et al.
Veröffentlicht: (2024)
von: Leng, Zikang, et al.
Veröffentlicht: (2024)
Beyond CLIP: Knowledge-Enhanced Multimodal Transformers for Cross-Modal Alignment in Diabetic Retinopathy Diagnosis
von: Samanta, Argha Kamal, et al.
Veröffentlicht: (2025)
von: Samanta, Argha Kamal, et al.
Veröffentlicht: (2025)
Multimodal Emotion Recognition with Vision-language Prompting and Modality Dropout
von: QI, Anbin, et al.
Veröffentlicht: (2024)
von: QI, Anbin, et al.
Veröffentlicht: (2024)
NocPlace: Nocturnal Visual Place Recognition via Generative and Inherited Knowledge Transfer
von: Liu, Bingxi, et al.
Veröffentlicht: (2024)
von: Liu, Bingxi, et al.
Veröffentlicht: (2024)
Knowledge-Enhanced Facial Expression Recognition with Emotional-to-Neutral Transformation
von: Li, Hangyu, et al.
Veröffentlicht: (2024)
von: Li, Hangyu, et al.
Veröffentlicht: (2024)
Towards Unified Multimodal Editing with Enhanced Knowledge Collaboration
von: Pan, Kaihang, et al.
Veröffentlicht: (2024)
von: Pan, Kaihang, et al.
Veröffentlicht: (2024)
Enhancing Knowledge Transfer in Hyperspectral Image Classification via Cross-scene Knowledge Integration
von: Huo, Lu, et al.
Veröffentlicht: (2025)
von: Huo, Lu, et al.
Veröffentlicht: (2025)
Beyond Emotion Recognition: A Multi-Turn Multimodal Emotion Understanding and Reasoning Benchmark
von: Hu, Jinpeng, et al.
Veröffentlicht: (2025)
von: Hu, Jinpeng, et al.
Veröffentlicht: (2025)
Non-target Divergence Hypothesis: Toward Understanding Domain Gaps in Cross-Modal Knowledge Distillation
von: Chen, Yilong, et al.
Veröffentlicht: (2024)
von: Chen, Yilong, et al.
Veröffentlicht: (2024)
DuwatBench: Bridging Language and Visual Heritage through an Arabic Calligraphy Benchmark for Multimodal Understanding
von: Patle, Shubham, et al.
Veröffentlicht: (2026)
von: Patle, Shubham, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
Eff-GRot: Efficient and Generalizable Rotation Estimation with Transformers
von: Mathioulakis, Fanis, et al.
Veröffentlicht: (2025) -
DAVE: Diagnostic benchmark for Audio Visual Evaluation
von: Radevski, Gorjan, et al.
Veröffentlicht: (2025) -
Classifying Novel 3D-Printed Objects without Retraining: Towards Post-Production Automation in Additive Manufacturing
von: Mathioulakis, Fanis, et al.
Veröffentlicht: (2026) -
Multimodal Knowledge Distillation for Egocentric Action Recognition Robust to Missing Modalities
von: Santos-Villafranca, Maria, et al.
Veröffentlicht: (2025) -
Learning Modality Knowledge Alignment for Cross-Modality Transfer
von: Ma, Wenxuan, et al.
Veröffentlicht: (2024)