Open-set Cross Modal Generalization via Multimodal Unified Representation
Fuente:
arXiv
Saved in:
| Main Authors: | Huang, Hai, Xia, Yan, Wang, Shulei, Wang, Hanting, Fang, Minghui, Ji, Shengpeng, Zhou, Sashuai, Jin, Tao, Zhao, Zhou |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Enhancing Multimodal Unified Representations for Cross Modal Generalization
by: Huang, Hai, et al.
Published: (2024)
by: Huang, Hai, et al.
Published: (2024)
Bridging Domain Generalization to Multimodal Domain Generalization via Unified Representations
by: Huang, Hai, et al.
Published: (2025)
by: Huang, Hai, et al.
Published: (2025)
IRBridge: Solving Image Restoration Bridge with Pre-trained Generative Diffusion Models
by: Wang, Hanting, et al.
Published: (2025)
by: Wang, Hanting, et al.
Published: (2025)
Continual Cross-Modal Generalization
by: Xia, Yan, et al.
Published: (2025)
by: Xia, Yan, et al.
Published: (2025)
Semantic Residual for Multimodal Unified Discrete Representation
by: Huang, Hai, et al.
Published: (2024)
by: Huang, Hai, et al.
Published: (2024)
Language-Codec: Bridging Discrete Codec Representations and Speech Language Models
by: Ji, Shengpeng, et al.
Published: (2024)
by: Ji, Shengpeng, et al.
Published: (2024)
TAP: Parameter-efficient Task-Aware Prompting for Adverse Weather Removal
by: Wang, Hanting, et al.
Published: (2025)
by: Wang, Hanting, et al.
Published: (2025)
Enhancing Multi-modal Models with Heterogeneous MoE Adapters for Fine-tuning
by: Zhou, Sashuai, et al.
Published: (2025)
by: Zhou, Sashuai, et al.
Published: (2025)
CART: A Generative Cross-Modal Retrieval Framework with Coarse-To-Fine Semantic Modeling
by: Fang, Minghui, et al.
Published: (2024)
by: Fang, Minghui, et al.
Published: (2024)
OmniSep: Unified Omni-Modality Sound Separation with Query-Mixup
by: Cheng, Xize, et al.
Published: (2024)
by: Cheng, Xize, et al.
Published: (2024)
Speech Watermarking with Discrete Intermediate Representations
by: Ji, Shengpeng, et al.
Published: (2024)
by: Ji, Shengpeng, et al.
Published: (2024)
T2A-Feedback: Improving Basic Capabilities of Text-to-Audio Generation via Fine-grained AI Feedback
by: Wang, Zehan, et al.
Published: (2025)
by: Wang, Zehan, et al.
Published: (2025)
MobileSpeech: A Fast and High-Fidelity Framework for Mobile Zero-Shot Text-to-Speech
by: Ji, Shengpeng, et al.
Published: (2024)
by: Ji, Shengpeng, et al.
Published: (2024)
Entropy-based Coarse and Compressed Semantic Speech Representation Learning
by: Zuo, Jialong, et al.
Published: (2025)
by: Zuo, Jialong, et al.
Published: (2025)
Towards Transformer-Based Aligned Generation with Self-Coherence Guidance
by: Wang, Shulei, et al.
Published: (2025)
by: Wang, Shulei, et al.
Published: (2025)
A Survey on Music Generation from Single-Modal, Cross-Modal, and Multi-Modal Perspectives
by: Li, Shuyu, et al.
Published: (2025)
by: Li, Shuyu, et al.
Published: (2025)
Dynamic Cross-Modal Prompt Generation for Multimodal Continual Instruction Tuning
by: Hu, Tao, et al.
Published: (2026)
by: Hu, Tao, et al.
Published: (2026)
RecBase: Generative Foundation Model Pretraining for Zero-Shot Recommendation
by: Zhou, Sashuai, et al.
Published: (2025)
by: Zhou, Sashuai, et al.
Published: (2025)
ControlSpeech: Towards Simultaneous and Independent Zero-shot Speaker Cloning and Zero-shot Language Style Control
by: Ji, Shengpeng, et al.
Published: (2024)
by: Ji, Shengpeng, et al.
Published: (2024)
Multimodal Prompt Learning with Missing Modalities for Sentiment Analysis and Emotion Recognition
by: Guo, Zirun, et al.
Published: (2024)
by: Guo, Zirun, et al.
Published: (2024)
Efficient Prompting for Continual Adaptation to Missing Modalities
by: Guo, Zirun, et al.
Published: (2025)
by: Guo, Zirun, et al.
Published: (2025)
XM-ALIGN: Unified Cross-Modal Embedding Alignment for Face-Voice Association
by: Fang, Zhihua, et al.
Published: (2025)
by: Fang, Zhihua, et al.
Published: (2025)
SpecBridge: Bridging Mass Spectrometry and Molecular Representations via Cross-Modal Alignment
by: Wang, Yinkai, et al.
Published: (2026)
by: Wang, Yinkai, et al.
Published: (2026)
OpenPSG: Open-set Panoptic Scene Graph Generation via Large Multimodal Models
by: Zhou, Zijian, et al.
Published: (2024)
by: Zhou, Zijian, et al.
Published: (2024)
CATD: Unified Representation Learning for EEG-to-fMRI Cross-Modal Generation
by: Yao, Weiheng, et al.
Published: (2024)
by: Yao, Weiheng, et al.
Published: (2024)
Unified Thinker: A General Reasoning Modular Core for Image Generation
by: Zhou, Sashuai, et al.
Published: (2026)
by: Zhou, Sashuai, et al.
Published: (2026)
TextrolSpeech: A Text Style Control Speech Corpus With Codec Language Text-to-Speech Models
by: Ji, Shengpeng, et al.
Published: (2023)
by: Ji, Shengpeng, et al.
Published: (2023)
FreeBind: Free Lunch in Unified Multimodal Space via Knowledge Fusion
by: Wang, Zehan, et al.
Published: (2024)
by: Wang, Zehan, et al.
Published: (2024)
UniAVGen: Unified Audio and Video Generation with Asymmetric Cross-Modal Interactions
by: Zhang, Guozhen, et al.
Published: (2025)
by: Zhang, Guozhen, et al.
Published: (2025)
WavRAG: Audio-Integrated Retrieval Augmented Generation for Spoken Dialogue Models
by: Chen, Yifu, et al.
Published: (2025)
by: Chen, Yifu, et al.
Published: (2025)
SCDM: Unified Representation Learning for EEG-to-fNIRS Cross-Modal Generation in MI-BCIs
by: Li, Yisheng, et al.
Published: (2024)
by: Li, Yisheng, et al.
Published: (2024)
SGDL: Smart contract vulnerability generation via deep learning
by: Hanting Chu, et al.
Published: (2024)
by: Hanting Chu, et al.
Published: (2024)
Mind the Gap: Learning Modality-Agnostic Representations with a Cross-Modality UNet
by: Niu, Xin, et al.
Published: (2026)
by: Niu, Xin, et al.
Published: (2026)
Enhancing Multimodal Emotion Recognition through Multi-Granularity Cross-Modal Alignment
by: Wang, Xuechen, et al.
Published: (2024)
by: Wang, Xuechen, et al.
Published: (2024)
Rhythm Controllable and Efficient Zero-Shot Voice Conversion via Shortcut Flow Matching
by: Zuo, Jialong, et al.
Published: (2025)
by: Zuo, Jialong, et al.
Published: (2025)
DocRetriever: A Plug-and-Play Framework for Multimodal Document Retrieval with Comprehensive Benchmark
by: Hu, Ruofan, et al.
Published: (2026)
by: Hu, Ruofan, et al.
Published: (2026)
PaperX: A Unified Framework for Multimodal Academic Presentation Generation with Scholar DAG
by: Yu, Tao, et al.
Published: (2026)
by: Yu, Tao, et al.
Published: (2026)
Harmonizing Visual Representations for Unified Multimodal Understanding and Generation
by: Wu, Size, et al.
Published: (2025)
by: Wu, Size, et al.
Published: (2025)
Cross-Modal Attention Network with Dual Graph Learning in Multimodal Recommendation
by: Dai, Ji, et al.
Published: (2026)
by: Dai, Ji, et al.
Published: (2026)
UnityVideo: Unified Multi-Modal Multi-Task Learning for Enhancing World-Aware Video Generation
by: Huang, Jiehui, et al.
Published: (2025)
by: Huang, Jiehui, et al.
Published: (2025)
Similar Items
-
Enhancing Multimodal Unified Representations for Cross Modal Generalization
by: Huang, Hai, et al.
Published: (2024) -
Bridging Domain Generalization to Multimodal Domain Generalization via Unified Representations
by: Huang, Hai, et al.
Published: (2025) -
IRBridge: Solving Image Restoration Bridge with Pre-trained Generative Diffusion Models
by: Wang, Hanting, et al.
Published: (2025) -
Continual Cross-Modal Generalization
by: Xia, Yan, et al.
Published: (2025) -
Semantic Residual for Multimodal Unified Discrete Representation
by: Huang, Hai, et al.
Published: (2024)