Open-set Cross Modal Generalization via Multimodal Unified Representation
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Huang, Hai, Xia, Yan, Wang, Shulei, Wang, Hanting, Fang, Minghui, Ji, Shengpeng, Zhou, Sashuai, Jin, Tao, Zhao, Zhou |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Enhancing Multimodal Unified Representations for Cross Modal Generalization
von: Huang, Hai, et al.
Veröffentlicht: (2024)
von: Huang, Hai, et al.
Veröffentlicht: (2024)
Bridging Domain Generalization to Multimodal Domain Generalization via Unified Representations
von: Huang, Hai, et al.
Veröffentlicht: (2025)
von: Huang, Hai, et al.
Veröffentlicht: (2025)
IRBridge: Solving Image Restoration Bridge with Pre-trained Generative Diffusion Models
von: Wang, Hanting, et al.
Veröffentlicht: (2025)
von: Wang, Hanting, et al.
Veröffentlicht: (2025)
Continual Cross-Modal Generalization
von: Xia, Yan, et al.
Veröffentlicht: (2025)
von: Xia, Yan, et al.
Veröffentlicht: (2025)
Semantic Residual for Multimodal Unified Discrete Representation
von: Huang, Hai, et al.
Veröffentlicht: (2024)
von: Huang, Hai, et al.
Veröffentlicht: (2024)
Language-Codec: Bridging Discrete Codec Representations and Speech Language Models
von: Ji, Shengpeng, et al.
Veröffentlicht: (2024)
von: Ji, Shengpeng, et al.
Veröffentlicht: (2024)
TAP: Parameter-efficient Task-Aware Prompting for Adverse Weather Removal
von: Wang, Hanting, et al.
Veröffentlicht: (2025)
von: Wang, Hanting, et al.
Veröffentlicht: (2025)
Enhancing Multi-modal Models with Heterogeneous MoE Adapters for Fine-tuning
von: Zhou, Sashuai, et al.
Veröffentlicht: (2025)
von: Zhou, Sashuai, et al.
Veröffentlicht: (2025)
CART: A Generative Cross-Modal Retrieval Framework with Coarse-To-Fine Semantic Modeling
von: Fang, Minghui, et al.
Veröffentlicht: (2024)
von: Fang, Minghui, et al.
Veröffentlicht: (2024)
OmniSep: Unified Omni-Modality Sound Separation with Query-Mixup
von: Cheng, Xize, et al.
Veröffentlicht: (2024)
von: Cheng, Xize, et al.
Veröffentlicht: (2024)
Speech Watermarking with Discrete Intermediate Representations
von: Ji, Shengpeng, et al.
Veröffentlicht: (2024)
von: Ji, Shengpeng, et al.
Veröffentlicht: (2024)
T2A-Feedback: Improving Basic Capabilities of Text-to-Audio Generation via Fine-grained AI Feedback
von: Wang, Zehan, et al.
Veröffentlicht: (2025)
von: Wang, Zehan, et al.
Veröffentlicht: (2025)
MobileSpeech: A Fast and High-Fidelity Framework for Mobile Zero-Shot Text-to-Speech
von: Ji, Shengpeng, et al.
Veröffentlicht: (2024)
von: Ji, Shengpeng, et al.
Veröffentlicht: (2024)
Entropy-based Coarse and Compressed Semantic Speech Representation Learning
von: Zuo, Jialong, et al.
Veröffentlicht: (2025)
von: Zuo, Jialong, et al.
Veröffentlicht: (2025)
Towards Transformer-Based Aligned Generation with Self-Coherence Guidance
von: Wang, Shulei, et al.
Veröffentlicht: (2025)
von: Wang, Shulei, et al.
Veröffentlicht: (2025)
A Survey on Music Generation from Single-Modal, Cross-Modal, and Multi-Modal Perspectives
von: Li, Shuyu, et al.
Veröffentlicht: (2025)
von: Li, Shuyu, et al.
Veröffentlicht: (2025)
Dynamic Cross-Modal Prompt Generation for Multimodal Continual Instruction Tuning
von: Hu, Tao, et al.
Veröffentlicht: (2026)
von: Hu, Tao, et al.
Veröffentlicht: (2026)
RecBase: Generative Foundation Model Pretraining for Zero-Shot Recommendation
von: Zhou, Sashuai, et al.
Veröffentlicht: (2025)
von: Zhou, Sashuai, et al.
Veröffentlicht: (2025)
ControlSpeech: Towards Simultaneous and Independent Zero-shot Speaker Cloning and Zero-shot Language Style Control
von: Ji, Shengpeng, et al.
Veröffentlicht: (2024)
von: Ji, Shengpeng, et al.
Veröffentlicht: (2024)
Multimodal Prompt Learning with Missing Modalities for Sentiment Analysis and Emotion Recognition
von: Guo, Zirun, et al.
Veröffentlicht: (2024)
von: Guo, Zirun, et al.
Veröffentlicht: (2024)
Efficient Prompting for Continual Adaptation to Missing Modalities
von: Guo, Zirun, et al.
Veröffentlicht: (2025)
von: Guo, Zirun, et al.
Veröffentlicht: (2025)
XM-ALIGN: Unified Cross-Modal Embedding Alignment for Face-Voice Association
von: Fang, Zhihua, et al.
Veröffentlicht: (2025)
von: Fang, Zhihua, et al.
Veröffentlicht: (2025)
SpecBridge: Bridging Mass Spectrometry and Molecular Representations via Cross-Modal Alignment
von: Wang, Yinkai, et al.
Veröffentlicht: (2026)
von: Wang, Yinkai, et al.
Veröffentlicht: (2026)
OpenPSG: Open-set Panoptic Scene Graph Generation via Large Multimodal Models
von: Zhou, Zijian, et al.
Veröffentlicht: (2024)
von: Zhou, Zijian, et al.
Veröffentlicht: (2024)
CATD: Unified Representation Learning for EEG-to-fMRI Cross-Modal Generation
von: Yao, Weiheng, et al.
Veröffentlicht: (2024)
von: Yao, Weiheng, et al.
Veröffentlicht: (2024)
Unified Thinker: A General Reasoning Modular Core for Image Generation
von: Zhou, Sashuai, et al.
Veröffentlicht: (2026)
von: Zhou, Sashuai, et al.
Veröffentlicht: (2026)
TextrolSpeech: A Text Style Control Speech Corpus With Codec Language Text-to-Speech Models
von: Ji, Shengpeng, et al.
Veröffentlicht: (2023)
von: Ji, Shengpeng, et al.
Veröffentlicht: (2023)
FreeBind: Free Lunch in Unified Multimodal Space via Knowledge Fusion
von: Wang, Zehan, et al.
Veröffentlicht: (2024)
von: Wang, Zehan, et al.
Veröffentlicht: (2024)
UniAVGen: Unified Audio and Video Generation with Asymmetric Cross-Modal Interactions
von: Zhang, Guozhen, et al.
Veröffentlicht: (2025)
von: Zhang, Guozhen, et al.
Veröffentlicht: (2025)
WavRAG: Audio-Integrated Retrieval Augmented Generation for Spoken Dialogue Models
von: Chen, Yifu, et al.
Veröffentlicht: (2025)
von: Chen, Yifu, et al.
Veröffentlicht: (2025)
SCDM: Unified Representation Learning for EEG-to-fNIRS Cross-Modal Generation in MI-BCIs
von: Li, Yisheng, et al.
Veröffentlicht: (2024)
von: Li, Yisheng, et al.
Veröffentlicht: (2024)
SGDL: Smart contract vulnerability generation via deep learning
von: Hanting Chu, et al.
Veröffentlicht: (2024)
von: Hanting Chu, et al.
Veröffentlicht: (2024)
Mind the Gap: Learning Modality-Agnostic Representations with a Cross-Modality UNet
von: Niu, Xin, et al.
Veröffentlicht: (2026)
von: Niu, Xin, et al.
Veröffentlicht: (2026)
Enhancing Multimodal Emotion Recognition through Multi-Granularity Cross-Modal Alignment
von: Wang, Xuechen, et al.
Veröffentlicht: (2024)
von: Wang, Xuechen, et al.
Veröffentlicht: (2024)
Rhythm Controllable and Efficient Zero-Shot Voice Conversion via Shortcut Flow Matching
von: Zuo, Jialong, et al.
Veröffentlicht: (2025)
von: Zuo, Jialong, et al.
Veröffentlicht: (2025)
DocRetriever: A Plug-and-Play Framework for Multimodal Document Retrieval with Comprehensive Benchmark
von: Hu, Ruofan, et al.
Veröffentlicht: (2026)
von: Hu, Ruofan, et al.
Veröffentlicht: (2026)
PaperX: A Unified Framework for Multimodal Academic Presentation Generation with Scholar DAG
von: Yu, Tao, et al.
Veröffentlicht: (2026)
von: Yu, Tao, et al.
Veröffentlicht: (2026)
Harmonizing Visual Representations for Unified Multimodal Understanding and Generation
von: Wu, Size, et al.
Veröffentlicht: (2025)
von: Wu, Size, et al.
Veröffentlicht: (2025)
Cross-Modal Attention Network with Dual Graph Learning in Multimodal Recommendation
von: Dai, Ji, et al.
Veröffentlicht: (2026)
von: Dai, Ji, et al.
Veröffentlicht: (2026)
UnityVideo: Unified Multi-Modal Multi-Task Learning for Enhancing World-Aware Video Generation
von: Huang, Jiehui, et al.
Veröffentlicht: (2025)
von: Huang, Jiehui, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Enhancing Multimodal Unified Representations for Cross Modal Generalization
von: Huang, Hai, et al.
Veröffentlicht: (2024) -
Bridging Domain Generalization to Multimodal Domain Generalization via Unified Representations
von: Huang, Hai, et al.
Veröffentlicht: (2025) -
IRBridge: Solving Image Restoration Bridge with Pre-trained Generative Diffusion Models
von: Wang, Hanting, et al.
Veröffentlicht: (2025) -
Continual Cross-Modal Generalization
von: Xia, Yan, et al.
Veröffentlicht: (2025) -
Semantic Residual for Multimodal Unified Discrete Representation
von: Huang, Hai, et al.
Veröffentlicht: (2024)