Saved in:
| Main Authors: | Wang, Zihua, Li, Ruibo, Du, Haozhe, Zhou, Joey Tianyi, Zhang, Yu, Yang, Xu |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2505.12728 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Noisy-Correspondence Learning for Text-to-Image Person Re-identification
by: Qin, Yang, et al.
Published: (2023)
by: Qin, Yang, et al.
Published: (2023)
Latent Reconstruction from Generated Data for Multimodal Misinformation Detection
by: Papadopoulos, Stefanos-Iordanis, et al.
Published: (2025)
by: Papadopoulos, Stefanos-Iordanis, et al.
Published: (2025)
Mitigating Audiovisual Mismatch in Visual-Guide Audio Captioning
by: Xu, Le, et al.
Published: (2025)
by: Xu, Le, et al.
Published: (2025)
Video2Roleplay: A Multimodal Dataset and Framework for Video-Guided Role-playing Agents
by: Zhang, Xueqiao, et al.
Published: (2025)
by: Zhang, Xueqiao, et al.
Published: (2025)
SAM as the Guide: Mastering Pseudo-Label Refinement in Semi-Supervised Referring Expression Segmentation
by: Yang, Danni, et al.
Published: (2024)
by: Yang, Danni, et al.
Published: (2024)
StableMoFusion: Towards Robust and Efficient Diffusion-based Motion Generation Framework
by: Huang, Yiheng, et al.
Published: (2024)
by: Huang, Yiheng, et al.
Published: (2024)
StableDub: Taming Diffusion Prior for Generalized and Efficient Visual Dubbing
by: Chen, Liyang, et al.
Published: (2025)
by: Chen, Liyang, et al.
Published: (2025)
Zero-shot Video Moment Retrieval via Off-the-shelf Multimodal Large Language Models
by: Xu, Yifang, et al.
Published: (2025)
by: Xu, Yifang, et al.
Published: (2025)
TextBraTS: Text-Guided Volumetric Brain Tumor Segmentation with Innovative Dataset Development and Fusion Module Exploration
by: Shi, Xiaoyu, et al.
Published: (2025)
by: Shi, Xiaoyu, et al.
Published: (2025)
Bridging Compressed Image Latents and Multimodal Large Language Models
by: Kao, Chia-Hao, et al.
Published: (2024)
by: Kao, Chia-Hao, et al.
Published: (2024)
MCIHN: A Hybrid Network Model Based on Multi-path Cross-modal Interaction for Multimodal Emotion Recognition
by: Zhang, Haoyang, et al.
Published: (2025)
by: Zhang, Haoyang, et al.
Published: (2025)
MagicAnime: A Hierarchically Annotated, Multimodal and Multitasking Dataset with Benchmarks for Cartoon Animation Generation
by: Xu, Shuolin, et al.
Published: (2025)
by: Xu, Shuolin, et al.
Published: (2025)
ReactDiff: Latent Diffusion for Facial Reaction Generation
by: Li, Jiaming, et al.
Published: (2025)
by: Li, Jiaming, et al.
Published: (2025)
Lyra: An Efficient and Speech-Centric Framework for Omni-Cognition
by: Zhong, Zhisheng, et al.
Published: (2024)
by: Zhong, Zhisheng, et al.
Published: (2024)
Towards Robust and Realible Multimodal Misinformation Recognition with Incomplete Modality
by: Zhou, Hengyang, et al.
Published: (2025)
by: Zhou, Hengyang, et al.
Published: (2025)
PanoGen++: Domain-Adapted Text-Guided Panoramic Environment Generation for Vision-and-Language Navigation
by: Wang, Sen, et al.
Published: (2025)
by: Wang, Sen, et al.
Published: (2025)
Rethinking Multi-Condition DiTs: Eliminating Redundant Attention via Position-Alignment and Keyword-Scoping
by: Zhou, Chao, et al.
Published: (2026)
by: Zhou, Chao, et al.
Published: (2026)
One Framework to Rule Them All: Unifying Multimodal Tasks with LLM Neural-Tuning
by: Sun, Hao, et al.
Published: (2024)
by: Sun, Hao, et al.
Published: (2024)
ALIEN: Analytic Latent Watermarking for Controllable Generation
by: Lei, Liangqi, et al.
Published: (2026)
by: Lei, Liangqi, et al.
Published: (2026)
Latent Feature-Guided Conditional Diffusion for Generative Image Semantic Communication
by: Chen, Zehao, et al.
Published: (2025)
by: Chen, Zehao, et al.
Published: (2025)
Video-Guided Foley Sound Generation with Multimodal Controls
by: Chen, Ziyang, et al.
Published: (2024)
by: Chen, Ziyang, et al.
Published: (2024)
Kubrick: Multimodal Agent Collaborations for Synthetic Video Generation
by: He, Liu, et al.
Published: (2024)
by: He, Liu, et al.
Published: (2024)
UniCode$^2$: Cascaded Large-scale Codebooks for Unified Multimodal Understanding and Generation
by: Chen, Yanzhe, et al.
Published: (2025)
by: Chen, Yanzhe, et al.
Published: (2025)
MAGE: Multimodal Alignment and Generation Enhancement via Bridging Visual and Semantic Spaces
by: E, Shaojun, et al.
Published: (2025)
by: E, Shaojun, et al.
Published: (2025)
LipGen: Viseme-Guided Lip Video Generation for Enhancing Visual Speech Recognition
by: Hao, Bowen, et al.
Published: (2025)
by: Hao, Bowen, et al.
Published: (2025)
Proxy-Tuning: Tailoring Multimodal Autoregressive Models for Subject-Driven Image Generation
by: Wu, Yi, et al.
Published: (2025)
by: Wu, Yi, et al.
Published: (2025)
Semi-supervised Chinese Poem-to-Painting Generation via Cycle-consistent Adversarial Networks
by: Lu, Zhengyang, et al.
Published: (2024)
by: Lu, Zhengyang, et al.
Published: (2024)
Multimodal Class-aware Semantic Enhancement Network for Audio-Visual Video Parsing
by: Zhao, Pengcheng, et al.
Published: (2024)
by: Zhao, Pengcheng, et al.
Published: (2024)
Grounded Chain-of-Thought for Multimodal Large Language Models
by: Wu, Qiong, et al.
Published: (2025)
by: Wu, Qiong, et al.
Published: (2025)
A Multimodal Transformer for Live Streaming Highlight Prediction
by: Deng, Jiaxin, et al.
Published: (2024)
by: Deng, Jiaxin, et al.
Published: (2024)
DAE-Talker: High Fidelity Speech-Driven Talking Face Generation with Diffusion Autoencoder
by: Du, Chenpeng, et al.
Published: (2023)
by: Du, Chenpeng, et al.
Published: (2023)
Wills Aligner: Multi-Subject Collaborative Brain Visual Decoding
by: Bao, Guangyin, et al.
Published: (2024)
by: Bao, Guangyin, et al.
Published: (2024)
DanceCamera3D: 3D Camera Movement Synthesis with Music and Dance
by: Wang, Zixuan, et al.
Published: (2024)
by: Wang, Zixuan, et al.
Published: (2024)
XEmoGPT: An Explainable Multimodal Emotion Recognition Framework with Cue-Level Perception and Reasoning
by: Zhang, Hanwen, et al.
Published: (2026)
by: Zhang, Hanwen, et al.
Published: (2026)
Graph-Driven Multimodal Feature Learning Framework for Apparent Personality Assessment
by: Wang, Kangsheng, et al.
Published: (2025)
by: Wang, Kangsheng, et al.
Published: (2025)
A Unit Enhancement and Guidance Framework for Audio-Driven Avatar Video Generation
by: Zhou, S. Z., et al.
Published: (2025)
by: Zhou, S. Z., et al.
Published: (2025)
Detached and Interactive Multimodal Learning
by: Fan, Yunfeng, et al.
Published: (2024)
by: Fan, Yunfeng, et al.
Published: (2024)
Synthetic Perception: Can Generated Images Unlock Latent Visual Prior for Text-Centric Reasoning?
by: Huang, Yuesheng, et al.
Published: (2025)
by: Huang, Yuesheng, et al.
Published: (2025)
Interpretable Concept-based Deep Learning Framework for Multimodal Human Behavior Modeling
by: Li, Xinyu, et al.
Published: (2025)
by: Li, Xinyu, et al.
Published: (2025)
A Novel Approach to Industrial Defect Generation through Blended Latent Diffusion Model with Online Adaptation
by: Li, Hanxi, et al.
Published: (2024)
by: Li, Hanxi, et al.
Published: (2024)
Similar Items
-
Noisy-Correspondence Learning for Text-to-Image Person Re-identification
by: Qin, Yang, et al.
Published: (2023) -
Latent Reconstruction from Generated Data for Multimodal Misinformation Detection
by: Papadopoulos, Stefanos-Iordanis, et al.
Published: (2025) -
Mitigating Audiovisual Mismatch in Visual-Guide Audio Captioning
by: Xu, Le, et al.
Published: (2025) -
Video2Roleplay: A Multimodal Dataset and Framework for Video-Guided Role-playing Agents
by: Zhang, Xueqiao, et al.
Published: (2025) -
SAM as the Guide: Mastering Pseudo-Label Refinement in Semi-Supervised Referring Expression Segmentation
by: Yang, Danni, et al.
Published: (2024)