CyIN: Cyclic Informative Latent Space for Bridging Complete and Incomplete Multimodal Learning
Fuente:
arXiv
Salvato in:
| Autori principali: | Lin, Ronghao, He, Qiaolin, Mai, Sijie, Zeng, Ying, Xiong, Aolin, Huang, Li, Tan, Yap-Peng, Hu, Haifeng |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Multi-source Multimodal Progressive Domain Adaption for Audio-Visual Deception Detection
di: Lin, Ronghao, et al.
Pubblicazione: (2025)
di: Lin, Ronghao, et al.
Pubblicazione: (2025)
MissMAC-Bench: Building Solid Benchmark for Missing Modality Issue in Robust Multimodal Affective Computing
di: Lin, Ronghao, et al.
Pubblicazione: (2026)
di: Lin, Ronghao, et al.
Pubblicazione: (2026)
E3RG: Building Explicit Emotion-driven Empathetic Response Generation System with Multimodal Large Language Model
di: Lin, Ronghao, et al.
Pubblicazione: (2025)
di: Lin, Ronghao, et al.
Pubblicazione: (2025)
End-to-end Semantic-centric Video-based Multimodal Affective Computing
di: Lin, Ronghao, et al.
Pubblicazione: (2024)
di: Lin, Ronghao, et al.
Pubblicazione: (2024)
Audio Super-Resolution with Latent Bridge Models
di: Li, Chang, et al.
Pubblicazione: (2025)
di: Li, Chang, et al.
Pubblicazione: (2025)
Generalized Audio Deepfake Detection Using Frame-level Latent Information Entropy
di: Zhao, Botao, et al.
Pubblicazione: (2025)
di: Zhao, Botao, et al.
Pubblicazione: (2025)
VS-Singer: Vision-Guided Stereo Singing Voice Synthesis with Consistency Schrödinger Bridge
di: Zhao, Zijing, et al.
Pubblicazione: (2025)
di: Zhao, Zijing, et al.
Pubblicazione: (2025)
LatentFlowSR: High-Fidelity Audio Super-Resolution via Noise-Robust Latent Flow Matching
di: Liu, Fei, et al.
Pubblicazione: (2026)
di: Liu, Fei, et al.
Pubblicazione: (2026)
VoiceBridge: General Speech Restoration with One-step Latent Bridge Models
di: Zhang, Chi, et al.
Pubblicazione: (2025)
di: Zhang, Chi, et al.
Pubblicazione: (2025)
MuseAgent-1: Interactive Grounded Multimodal Understanding of Music Scores and Performance Audio
di: Zhao, Qihao, et al.
Pubblicazione: (2026)
di: Zhao, Qihao, et al.
Pubblicazione: (2026)
LongCat-AudioDiT: High-Fidelity Diffusion Text-to-Speech in the Waveform Latent Space
di: Xin, Detai, et al.
Pubblicazione: (2026)
di: Xin, Detai, et al.
Pubblicazione: (2026)
Learning Interpretable Features in Audio Latent Spaces via Sparse Autoencoders
di: Paek, Nathan, et al.
Pubblicazione: (2025)
di: Paek, Nathan, et al.
Pubblicazione: (2025)
SightSound-R1: Cross-Modal Reasoning Distillation from Vision to Audio Language Models
di: Wang, Qiaolin, et al.
Pubblicazione: (2025)
di: Wang, Qiaolin, et al.
Pubblicazione: (2025)
CORD: Bridging the Audio-Text Reasoning Gap via Weighted On-policy Cross-modal Distillation
di: Hu, Jing, et al.
Pubblicazione: (2026)
di: Hu, Jing, et al.
Pubblicazione: (2026)
GSDNet: Revisiting Incomplete Multimodal-Diffusion from Graph Spectrum Perspective for Conversation Emotion Recognition
di: Shou, Yuntao, et al.
Pubblicazione: (2025)
di: Shou, Yuntao, et al.
Pubblicazione: (2025)
Evaluating Latent Space Structure in Timbre VAEs: A Comparative Study of Unsupervised, Descriptor-Conditioned, and Perceptual Feature-Conditioned Models
di: Cameron, Joseph, et al.
Pubblicazione: (2026)
di: Cameron, Joseph, et al.
Pubblicazione: (2026)
Composer Vector: Style-steering Symbolic Music Generation in a Latent Space
di: Jiang, Xunyi, et al.
Pubblicazione: (2026)
di: Jiang, Xunyi, et al.
Pubblicazione: (2026)
Advancing Singlish Understanding: Bridging the Gap with Datasets and Multimodal Models
di: Wang, Bin, et al.
Pubblicazione: (2025)
di: Wang, Bin, et al.
Pubblicazione: (2025)
Exploring Token-Space Manipulation in Latent Audio Tokenizers
di: Paissan, Francesco, et al.
Pubblicazione: (2026)
di: Paissan, Francesco, et al.
Pubblicazione: (2026)
Generalizable Audio Deepfake Detection via Latent Space Refinement and Augmentation
di: Huang, Wen, et al.
Pubblicazione: (2025)
di: Huang, Wen, et al.
Pubblicazione: (2025)
Meta-Learn Unimodal Signals with Weak Supervision for Multimodal Sentiment Analysis
di: Mai, Sijie, et al.
Pubblicazione: (2024)
di: Mai, Sijie, et al.
Pubblicazione: (2024)
SemanticVocoder: Bridging Audio Generation and Audio Understanding via Semantic Latents
di: Xie, Zeyu, et al.
Pubblicazione: (2026)
di: Xie, Zeyu, et al.
Pubblicazione: (2026)
Sparsely Shared LoRA on Whisper for Child Speech Recognition
di: Liu, Wei, et al.
Pubblicazione: (2023)
di: Liu, Wei, et al.
Pubblicazione: (2023)
ZSVC: Zero-shot Style Voice Conversion with Disentangled Latent Diffusion Models and Adversarial Training
di: Zhu, Xinfa, et al.
Pubblicazione: (2025)
di: Zhu, Xinfa, et al.
Pubblicazione: (2025)
Schrodinger Bridges Beat Diffusion Models on Text-to-Speech Synthesis
di: Chen, Zehua, et al.
Pubblicazione: (2023)
di: Chen, Zehua, et al.
Pubblicazione: (2023)
CodecFlow: Efficient Bandwidth Extension via Conditional Flow Matching in Neural Codec Latent Space
di: Zhang, Bowen, et al.
Pubblicazione: (2026)
di: Zhang, Bowen, et al.
Pubblicazione: (2026)
Audio-Plane: Audio Factorization Plane Gaussian Splatting for Real-Time Talking Head Synthesis
di: Shen, Shuai, et al.
Pubblicazione: (2025)
di: Shen, Shuai, et al.
Pubblicazione: (2025)
Investigating the Invertibility of Multimodal Latent Spaces: Limitations of Optimization-Based Methods
di: Park, Siwoo
Pubblicazione: (2025)
di: Park, Siwoo
Pubblicazione: (2025)
MAGIC-TTS: Fine-Grained Controllable Speech Synthesis with Explicit Local Duration and Pause Control
di: Mai, Jialong, et al.
Pubblicazione: (2026)
di: Mai, Jialong, et al.
Pubblicazione: (2026)
Text-Driven Voice Conversion via Latent State-Space Modeling
di: Li, Wen, et al.
Pubblicazione: (2025)
di: Li, Wen, et al.
Pubblicazione: (2025)
When Tone and Words Disagree: Towards Robust Speech Emotion Recognition under Acoustic-Semantic Conflict
di: Huang, Dawei, et al.
Pubblicazione: (2026)
di: Huang, Dawei, et al.
Pubblicazione: (2026)
Latent-Mark: An Audio Watermark Robust to Neural Resynthesis
di: Chen, Yen-Shan, et al.
Pubblicazione: (2026)
di: Chen, Yen-Shan, et al.
Pubblicazione: (2026)
Enterprise Sales Copilot: Enabling Real-Time AI Support with Automatic Information Retrieval in Live Sales Calls
di: Qiu, Jielin, et al.
Pubblicazione: (2026)
di: Qiu, Jielin, et al.
Pubblicazione: (2026)
Latent Diffusion Bridges for Unsupervised Musical Audio Timbre Transfer
di: Mancusi, Michele, et al.
Pubblicazione: (2024)
di: Mancusi, Michele, et al.
Pubblicazione: (2024)
TLDiffGAN: A Latent Diffusion-GAN Framework with Temporal Information Fusion for Anomalous Sound Detection
di: Ma, Chengyuan, et al.
Pubblicazione: (2026)
di: Ma, Chengyuan, et al.
Pubblicazione: (2026)
EE-TTS: Emphatic Expressive TTS with Linguistic Information
di: Zhong, Yi, et al.
Pubblicazione: (2023)
di: Zhong, Yi, et al.
Pubblicazione: (2023)
Latent Multi-view Learning for Robust Environmental Sound Representations
di: Ding, Sivan, et al.
Pubblicazione: (2025)
di: Ding, Sivan, et al.
Pubblicazione: (2025)
Multimodal Contrastive Learning via Uni-Modal Coding and Cross-Modal Prediction for Multimodal Sentiment Analysis
di: Lin, Ronghao, et al.
Pubblicazione: (2022)
di: Lin, Ronghao, et al.
Pubblicazione: (2022)
Variational Bayesian Adaptive Learning of Deep Latent Variables for Acoustic Knowledge Transfer
di: Hu, Hu, et al.
Pubblicazione: (2025)
di: Hu, Hu, et al.
Pubblicazione: (2025)
Tiny-Align: Bridging Automatic Speech Recognition and Large Language Model on the Edge
di: Qin, Ruiyang, et al.
Pubblicazione: (2024)
di: Qin, Ruiyang, et al.
Pubblicazione: (2024)
Documenti analoghi
-
Multi-source Multimodal Progressive Domain Adaption for Audio-Visual Deception Detection
di: Lin, Ronghao, et al.
Pubblicazione: (2025) -
MissMAC-Bench: Building Solid Benchmark for Missing Modality Issue in Robust Multimodal Affective Computing
di: Lin, Ronghao, et al.
Pubblicazione: (2026) -
E3RG: Building Explicit Emotion-driven Empathetic Response Generation System with Multimodal Large Language Model
di: Lin, Ronghao, et al.
Pubblicazione: (2025) -
End-to-end Semantic-centric Video-based Multimodal Affective Computing
di: Lin, Ronghao, et al.
Pubblicazione: (2024) -
Audio Super-Resolution with Latent Bridge Models
di: Li, Chang, et al.
Pubblicazione: (2025)