Semantic Residual for Multimodal Unified Discrete Representation
Fuente:
arXiv
Saved in:
| Main Authors: | Huang, Hai, Wang, Shulei, Xia, Yan |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Bridging Domain Generalization to Multimodal Domain Generalization via Unified Representations
by: Huang, Hai, et al.
Published: (2025)
by: Huang, Hai, et al.
Published: (2025)
Open-set Cross Modal Generalization via Multimodal Unified Representation
by: Huang, Hai, et al.
Published: (2025)
by: Huang, Hai, et al.
Published: (2025)
Peeking Behind the Curtains of Residual Learning
by: Zhang, Tunhou, et al.
Published: (2024)
by: Zhang, Tunhou, et al.
Published: (2024)
Toward Unified Multimodal Representation Learning for Autonomous Driving
by: Tao, Ximeng, et al.
Published: (2026)
by: Tao, Ximeng, et al.
Published: (2026)
Unified Multimodal Discrete Diffusion
by: Swerdlow, Alexander, et al.
Published: (2025)
by: Swerdlow, Alexander, et al.
Published: (2025)
Enhancing Multimodal Unified Representations for Cross Modal Generalization
by: Huang, Hai, et al.
Published: (2024)
by: Huang, Hai, et al.
Published: (2024)
On the Role of Discrete Tokenization in Visual Representation Learning
by: Du, Tianqi, et al.
Published: (2024)
by: Du, Tianqi, et al.
Published: (2024)
Towards Robust Multimodal Representation: A Unified Approach with Adaptive Experts and Alignment
by: Moradinasab, Nazanin, et al.
Published: (2025)
by: Moradinasab, Nazanin, et al.
Published: (2025)
Grouped Discrete Representation for Object-Centric Learning
by: Zhao, Rongzhen, et al.
Published: (2024)
by: Zhao, Rongzhen, et al.
Published: (2024)
AnyGPT: Unified Multimodal LLM with Discrete Sequence Modeling
by: Zhan, Jun, et al.
Published: (2024)
by: Zhan, Jun, et al.
Published: (2024)
Principled Multimodal Representation Learning
by: Liu, Xiaohao, et al.
Published: (2025)
by: Liu, Xiaohao, et al.
Published: (2025)
Semantic Concentration for Self-Supervised Dense Representations Learning
by: Wen, Peisong, et al.
Published: (2025)
by: Wen, Peisong, et al.
Published: (2025)
Muddit: Liberating Generation Beyond Text-to-Image with a Unified Discrete Diffusion Model
by: Shi, Qingyu, et al.
Published: (2025)
by: Shi, Qingyu, et al.
Published: (2025)
Unified Multimodal Uncertain Inference
by: Zhang, Dengjia, et al.
Published: (2026)
by: Zhang, Dengjia, et al.
Published: (2026)
Predictive Regularization Against Visual Representation Degradation in Multimodal Large Language Models
by: Wang, Enguang, et al.
Published: (2026)
by: Wang, Enguang, et al.
Published: (2026)
Text-Conditional JEPA for Learning Semantically Rich Visual Representations
by: Huang, Chen, et al.
Published: (2026)
by: Huang, Chen, et al.
Published: (2026)
A Geometric View of SRC: Learning Representations for Stable Residual Inference
by: Oikonomou, Vangelis P.
Published: (2026)
by: Oikonomou, Vangelis P.
Published: (2026)
Understanding Multimodal Hallucination with Parameter-Free Representation Alignment
by: Wang, Yueqian, et al.
Published: (2024)
by: Wang, Yueqian, et al.
Published: (2024)
Holistic Semantic Representation for Navigational Trajectory Generation
by: Cao, Ji, et al.
Published: (2025)
by: Cao, Ji, et al.
Published: (2025)
Calibrated Multimodal Representation Learning with Missing Modalities
by: Liu, Xiaohao, et al.
Published: (2025)
by: Liu, Xiaohao, et al.
Published: (2025)
Learning Unified Representations from Heterogeneous Data for Robust Heart Rate Modeling
by: Huang, Zhengdong, et al.
Published: (2025)
by: Huang, Zhengdong, et al.
Published: (2025)
Learning Visual-Semantic Subspace Representations
by: Moreira, Gabriel, et al.
Published: (2024)
by: Moreira, Gabriel, et al.
Published: (2024)
Efficient Prompting for Continual Adaptation to Missing Modalities
by: Guo, Zirun, et al.
Published: (2025)
by: Guo, Zirun, et al.
Published: (2025)
FORESEE: Multimodal and Multi-view Representation Learning for Robust Prediction of Cancer Survival
by: Pan, Liangrui, et al.
Published: (2024)
by: Pan, Liangrui, et al.
Published: (2024)
Self-Organising Neural Discrete Representation Learning à la Kohonen
by: Irie, Kazuki, et al.
Published: (2023)
by: Irie, Kazuki, et al.
Published: (2023)
Investigating Permutation-Invariant Discrete Representation Learning for Spatially Aligned Images
by: Stirling, Jamie S. J., et al.
Published: (2026)
by: Stirling, Jamie S. J., et al.
Published: (2026)
FuseLIP: Multimodal Embeddings via Early Fusion of Discrete Tokens
by: Schlarmann, Christian, et al.
Published: (2025)
by: Schlarmann, Christian, et al.
Published: (2025)
Riemannian Motion Generation: A Unified Framework for Human Motion Representation and Generation via Riemannian Flow Matching
by: Miao, Fangran, et al.
Published: (2026)
by: Miao, Fangran, et al.
Published: (2026)
MMSFormer: Multimodal Transformer for Material and Semantic Segmentation
by: Reza, Md Kaykobad, et al.
Published: (2023)
by: Reza, Md Kaykobad, et al.
Published: (2023)
Multimodal Adaptive Retrieval Augmented Generation through Internal Representation Learning
by: Du, Ruoshuang, et al.
Published: (2026)
by: Du, Ruoshuang, et al.
Published: (2026)
PRCL: Probabilistic Representation Contrastive Learning for Semi-Supervised Semantic Segmentation
by: Xie, Haoyu, et al.
Published: (2024)
by: Xie, Haoyu, et al.
Published: (2024)
StableSemantics: A Synthetic Language-Vision Dataset of Semantic Representations in Naturalistic Images
by: Zawar, Rushikesh, et al.
Published: (2024)
by: Zawar, Rushikesh, et al.
Published: (2024)
Less Data, Faster Convergence: Goal-Driven Data Optimization for Multimodal Instruction Tuning
by: Wu, Rujie, et al.
Published: (2026)
by: Wu, Rujie, et al.
Published: (2026)
Multimodal Representation Learning by Alternating Unimodal Adaptation
by: Zhang, Xiaohui, et al.
Published: (2023)
by: Zhang, Xiaohui, et al.
Published: (2023)
Residual Denoising Diffusion Models
by: Liu, Jiawei, et al.
Published: (2023)
by: Liu, Jiawei, et al.
Published: (2023)
Capturing Context-Aware Route Choice Semantics for Trajectory Representation Learning
by: Cao, Ji, et al.
Published: (2025)
by: Cao, Ji, et al.
Published: (2025)
Automated Learning of Semantic Embedding Representations for Diffusion Models
by: Jiang, Limai, et al.
Published: (2025)
by: Jiang, Limai, et al.
Published: (2025)
A Lightweight Brain-Inspired Machine Learning Framework for Coronary Angiography: Hybrid Neural Representation and Robust Learning Strategies
by: Xia, Jingsong, et al.
Published: (2026)
by: Xia, Jingsong, et al.
Published: (2026)
Anchors Aweigh! Sail for Optimal Unified Multi-Modal Representations
by: Jeong, Minoh, et al.
Published: (2024)
by: Jeong, Minoh, et al.
Published: (2024)
STARFlow2: Bridging Language Models and Normalizing Flows for Unified Multimodal Generation
by: Shen, Ying, et al.
Published: (2026)
by: Shen, Ying, et al.
Published: (2026)
Similar Items
-
Bridging Domain Generalization to Multimodal Domain Generalization via Unified Representations
by: Huang, Hai, et al.
Published: (2025) -
Open-set Cross Modal Generalization via Multimodal Unified Representation
by: Huang, Hai, et al.
Published: (2025) -
Peeking Behind the Curtains of Residual Learning
by: Zhang, Tunhou, et al.
Published: (2024) -
Toward Unified Multimodal Representation Learning for Autonomous Driving
by: Tao, Ximeng, et al.
Published: (2026) -
Unified Multimodal Discrete Diffusion
by: Swerdlow, Alexander, et al.
Published: (2025)