Principled Multimodal Representation Learning
Fuente:
arXiv
Saved in:
| Main Authors: | Liu, Xiaohao, Xia, Xiaobo, Ng, See-Kiong, Chua, Tat-Seng |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Calibrated Multimodal Representation Learning with Missing Modalities
by: Liu, Xiaohao, et al.
Published: (2025)
by: Liu, Xiaohao, et al.
Published: (2025)
Towards Modality Generalization: A Benchmark and Prospective Analysis
by: Liu, Xiaohao, et al.
Published: (2024)
by: Liu, Xiaohao, et al.
Published: (2024)
Extending Visual Dynamics for Video-to-Music Generation
by: Liu, Xiaohao, et al.
Published: (2025)
by: Liu, Xiaohao, et al.
Published: (2025)
Continual Multimodal Contrastive Learning
by: Liu, Xiaohao, et al.
Published: (2025)
by: Liu, Xiaohao, et al.
Published: (2025)
SILMM: Self-Improving Large Multimodal Models for Compositional Text-to-Image Generation
by: Qu, Leigang, et al.
Published: (2024)
by: Qu, Leigang, et al.
Published: (2024)
TIGeR: Unifying Text-to-Image Generation and Retrieval with Large Multimodal Models
by: Qu, Leigang, et al.
Published: (2024)
by: Qu, Leigang, et al.
Published: (2024)
D-Judge: How Far Are We? Assessing the Discrepancies Between AI-synthesized and Natural Images through Multimodal Guidance
by: Liu, Renyang, et al.
Published: (2024)
by: Liu, Renyang, et al.
Published: (2024)
FutureOmni: Evaluating Future Forecasting from Omni-Modal Context for Multimodal LLMs
by: Chen, Qian, et al.
Published: (2026)
by: Chen, Qian, et al.
Published: (2026)
Dynamic Multimodal Fusion via Meta-Learning Towards Micro-Video Recommendation
by: Liu, Han, et al.
Published: (2025)
by: Liu, Han, et al.
Published: (2025)
TTOM: Test-Time Optimization and Memorization for Compositional Video Generation
by: Qu, Leigang, et al.
Published: (2025)
by: Qu, Leigang, et al.
Published: (2025)
Turing Patterns for Multimedia: Reaction-Diffusion Multi-Modal Fusion for Language-Guided Video Moment Retrieval
by: Fang, Xiang, et al.
Published: (2026)
by: Fang, Xiang, et al.
Published: (2026)
Data relativistic uncertainty framework for low-illumination anime scenery image enhancement
by: Gao, Yiquan, et al.
Published: (2025)
by: Gao, Yiquan, et al.
Published: (2025)
Logic Unseen: Revealing the Logical Blindspots of Vision-Language Models
by: Zhou, Yuchen, et al.
Published: (2025)
by: Zhou, Yuchen, et al.
Published: (2025)
Towards Natural Language-Guided Drones: GeoText-1652 Benchmark with Spatial Relation Matching
by: Chu, Meng, et al.
Published: (2023)
by: Chu, Meng, et al.
Published: (2023)
GalleryGPT: Analyzing Paintings with Large Multimodal Models
by: Bin, Yi, et al.
Published: (2024)
by: Bin, Yi, et al.
Published: (2024)
ExpLLM: Towards Chain of Thought for Facial Expression Recognition
by: Lan, Xing, et al.
Published: (2024)
by: Lan, Xing, et al.
Published: (2024)
Can I Trust Your Answer? Visually Grounded Video Question Answering
by: Xiao, Junbin, et al.
Published: (2023)
by: Xiao, Junbin, et al.
Published: (2023)
Hierarchical Adaptive Expert for Multimodal Sentiment Analysis
by: Qin, Jiahao, et al.
Published: (2025)
by: Qin, Jiahao, et al.
Published: (2025)
Multimodal Learning on Low-Quality Data with Conformal Predictive Self-Calibration
by: Jiang, Xun, et al.
Published: (2026)
by: Jiang, Xun, et al.
Published: (2026)
VINCIE: Unlocking In-context Image Editing from Video
by: Qu, Leigang, et al.
Published: (2025)
by: Qu, Leigang, et al.
Published: (2025)
DeCo-VAE: Learning Compact Latents for Video Reconstruction via Decoupled Representation
by: Yin, Xiangchen, et al.
Published: (2025)
by: Yin, Xiangchen, et al.
Published: (2025)
Learning from Mistakes: Self-Regularizing Hierarchical Representations in Point Cloud Semantic Segmentation
by: Camuffo, Elena, et al.
Published: (2023)
by: Camuffo, Elena, et al.
Published: (2023)
NExT-OMNI: Towards Any-to-Any Omnimodal Foundation Models with Discrete Flow Matching
by: Luo, Run, et al.
Published: (2025)
by: Luo, Run, et al.
Published: (2025)
PetalView: Fine-grained Location and Orientation Extraction of Street-view Images via Cross-view Local Search with Supplementary Materials
by: Hu, Wenmiao, et al.
Published: (2024)
by: Hu, Wenmiao, et al.
Published: (2024)
InstructVid2Vid: Controllable Video Editing with Natural Language Instructions
by: Qin, Bosheng, et al.
Published: (2023)
by: Qin, Bosheng, et al.
Published: (2023)
Multimodal Transformer With a Low-Computational-Cost Guarantee
by: Park, Sungjin, et al.
Published: (2024)
by: Park, Sungjin, et al.
Published: (2024)
Scene-Text Grounding for Text-Based Video Question Answering
by: Zhou, Sheng, et al.
Published: (2024)
by: Zhou, Sheng, et al.
Published: (2024)
Bridging Compressed Image Latents and Multimodal Large Language Models
by: Kao, Chia-Hao, et al.
Published: (2024)
by: Kao, Chia-Hao, et al.
Published: (2024)
Embracing Collaboration Over Competition: Condensing Multiple Prompts for Visual In-Context Learning
by: Wang, Jinpeng, et al.
Published: (2025)
by: Wang, Jinpeng, et al.
Published: (2025)
Catalogue Grounded Multimodal Attribution for Museum Video under Resource and Regulatory Constraints
by: Nanang, Minsak, et al.
Published: (2026)
by: Nanang, Minsak, et al.
Published: (2026)
4D Multimodal Co-attention Fusion Network with Latent Contrastive Alignment for Alzheimer's Diagnosis
by: Wei, Yuxiang, et al.
Published: (2025)
by: Wei, Yuxiang, et al.
Published: (2025)
EgoTextVQA: Towards Egocentric Scene-Text Aware Video Question Answering
by: Zhou, Sheng, et al.
Published: (2025)
by: Zhou, Sheng, et al.
Published: (2025)
On-the-fly Modulation for Balanced Multimodal Learning
by: Wei, Yake, et al.
Published: (2024)
by: Wei, Yake, et al.
Published: (2024)
ReFiNe: Recursive Field Networks for Cross-modal Multi-scene Representation
by: Zakharov, Sergey, et al.
Published: (2024)
by: Zakharov, Sergey, et al.
Published: (2024)
MMPareto: Boosting Multimodal Learning with Innocent Unimodal Assistance
by: Wei, Yake, et al.
Published: (2024)
by: Wei, Yake, et al.
Published: (2024)
Deep Learning-based Text-in-Image Watermarking
by: Karki, Bishwa, et al.
Published: (2024)
by: Karki, Bishwa, et al.
Published: (2024)
Omnidirectional Video Super-Resolution using Deep Learning
by: Baniya, Arbind Agrahari, et al.
Published: (2025)
by: Baniya, Arbind Agrahari, et al.
Published: (2025)
X-Former: Unifying Contrastive and Reconstruction Learning for MLLMs
by: Swetha, Sirnam, et al.
Published: (2024)
by: Swetha, Sirnam, et al.
Published: (2024)
Calibrating Multimodal Consensus for Emotion Recognition
by: Zhong, Guowei, et al.
Published: (2025)
by: Zhong, Guowei, et al.
Published: (2025)
Text-Guided Image Invariant Feature Learning for Robust Image Watermarking
by: Ahtesham, Muhammad, et al.
Published: (2025)
by: Ahtesham, Muhammad, et al.
Published: (2025)
Similar Items
-
Calibrated Multimodal Representation Learning with Missing Modalities
by: Liu, Xiaohao, et al.
Published: (2025) -
Towards Modality Generalization: A Benchmark and Prospective Analysis
by: Liu, Xiaohao, et al.
Published: (2024) -
Extending Visual Dynamics for Video-to-Music Generation
by: Liu, Xiaohao, et al.
Published: (2025) -
Continual Multimodal Contrastive Learning
by: Liu, Xiaohao, et al.
Published: (2025) -
SILMM: Self-Improving Large Multimodal Models for Compositional Text-to-Image Generation
by: Qu, Leigang, et al.
Published: (2024)