Explore More, Learn Better: Parallel MLLM Embeddings under Mutual Information Minimization
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Wang, Zhicheng, Ju, Chen, Chen, Xu, Xiao, Shuai, Lan, Jinsong, Zhu, Xiaoyong, Chen, Ying, Cao, Zhiguo |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Wave-Particle (Continuous-Discrete) Dualistic Visual Tokenization for Unified Understanding and Generation
von: Chen, Yizhu, et al.
Veröffentlicht: (2025)
von: Chen, Yizhu, et al.
Veröffentlicht: (2025)
Cell Variational Information Bottleneck Network
von: Zhai, Zhonghua, et al.
Veröffentlicht: (2024)
von: Zhai, Zhonghua, et al.
Veröffentlicht: (2024)
Wear-Any-Way: Manipulable Virtual Try-on via Sparse Correspondence Alignment
von: Chen, Mengting, et al.
Veröffentlicht: (2024)
von: Chen, Mengting, et al.
Veröffentlicht: (2024)
Turbo: Informativity-Driven Acceleration Plug-In for Vision-Language Large Models
von: Ju, Chen, et al.
Veröffentlicht: (2024)
von: Ju, Chen, et al.
Veröffentlicht: (2024)
Pailitao-VL: Unified Embedding and Reranker for Real-Time Multi-Modal Industrial Search
von: Chen, Lei, et al.
Veröffentlicht: (2026)
von: Chen, Lei, et al.
Veröffentlicht: (2026)
Cross-modal Identity Mapping: Minimizing Information Loss in Modality Conversion via Reinforcement Learning
von: Jia, Haonan, et al.
Veröffentlicht: (2026)
von: Jia, Haonan, et al.
Veröffentlicht: (2026)
Improving Human Image Animation via Semantic Representation Alignment
von: Liu, Chang, et al.
Veröffentlicht: (2026)
von: Liu, Chang, et al.
Veröffentlicht: (2026)
FashionChameleon: Towards Real-Time and Interactive Human-Garment Video Customization
von: Song, Quanjian, et al.
Veröffentlicht: (2026)
von: Song, Quanjian, et al.
Veröffentlicht: (2026)
Vision Transformer Off-the-Shelf: A Surprising Baseline for Few-Shot Class-Agnostic Counting
von: Wang, Zhicheng, et al.
Veröffentlicht: (2023)
von: Wang, Zhicheng, et al.
Veröffentlicht: (2023)
Exploring MLLM-Diffusion Information Transfer with MetaCanvas
von: Lin, Han, et al.
Veröffentlicht: (2025)
von: Lin, Han, et al.
Veröffentlicht: (2025)
Advancing Myopia To Holism: Fully Contrastive Language-Image Pre-training
von: Wang, Haicheng, et al.
Veröffentlicht: (2024)
von: Wang, Haicheng, et al.
Veröffentlicht: (2024)
Exploring Contextual Attribute Density in Referring Expression Counting
von: Wang, Zhicheng, et al.
Veröffentlicht: (2025)
von: Wang, Zhicheng, et al.
Veröffentlicht: (2025)
Better, Stronger, Faster: Tackling the Trilemma in MLLM-based Segmentation with Simultaneous Textual Mask Prediction
von: Liu, Jiazhen, et al.
Veröffentlicht: (2025)
von: Liu, Jiazhen, et al.
Veröffentlicht: (2025)
Mutual Information guided Visual Contrastive Learning
von: Chen, Hanyang, et al.
Veröffentlicht: (2025)
von: Chen, Hanyang, et al.
Veröffentlicht: (2025)
iTryOn: Mastering Interactive Video Virtual Try-On with Spatial-Semantic Guidance
von: Zheng, Jun, et al.
Veröffentlicht: (2026)
von: Zheng, Jun, et al.
Veröffentlicht: (2026)
Squeeze Out Tokens from Sample for Finer-Grained Data Governance
von: Lin, Weixiong, et al.
Veröffentlicht: (2025)
von: Lin, Weixiong, et al.
Veröffentlicht: (2025)
LIME: Less Is More for MLLM Evaluation
von: Zhu, King, et al.
Veröffentlicht: (2024)
von: Zhu, King, et al.
Veröffentlicht: (2024)
More is Better: Deep Domain Adaptation with Multiple Sources
von: Zhao, Sicheng, et al.
Veröffentlicht: (2024)
von: Zhao, Sicheng, et al.
Veröffentlicht: (2024)
Learning Disentangled Representations for Perceptual Point Cloud Quality Assessment via Mutual Information Minimization
von: Shan, Ziyu, et al.
Veröffentlicht: (2024)
von: Shan, Ziyu, et al.
Veröffentlicht: (2024)
MR-MLLM: Mutual Reinforcement of Multimodal Comprehension and Vision Perception
von: Wang, Guanqun, et al.
Veröffentlicht: (2024)
von: Wang, Guanqun, et al.
Veröffentlicht: (2024)
The Evolution of Video Anomaly Detection: A Unified Framework from DNN to MLLM
von: Gao, Shibo, et al.
Veröffentlicht: (2025)
von: Gao, Shibo, et al.
Veröffentlicht: (2025)
Counterfactual Learning-Driven Representation Disentanglement for Search-Enhanced Recommendation
von: Cui, Jiajun, et al.
Veröffentlicht: (2024)
von: Cui, Jiajun, et al.
Veröffentlicht: (2024)
MoCa: Modality-aware Continual Pre-training Makes Better Bidirectional Multimodal Embeddings
von: Chen, Haonan, et al.
Veröffentlicht: (2025)
von: Chen, Haonan, et al.
Veröffentlicht: (2025)
Reasoning Guided Embeddings: Leveraging MLLM Reasoning for Improved Multimodal Retrieval
von: Liu, Chunxu, et al.
Veröffentlicht: (2025)
von: Liu, Chunxu, et al.
Veröffentlicht: (2025)
Complementary Information Mutual Learning for Multimodality Medical Image Segmentation
von: Shen, Chuyun, et al.
Veröffentlicht: (2024)
von: Shen, Chuyun, et al.
Veröffentlicht: (2024)
IPCV: Information-Preserving Compression for MLLM Visual Encoders
von: Chen, Yuan, et al.
Veröffentlicht: (2025)
von: Chen, Yuan, et al.
Veröffentlicht: (2025)
PointSmile: Point Self-supervised Learning via Curriculum Mutual Information
von: Li, Xin, et al.
Veröffentlicht: (2023)
von: Li, Xin, et al.
Veröffentlicht: (2023)
Tunnel Try-on: Excavating Spatial-temporal Tunnels for High-quality Virtual Try-on in Videos
von: Xu, Zhengze, et al.
Veröffentlicht: (2024)
von: Xu, Zhengze, et al.
Veröffentlicht: (2024)
Efficient Diffusion Distillation via Embedding Loss
von: Ying, Jincheng, et al.
Veröffentlicht: (2026)
von: Ying, Jincheng, et al.
Veröffentlicht: (2026)
Co-Seg: Mutual Prompt-Guided Collaborative Learning for Tissue and Nuclei Segmentation
von: Xu, Qing, et al.
Veröffentlicht: (2025)
von: Xu, Qing, et al.
Veröffentlicht: (2025)
Empowering Source-Free Domain Adaptation via MLLM-Guided Reliability-Based Curriculum Learning
von: Chen, Dongjie, et al.
Veröffentlicht: (2024)
von: Chen, Dongjie, et al.
Veröffentlicht: (2024)
ABC: Achieving Better Control of Multimodal Embeddings using VLMs
von: Schneider, Benjamin, et al.
Veröffentlicht: (2025)
von: Schneider, Benjamin, et al.
Veröffentlicht: (2025)
More Is Better: A MoE-Based Emotion Recognition Framework with Human Preference Alignment
von: Xie, Jun, et al.
Veröffentlicht: (2025)
von: Xie, Jun, et al.
Veröffentlicht: (2025)
ForensicZip: More Tokens are Better but Not Necessary in Forensic Vision-Language Models
von: Lai, Yingxin, et al.
Veröffentlicht: (2026)
von: Lai, Yingxin, et al.
Veröffentlicht: (2026)
Leveraging MLLM Embeddings and Attribute Smoothing for Compositional Zero-Shot Learning
von: Yan, Xudong, et al.
Veröffentlicht: (2024)
von: Yan, Xudong, et al.
Veröffentlicht: (2024)
LLaVE: Large Language and Vision Embedding Models with Hardness-Weighted Contrastive Learning
von: Lan, Zhibin, et al.
Veröffentlicht: (2025)
von: Lan, Zhibin, et al.
Veröffentlicht: (2025)
Improving Joint Embedding Predictive Architecture with Diffusion Noise
von: Qiu, Yuping, et al.
Veröffentlicht: (2025)
von: Qiu, Yuping, et al.
Veröffentlicht: (2025)
End-to-end Video Gaze Estimation via Capturing Head-face-eye Spatial-temporal Interaction Context
von: Guan, Yiran, et al.
Veröffentlicht: (2023)
von: Guan, Yiran, et al.
Veröffentlicht: (2023)
Creation-MMBench: Assessing Context-Aware Creative Intelligence in MLLM
von: Fang, Xinyu, et al.
Veröffentlicht: (2025)
von: Fang, Xinyu, et al.
Veröffentlicht: (2025)
Co-Seg++: Mutual Prompt-Guided Collaborative Learning for Versatile Medical Segmentation
von: Xu, Qing, et al.
Veröffentlicht: (2025)
von: Xu, Qing, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Wave-Particle (Continuous-Discrete) Dualistic Visual Tokenization for Unified Understanding and Generation
von: Chen, Yizhu, et al.
Veröffentlicht: (2025) -
Cell Variational Information Bottleneck Network
von: Zhai, Zhonghua, et al.
Veröffentlicht: (2024) -
Wear-Any-Way: Manipulable Virtual Try-on via Sparse Correspondence Alignment
von: Chen, Mengting, et al.
Veröffentlicht: (2024) -
Turbo: Informativity-Driven Acceleration Plug-In for Vision-Language Large Models
von: Ju, Chen, et al.
Veröffentlicht: (2024) -
Pailitao-VL: Unified Embedding and Reranker for Real-Time Multi-Modal Industrial Search
von: Chen, Lei, et al.
Veröffentlicht: (2026)