Breaking the Modality Barrier: Universal Embedding Learning with Multimodal LLMs
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Gu, Tiancheng, Yang, Kaicheng, Feng, Ziyong, Wang, Xingjun, Zhang, Yanzhao, Long, Dingkun, Chen, Yingda, Cai, Weidong, Deng, Jiankang |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
UniME-V2: MLLM-as-a-Judge for Universal Multimodal Embedding Learning
von: Gu, Tiancheng, et al.
Veröffentlicht: (2025)
von: Gu, Tiancheng, et al.
Veröffentlicht: (2025)
RWKV-CLIP: A Robust Vision-Language Representation Learner
von: Gu, Tiancheng, et al.
Veröffentlicht: (2024)
von: Gu, Tiancheng, et al.
Veröffentlicht: (2024)
RealSyn: An Effective and Scalable Multimodal Interleaved Document Transformation Paradigm
von: Gu, Tiancheng, et al.
Veröffentlicht: (2025)
von: Gu, Tiancheng, et al.
Veröffentlicht: (2025)
CLIP-CID: Efficient CLIP Distillation via Cluster-Instance Discrimination
von: Yang, Kaicheng, et al.
Veröffentlicht: (2024)
von: Yang, Kaicheng, et al.
Veröffentlicht: (2024)
ORID: Organ-Regional Information Driven Framework for Radiology Report Generation
von: Gu, Tiancheng, et al.
Veröffentlicht: (2024)
von: Gu, Tiancheng, et al.
Veröffentlicht: (2024)
Multi-label Cluster Discrimination for Visual Representation Learning
von: An, Xiang, et al.
Veröffentlicht: (2024)
von: An, Xiang, et al.
Veröffentlicht: (2024)
LaPA: Latent Prompt Assist Model For Medical Visual Question Answering
von: Gu, Tiancheng, et al.
Veröffentlicht: (2024)
von: Gu, Tiancheng, et al.
Veröffentlicht: (2024)
Towards Universal Video Retrieval: Generalizing Video Embedding via Synthesized Multimodal Pyramid Curriculum
von: Guo, Zhuoning, et al.
Veröffentlicht: (2025)
von: Guo, Zhuoning, et al.
Veröffentlicht: (2025)
High-Fidelity Facial Albedo Estimation via Texture Quantization
von: Ran, Zimin, et al.
Veröffentlicht: (2024)
von: Ran, Zimin, et al.
Veröffentlicht: (2024)
PaCo-FR: Patch-Pixel Aligned End-to-End Codebook Learning for Facial Representation Pre-training
von: Xie, Yin, et al.
Veröffentlicht: (2025)
von: Xie, Yin, et al.
Veröffentlicht: (2025)
UniDoc-RL: Coarse-to-Fine Visual RAG with Hierarchical Actions and Dense Rewards
von: Wang, Jun, et al.
Veröffentlicht: (2026)
von: Wang, Jun, et al.
Veröffentlicht: (2026)
Region-based Cluster Discrimination for Visual Representation Learning
von: Xie, Yin, et al.
Veröffentlicht: (2025)
von: Xie, Yin, et al.
Veröffentlicht: (2025)
ViCToR: Improving Visual Comprehension via Token Reconstruction for Pretraining LMMs
von: Xie, Yin, et al.
Veröffentlicht: (2024)
von: Xie, Yin, et al.
Veröffentlicht: (2024)
DanQing: An Up-to-Date Large-Scale Chinese Vision-Language Pre-training Dataset
von: Shen, Hengyu, et al.
Veröffentlicht: (2026)
von: Shen, Hengyu, et al.
Veröffentlicht: (2026)
IDAdapter: Learning Mixed Features for Tuning-Free Personalization of Text-to-Image Models
von: Cui, Siying, et al.
Veröffentlicht: (2024)
von: Cui, Siying, et al.
Veröffentlicht: (2024)
1st Place Solution to the 1st SkatingVerse Challenge
von: Sun, Tao, et al.
Veröffentlicht: (2024)
von: Sun, Tao, et al.
Veröffentlicht: (2024)
Gradient-Attention Guided Dual-Masking Synergetic Framework for Robust Text-based Person Retrieval
von: Zheng, Tianlu, et al.
Veröffentlicht: (2025)
von: Zheng, Tianlu, et al.
Veröffentlicht: (2025)
Decoupled Global-Local Alignment for Improving Compositional Understanding
von: Hu, Xiaoxing, et al.
Veröffentlicht: (2025)
von: Hu, Xiaoxing, et al.
Veröffentlicht: (2025)
Efficient, Validation-Free Intrinsic Quality Estimation for Large-Scale Face Recognition Datasets
von: Chen, Zhichao, et al.
Veröffentlicht: (2026)
von: Chen, Zhichao, et al.
Veröffentlicht: (2026)
OneVision-Encoder: Codec-Aligned Sparsity as a Foundational Principle for Multimodal Intelligence
von: Tang, Feilong, et al.
Veröffentlicht: (2026)
von: Tang, Feilong, et al.
Veröffentlicht: (2026)
Supervised Fine-Tuning or Contrastive Learning? Towards Better Multimodal LLM Reranking
von: Dai, Ziqi, et al.
Veröffentlicht: (2025)
von: Dai, Ziqi, et al.
Veröffentlicht: (2025)
ForCenNet: Foreground-Centric Network for Document Image Rectification
von: Cai, Peng, et al.
Veröffentlicht: (2025)
von: Cai, Peng, et al.
Veröffentlicht: (2025)
EliGen: Entity-Level Controlled Image Generation with Regional Attention
von: Zhang, Hong, et al.
Veröffentlicht: (2025)
von: Zhang, Hong, et al.
Veröffentlicht: (2025)
Modality Curation: Building Universal Embeddings for Advanced Multimodal Information Retrieval
von: Kong, Fanheng, et al.
Veröffentlicht: (2025)
von: Kong, Fanheng, et al.
Veröffentlicht: (2025)
Optimizing Multimodal LLMs for Egocentric Video Understanding: A Solution for the HD-EPIC VQA Challenge
von: Yang, Sicheng, et al.
Veröffentlicht: (2026)
von: Yang, Sicheng, et al.
Veröffentlicht: (2026)
Multimodal LLMs under Pairwise Modalities
von: Li, Yan, et al.
Veröffentlicht: (2026)
von: Li, Yan, et al.
Veröffentlicht: (2026)
ProCLIP: Progressive Vision-Language Alignment via LLM-based Embedder
von: Hu, Xiaoxing, et al.
Veröffentlicht: (2025)
von: Hu, Xiaoxing, et al.
Veröffentlicht: (2025)
LLaVA-OneVision-1.5: Fully Open Framework for Democratized Multimodal Training
von: An, Xiang, et al.
Veröffentlicht: (2025)
von: An, Xiang, et al.
Veröffentlicht: (2025)
Tracing and Mitigating Hallucinations in Multimodal LLMs via Dynamic Attention Localization
von: Yang, Tiancheng, et al.
Veröffentlicht: (2025)
von: Yang, Tiancheng, et al.
Veröffentlicht: (2025)
Multimodal Causal Reasoning Benchmark: Challenging Vision Large Language Models to Discern Causal Links Across Modalities
von: Li, Zhiyuan, et al.
Veröffentlicht: (2024)
von: Li, Zhiyuan, et al.
Veröffentlicht: (2024)
Nexus-Gen: Unified Image Understanding, Generation, and Editing via Prefilled Autoregression in Shared Embedding Space
von: Zhang, Hong, et al.
Veröffentlicht: (2025)
von: Zhang, Hong, et al.
Veröffentlicht: (2025)
Elevating Visual Perception in Multimodal LLMs with Visual Embedding Distillation
von: Jain, Jitesh, et al.
Veröffentlicht: (2024)
von: Jain, Jitesh, et al.
Veröffentlicht: (2024)
RedDiffuser: Auditing Multimodal Safety Failures in Vision-Language Models via Reinforced Diffusion
von: Wang, Ruofan, et al.
Veröffentlicht: (2025)
von: Wang, Ruofan, et al.
Veröffentlicht: (2025)
Plug-and-Play Grounding of Reasoning in Multimodal Large Language Models
von: Chen, Jiaxing, et al.
Veröffentlicht: (2024)
von: Chen, Jiaxing, et al.
Veröffentlicht: (2024)
MM-Embed: Universal Multimodal Retrieval with Multimodal LLMs
von: Lin, Sheng-Chieh, et al.
Veröffentlicht: (2024)
von: Lin, Sheng-Chieh, et al.
Veröffentlicht: (2024)
ObjEmbed: Towards Universal Multimodal Object Embeddings
von: Fu, Shenghao, et al.
Veröffentlicht: (2026)
von: Fu, Shenghao, et al.
Veröffentlicht: (2026)
Breaking the Discretization Barrier of Continuous Physics Simulation Learning
von: Xu, Fan, et al.
Veröffentlicht: (2025)
von: Xu, Fan, et al.
Veröffentlicht: (2025)
Learning Contrastive Multimodal Fusion with Improved Modality Dropout for Disease Detection and Prediction
von: Gu, Yi, et al.
Veröffentlicht: (2025)
von: Gu, Yi, et al.
Veröffentlicht: (2025)
PLUME: Latent Reasoning Based Universal Multimodal Embedding
von: He, Chenwei, et al.
Veröffentlicht: (2026)
von: He, Chenwei, et al.
Veröffentlicht: (2026)
Magic-MM-Embedding: Towards Visual-Token-Efficient Universal Multimodal Embedding with MLLMs
von: Li, Qi, et al.
Veröffentlicht: (2026)
von: Li, Qi, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
UniME-V2: MLLM-as-a-Judge for Universal Multimodal Embedding Learning
von: Gu, Tiancheng, et al.
Veröffentlicht: (2025) -
RWKV-CLIP: A Robust Vision-Language Representation Learner
von: Gu, Tiancheng, et al.
Veröffentlicht: (2024) -
RealSyn: An Effective and Scalable Multimodal Interleaved Document Transformation Paradigm
von: Gu, Tiancheng, et al.
Veröffentlicht: (2025) -
CLIP-CID: Efficient CLIP Distillation via Cluster-Instance Discrimination
von: Yang, Kaicheng, et al.
Veröffentlicht: (2024) -
ORID: Organ-Regional Information Driven Framework for Radiology Report Generation
von: Gu, Tiancheng, et al.
Veröffentlicht: (2024)