Multi-level Mixture of Experts for Multimodal Entity Linking
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Hu, Zhiwei, Gutiérrez-Basulto, Víctor, Xiang, Zhiliang, Li, Ru, Pan, Jeff Z. |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Multi-level Matching Network for Multimodal Entity Linking
von: Hu, Zhiwei, et al.
Veröffentlicht: (2024)
von: Hu, Zhiwei, et al.
Veröffentlicht: (2024)
Uni-MoE: Scaling Unified Multimodal LLMs with Mixture of Experts
von: Li, Yunxin, et al.
Veröffentlicht: (2024)
von: Li, Yunxin, et al.
Veröffentlicht: (2024)
Reverse Region-to-Entity Annotation for Pixel-Level Visual Entity Linking
von: Xu, Zhengfei, et al.
Veröffentlicht: (2024)
von: Xu, Zhengfei, et al.
Veröffentlicht: (2024)
Advancing Grounded Multimodal Named Entity Recognition via LLM-Based Reformulation and Box-Based Segmentation
von: Li, Jinyuan, et al.
Veröffentlicht: (2024)
von: Li, Jinyuan, et al.
Veröffentlicht: (2024)
Video Summarization: Towards Entity-Aware Captions
von: Ayyubi, Hammad A., et al.
Veröffentlicht: (2023)
von: Ayyubi, Hammad A., et al.
Veröffentlicht: (2023)
Can Multimodal LLMs See Materials Clearly? A Multimodal Benchmark on Materials Characterization
von: Lai, Zhengzhao, et al.
Veröffentlicht: (2025)
von: Lai, Zhengzhao, et al.
Veröffentlicht: (2025)
COTET: Cross-view Optimal Transport for Knowledge Graph Entity Typing
von: Hu, Zhiwei, et al.
Veröffentlicht: (2024)
von: Hu, Zhiwei, et al.
Veröffentlicht: (2024)
Can't See the Forest for the Trees: Benchmarking Multimodal Safety Awareness for Multimodal LLMs
von: Wang, Wenxuan, et al.
Veröffentlicht: (2025)
von: Wang, Wenxuan, et al.
Veröffentlicht: (2025)
Veagle: Advancements in Multimodal Representation Learning
von: Chawla, Rajat, et al.
Veröffentlicht: (2024)
von: Chawla, Rajat, et al.
Veröffentlicht: (2024)
Recurrence Meets Transformers for Universal Multimodal Retrieval
von: Caffagni, Davide, et al.
Veröffentlicht: (2025)
von: Caffagni, Davide, et al.
Veröffentlicht: (2025)
The Revolution of Multimodal Large Language Models: A Survey
von: Caffagni, Davide, et al.
Veröffentlicht: (2024)
von: Caffagni, Davide, et al.
Veröffentlicht: (2024)
Multimodal LLM-Guided Semantic Correction in Text-to-Image Diffusion
von: Lv, Zheqi, et al.
Veröffentlicht: (2025)
von: Lv, Zheqi, et al.
Veröffentlicht: (2025)
TimeLens: Rethinking Video Temporal Grounding with Multimodal LLMs
von: Zhang, Jun, et al.
Veröffentlicht: (2025)
von: Zhang, Jun, et al.
Veröffentlicht: (2025)
Multimodal LLMs Can Reason about Aesthetics in Zero-Shot
von: Jiang, Ruixiang, et al.
Veröffentlicht: (2025)
von: Jiang, Ruixiang, et al.
Veröffentlicht: (2025)
Learning Compact Vision Tokens for Efficient Large Multimodal Models
von: Tang, Hao, et al.
Veröffentlicht: (2025)
von: Tang, Hao, et al.
Veröffentlicht: (2025)
Recurrence-Enhanced Vision-and-Language Transformers for Robust Multimodal Document Retrieval
von: Caffagni, Davide, et al.
Veröffentlicht: (2025)
von: Caffagni, Davide, et al.
Veröffentlicht: (2025)
Mitigating Hallucinations in Multimodal LLMs via Object-aware Preference Optimization
von: Compagnoni, Alberto, et al.
Veröffentlicht: (2025)
von: Compagnoni, Alberto, et al.
Veröffentlicht: (2025)
Wiki-LLaVA: Hierarchical Retrieval-Augmented Generation for Multimodal LLMs
von: Caffagni, Davide, et al.
Veröffentlicht: (2024)
von: Caffagni, Davide, et al.
Veröffentlicht: (2024)
Magnifier Prompt: Tackling Multimodal Hallucination via Extremely Simple Instructions
von: Fu, Yuhan, et al.
Veröffentlicht: (2024)
von: Fu, Yuhan, et al.
Veröffentlicht: (2024)
PIN: A Knowledge-Intensive Dataset for Paired and Interleaved Multimodal Documents
von: Wang, Junjie, et al.
Veröffentlicht: (2024)
von: Wang, Junjie, et al.
Veröffentlicht: (2024)
Analyzing Diffusion and Autoregressive Vision Language Models in Multimodal Embedding Space
von: Wang, Zihang, et al.
Veröffentlicht: (2026)
von: Wang, Zihang, et al.
Veröffentlicht: (2026)
Personalizing Multimodal Large Language Models for Image Captioning: An Experimental Analysis
von: Bucciarelli, Davide, et al.
Veröffentlicht: (2024)
von: Bucciarelli, Davide, et al.
Veröffentlicht: (2024)
How to Bridge the Gap between Modalities: Survey on Multimodal Large Language Model
von: Song, Shezheng, et al.
Veröffentlicht: (2023)
von: Song, Shezheng, et al.
Veröffentlicht: (2023)
VideoConviction: A Multimodal Benchmark for Human Conviction and Stock Market Recommendations
von: Galarnyk, Michael, et al.
Veröffentlicht: (2025)
von: Galarnyk, Michael, et al.
Veröffentlicht: (2025)
RaTA-Tool: Retrieval-based Tool Selection with Multimodal Large Language Models
von: Mattioli, Gabriele, et al.
Veröffentlicht: (2026)
von: Mattioli, Gabriele, et al.
Veröffentlicht: (2026)
Both Text and Images Leaked! A Systematic Analysis of Data Contamination in Multimodal LLM
von: Song, Dingjie, et al.
Veröffentlicht: (2024)
von: Song, Dingjie, et al.
Veröffentlicht: (2024)
Seeing Beyond Words: Self-Supervised Visual Learning for Multimodal Large Language Models
von: Caffagni, Davide, et al.
Veröffentlicht: (2025)
von: Caffagni, Davide, et al.
Veröffentlicht: (2025)
Augmenting Multimodal LLMs with Self-Reflective Tokens for Knowledge-based Visual Question Answering
von: Cocchi, Federico, et al.
Veröffentlicht: (2024)
von: Cocchi, Federico, et al.
Veröffentlicht: (2024)
TheoremExplainAgent: Towards Video-based Multimodal Explanations for LLM Theorem Understanding
von: Ku, Max, et al.
Veröffentlicht: (2025)
von: Ku, Max, et al.
Veröffentlicht: (2025)
A Survey of Multimodal Large Language Model from A Data-centric Perspective
von: Bai, Tianyi, et al.
Veröffentlicht: (2024)
von: Bai, Tianyi, et al.
Veröffentlicht: (2024)
What Changed? Detecting and Evaluating Instruction-Guided Image Edits with Multimodal Large Language Models
von: Baraldi, Lorenzo, et al.
Veröffentlicht: (2025)
von: Baraldi, Lorenzo, et al.
Veröffentlicht: (2025)
Mitigating Hallucination in Multimodal Large Language Model via Hallucination-targeted Direct Preference Optimization
von: Fu, Yuhan, et al.
Veröffentlicht: (2024)
von: Fu, Yuhan, et al.
Veröffentlicht: (2024)
VideoAutoArena: An Automated Arena for Evaluating Large Multimodal Models in Video Analysis through User Simulation
von: Luo, Ziyang, et al.
Veröffentlicht: (2024)
von: Luo, Ziyang, et al.
Veröffentlicht: (2024)
MultiMed: Massively Multimodal and Multitask Medical Understanding
von: Mo, Shentong, et al.
Veröffentlicht: (2024)
von: Mo, Shentong, et al.
Veröffentlicht: (2024)
GalleryGPT: Analyzing Paintings with Large Multimodal Models
von: Bin, Yi, et al.
Veröffentlicht: (2024)
von: Bin, Yi, et al.
Veröffentlicht: (2024)
Mitigating Hallucinations in Large Vision-Language Models with Instruction Contrastive Decoding
von: Wang, Xintong, et al.
Veröffentlicht: (2024)
von: Wang, Xintong, et al.
Veröffentlicht: (2024)
MaVEn: An Effective Multi-granularity Hybrid Visual Encoding Framework for Multimodal Large Language Model
von: Jiang, Chaoya, et al.
Veröffentlicht: (2024)
von: Jiang, Chaoya, et al.
Veröffentlicht: (2024)
VIEScore: Towards Explainable Metrics for Conditional Image Synthesis Evaluation
von: Ku, Max, et al.
Veröffentlicht: (2023)
von: Ku, Max, et al.
Veröffentlicht: (2023)
Distilling Neuro-Symbolic Programs into 3D Multi-modal LLMs
von: Mo, Wentao, et al.
Veröffentlicht: (2026)
von: Mo, Wentao, et al.
Veröffentlicht: (2026)
Every Part Matters: Integrity Verification of Scientific Figures Based on Multimodal Large Language Models
von: Shi, Xiang, et al.
Veröffentlicht: (2024)
von: Shi, Xiang, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Multi-level Matching Network for Multimodal Entity Linking
von: Hu, Zhiwei, et al.
Veröffentlicht: (2024) -
Uni-MoE: Scaling Unified Multimodal LLMs with Mixture of Experts
von: Li, Yunxin, et al.
Veröffentlicht: (2024) -
Reverse Region-to-Entity Annotation for Pixel-Level Visual Entity Linking
von: Xu, Zhengfei, et al.
Veröffentlicht: (2024) -
Advancing Grounded Multimodal Named Entity Recognition via LLM-Based Reformulation and Box-Based Segmentation
von: Li, Jinyuan, et al.
Veröffentlicht: (2024) -
Video Summarization: Towards Entity-Aware Captions
von: Ayyubi, Hammad A., et al.
Veröffentlicht: (2023)