Multi-level Matching Network for Multimodal Entity Linking
Fuente:
arXiv
Saved in:
| Main Authors: | Hu, Zhiwei, Gutiérrez-Basulto, Víctor, Li, Ru, Pan, Jeff Z. |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Multi-level Mixture of Experts for Multimodal Entity Linking
by: Hu, Zhiwei, et al.
Published: (2025)
by: Hu, Zhiwei, et al.
Published: (2025)
Uncovering and Mitigating Transient Blindness in Multimodal Model Editing
by: Han, Xiaoqi, et al.
Published: (2025)
by: Han, Xiaoqi, et al.
Published: (2025)
DWE+: Dual-Way Matching Enhanced Framework for Multimodal Entity Linking
by: Song, Shezheng, et al.
Published: (2024)
by: Song, Shezheng, et al.
Published: (2024)
A Dual-way Enhanced Framework from Text Matching Point of View for Multimodal Entity Linking
by: Song, Shezheng, et al.
Published: (2023)
by: Song, Shezheng, et al.
Published: (2023)
COTET: Cross-view Optimal Transport for Knowledge Graph Entity Typing
by: Hu, Zhiwei, et al.
Published: (2024)
by: Hu, Zhiwei, et al.
Published: (2024)
Knowledge-Aware Neuron Interpretation for Scene Classification
by: Guan, Yong, et al.
Published: (2024)
by: Guan, Yong, et al.
Published: (2024)
Leveraging Intra-modal and Inter-modal Interaction for Multi-Modal Entity Alignment
by: Hu, Zhiwei, et al.
Published: (2024)
by: Hu, Zhiwei, et al.
Published: (2024)
Consistency-Aware Editing for Entity-level Unlearning in Language Models
by: Han, Xiaoqi, et al.
Published: (2025)
by: Han, Xiaoqi, et al.
Published: (2025)
Enhancing Multimodal Entity Linking with Jaccard Distance-based Conditional Contrastive Learning and Contextual Visual Augmentation
by: Nguyen, Cong-Duy, et al.
Published: (2025)
by: Nguyen, Cong-Duy, et al.
Published: (2025)
M$^3$-VQA: A Benchmark for Multimodal, Multi-Entity, Multi-Hop Visual Question Answering
by: Ma, Jiatong, et al.
Published: (2026)
by: Ma, Jiatong, et al.
Published: (2026)
$M^3EL$: A Multi-task Multi-topic Dataset for Multi-modal Entity Linking
by: Wang, Fang, et al.
Published: (2024)
by: Wang, Fang, et al.
Published: (2024)
EntityBench: Towards Entity-Consistent Long-Range Multi-Shot Video Generation
by: He, Ruozhen, et al.
Published: (2026)
by: He, Ruozhen, et al.
Published: (2026)
Multi-Grained Query-Guided Set Prediction Network for Grounded Multimodal Named Entity Recognition
by: Tang, Jielong, et al.
Published: (2024)
by: Tang, Jielong, et al.
Published: (2024)
Reverse Region-to-Entity Annotation for Pixel-Level Visual Entity Linking
by: Xu, Zhengfei, et al.
Published: (2024)
by: Xu, Zhengfei, et al.
Published: (2024)
Hierarchical Long Video Understanding with Audiovisual Entity Cohesion and Agentic Search
by: Yin, Xinlei, et al.
Published: (2026)
by: Yin, Xinlei, et al.
Published: (2026)
RoBus: A Multimodal Dataset for Controllable Road Networks and Building Layouts Generation
by: Li, Tao, et al.
Published: (2024)
by: Li, Tao, et al.
Published: (2024)
Knowledge Completes the Vision: A Multimodal Entity-aware Retrieval-Augmented Generation Framework for News Image Captioning
by: You, Xiaoxing, et al.
Published: (2025)
by: You, Xiaoxing, et al.
Published: (2025)
Distill-then-prune: An Efficient Compression Framework for Real-time Stereo Matching Network on Edge Devices
by: Pan, Baiyu, et al.
Published: (2024)
by: Pan, Baiyu, et al.
Published: (2024)
QAPruner: Quantization-Aware Vision Token Pruning for Multimodal Large Language Models
by: Wang, Xinhao, et al.
Published: (2026)
by: Wang, Xinhao, et al.
Published: (2026)
Multimodal Distribution Matching for Vision-Language Dataset Distillation
by: Jeong, Jongoh, et al.
Published: (2026)
by: Jeong, Jongoh, et al.
Published: (2026)
TiS-TSL: Image-Label Supervised Surgical Video Stereo Matching via Time-Switchable Teacher-Student Learning
by: Wang, Rui, et al.
Published: (2025)
by: Wang, Rui, et al.
Published: (2025)
MT-Video-Bench: A Holistic Video Understanding Benchmark for Evaluating Multimodal LLMs in Multi-Turn Dialogues
by: Pan, Yaning, et al.
Published: (2025)
by: Pan, Yaning, et al.
Published: (2025)
ScaleNet: Scaling up Pretrained Neural Networks with Incremental Parameters
by: Hao, Zhiwei, et al.
Published: (2025)
by: Hao, Zhiwei, et al.
Published: (2025)
MediRound: Multi-Round Entity-Level Reasoning Segmentation in Medical Images
by: Tong, Qinyue, et al.
Published: (2025)
by: Tong, Qinyue, et al.
Published: (2025)
MELT: Improve Composed Image Retrieval via the Modification Frequentation-Rarity Balance Network
by: Qiu, Guozhi, et al.
Published: (2026)
by: Qiu, Guozhi, et al.
Published: (2026)
MCN-CL: Multimodal Cross-Attention Network and Contrastive Learning for Multimodal Emotion Recognition
by: Li, Feng, et al.
Published: (2025)
by: Li, Feng, et al.
Published: (2025)
Principled RL for Flow Matching Emerges from the Chunk-level Policy Optimization
by: Luo, Yifu, et al.
Published: (2025)
by: Luo, Yifu, et al.
Published: (2025)
TextMatch: Enhancing Image-Text Consistency Through Multimodal Optimization
by: Luo, Yucong, et al.
Published: (2024)
by: Luo, Yucong, et al.
Published: (2024)
Few-Shot Precise Event Spotting via Unified Multi-Entity Graph and Distillation
by: Liu, Zhaoyu, et al.
Published: (2025)
by: Liu, Zhaoyu, et al.
Published: (2025)
MicroscopyMatching: Towards a Ready-to-use Framework for Microscopy Image Analysis in Diverse Conditions
by: Hui, Xiaofei, et al.
Published: (2026)
by: Hui, Xiaofei, et al.
Published: (2026)
MogaNet: Multi-order Gated Aggregation Network
by: Li, Siyuan, et al.
Published: (2022)
by: Li, Siyuan, et al.
Published: (2022)
Fixed-length Dense Descriptor for Efficient Fingerprint Matching
by: Pan, Zhiyu, et al.
Published: (2023)
by: Pan, Zhiyu, et al.
Published: (2023)
EM-Vid: Training-Free Entity-Centric Memory for Efficient and Consistent Multi-Shot Video Generation
by: Vandersanden, Jente, et al.
Published: (2026)
by: Vandersanden, Jente, et al.
Published: (2026)
Online,Target-Free LiDAR-Camera Extrinsic Calibration via Cross-Modal Mask Matching
by: Huang, Zhiwei, et al.
Published: (2024)
by: Huang, Zhiwei, et al.
Published: (2024)
Mitigating Hallucinations in Large Vision-Language Models via Entity-Centric Multimodal Preference Optimization
by: Wu, Jiulong, et al.
Published: (2025)
by: Wu, Jiulong, et al.
Published: (2025)
Pedestrian Crossing Intention Prediction Using Multimodal Fusion Network
by: Li, Yuanzhe, et al.
Published: (2025)
by: Li, Yuanzhe, et al.
Published: (2025)
Continuous-Time Distribution Matching for Few-Step Diffusion Distillation
by: Liu, Tao, et al.
Published: (2026)
by: Liu, Tao, et al.
Published: (2026)
Mosaic: Multimodal Jailbreak against Closed-Source VLMs via Multi-View Ensemble Optimization
by: Lan, Yuqin, et al.
Published: (2026)
by: Lan, Yuqin, et al.
Published: (2026)
Multimodal Causal Reasoning Benchmark: Challenging Vision Large Language Models to Discern Causal Links Across Modalities
by: Li, Zhiyuan, et al.
Published: (2024)
by: Li, Zhiyuan, et al.
Published: (2024)
Reviving the Context: Camera Trap Species Classification as Link Prediction on Multimodal Knowledge Graphs
by: Pahuja, Vardaan, et al.
Published: (2023)
by: Pahuja, Vardaan, et al.
Published: (2023)
Similar Items
-
Multi-level Mixture of Experts for Multimodal Entity Linking
by: Hu, Zhiwei, et al.
Published: (2025) -
Uncovering and Mitigating Transient Blindness in Multimodal Model Editing
by: Han, Xiaoqi, et al.
Published: (2025) -
DWE+: Dual-Way Matching Enhanced Framework for Multimodal Entity Linking
by: Song, Shezheng, et al.
Published: (2024) -
A Dual-way Enhanced Framework from Text Matching Point of View for Multimodal Entity Linking
by: Song, Shezheng, et al.
Published: (2023) -
COTET: Cross-view Optimal Transport for Knowledge Graph Entity Typing
by: Hu, Zhiwei, et al.
Published: (2024)