A Dual-way Enhanced Framework from Text Matching Point of View for Multimodal Entity Linking
Fuente:
arXiv
Salvato in:
| Autori principali: | Song, Shezheng, Zhao, Shan, Wang, Chengyu, Yan, Tianwei, Li, Shasha, Mao, Xiaoguang, Wang, Meng |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2023
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
DWE+: Dual-Way Matching Enhanced Framework for Multimodal Entity Linking
di: Song, Shezheng, et al.
Pubblicazione: (2024)
di: Song, Shezheng, et al.
Pubblicazione: (2024)
MOSABench: Multi-Object Sentiment Analysis Benchmark for Evaluating Multimodal Large Language Models Understanding of Complex Image
di: Song, Shezheng, et al.
Pubblicazione: (2024)
di: Song, Shezheng, et al.
Pubblicazione: (2024)
How to Bridge the Gap between Modalities: Survey on Multimodal Large Language Model
di: Song, Shezheng, et al.
Pubblicazione: (2023)
di: Song, Shezheng, et al.
Pubblicazione: (2023)
DcMatch: Unsupervised Multi-Shape Matching with Dual-Level Consistency
di: Ye, Tianwei, et al.
Pubblicazione: (2025)
di: Ye, Tianwei, et al.
Pubblicazione: (2025)
Where Does Vision Meet Language? Understanding and Refining Visual Fusion in MLLMs via Contrastive Attention
di: Song, Shezheng, et al.
Pubblicazione: (2026)
di: Song, Shezheng, et al.
Pubblicazione: (2026)
Seeing Right but Saying Wrong: Inter- and Intra-Layer Refinement in MLLMs without Training
di: Song, Shezheng, et al.
Pubblicazione: (2026)
di: Song, Shezheng, et al.
Pubblicazione: (2026)
EntityCLIP: Entity-Centric Image-Text Matching via Multimodal Attentive Contrastive Learning
di: Wang, Yaxiong, et al.
Pubblicazione: (2024)
di: Wang, Yaxiong, et al.
Pubblicazione: (2024)
DIM: Dynamic Integration of Multimodal Entity Linking with Large Language Model
di: Song, Shezheng, et al.
Pubblicazione: (2024)
di: Song, Shezheng, et al.
Pubblicazione: (2024)
PTA: Enhancing Multimodal Sentiment Analysis through Pipelined Prediction and Translation-based Alignment
di: Song, Shezheng, et al.
Pubblicazione: (2024)
di: Song, Shezheng, et al.
Pubblicazione: (2024)
Multi-level Matching Network for Multimodal Entity Linking
di: Hu, Zhiwei, et al.
Pubblicazione: (2024)
di: Hu, Zhiwei, et al.
Pubblicazione: (2024)
CLII: Visual-Text Inpainting via Cross-Modal Predictive Interaction
di: Zhao, Liang, et al.
Pubblicazione: (2024)
di: Zhao, Liang, et al.
Pubblicazione: (2024)
SGMatch: Semantic-Guided Non-Rigid Shape Matching with Flow Regularization
di: Ye, Tianwei, et al.
Pubblicazione: (2026)
di: Ye, Tianwei, et al.
Pubblicazione: (2026)
Compositional Image-Text Matching and Retrieval by Grounding Entities
di: Vongala, Madhukar Reddy, et al.
Pubblicazione: (2025)
di: Vongala, Madhukar Reddy, et al.
Pubblicazione: (2025)
$M^3EL$: A Multi-task Multi-topic Dataset for Multi-modal Entity Linking
di: Wang, Fang, et al.
Pubblicazione: (2024)
di: Wang, Fang, et al.
Pubblicazione: (2024)
VP-MEL: Visual Prompts Guided Multimodal Entity Linking
di: Mi, Hongze, et al.
Pubblicazione: (2024)
di: Mi, Hongze, et al.
Pubblicazione: (2024)
DPSeg: Dual-Prompt Cost Volume Learning for Open-Vocabulary Semantic Segmentation
di: Zhao, Ziyu, et al.
Pubblicazione: (2025)
di: Zhao, Ziyu, et al.
Pubblicazione: (2025)
Small, Versatile and Mighty: A Range-View Perception Framework
di: Meng, Qiang, et al.
Pubblicazione: (2024)
di: Meng, Qiang, et al.
Pubblicazione: (2024)
MatchAttention: Matching the Relative Positions for High-Resolution Cross-View Matching
di: Yan, Tingman, et al.
Pubblicazione: (2025)
di: Yan, Tingman, et al.
Pubblicazione: (2025)
Enhancing Multimodal Entity Linking with Jaccard Distance-based Conditional Contrastive Learning and Contextual Visual Augmentation
di: Nguyen, Cong-Duy, et al.
Pubblicazione: (2025)
di: Nguyen, Cong-Duy, et al.
Pubblicazione: (2025)
TextMatch: Enhancing Image-Text Consistency Through Multimodal Optimization
di: Luo, Yucong, et al.
Pubblicazione: (2024)
di: Luo, Yucong, et al.
Pubblicazione: (2024)
I2CR: Intra- and Inter-modal Collaborative Reflections for Multimodal Entity Linking
di: Liu, Ziyan, et al.
Pubblicazione: (2025)
di: Liu, Ziyan, et al.
Pubblicazione: (2025)
TextMaster: A Unified Framework for Realistic Text Editing via Glyph-Style Dual-Control
di: Yan, Zhenyu, et al.
Pubblicazione: (2024)
di: Yan, Zhenyu, et al.
Pubblicazione: (2024)
PolarBEVDet: Exploring Polar Representation for Multi-View 3D Object Detection in Bird's-Eye-View
di: Yu, Zichen, et al.
Pubblicazione: (2024)
di: Yu, Zichen, et al.
Pubblicazione: (2024)
PUFM++: Point Cloud Upsampling via Enhanced Flow Matching
di: Liu, Zhi-Song, et al.
Pubblicazione: (2025)
di: Liu, Zhi-Song, et al.
Pubblicazione: (2025)
Instruction-augmented Multimodal Alignment for Image-Text and Element Matching
di: Yue, Xinli, et al.
Pubblicazione: (2025)
di: Yue, Xinli, et al.
Pubblicazione: (2025)
Dual-Level Precision Edges Guided Multi-View Stereo with Accurate Planarization
di: Chen, Kehua, et al.
Pubblicazione: (2024)
di: Chen, Kehua, et al.
Pubblicazione: (2024)
3D Landmark Detection on Human Point Clouds: A Benchmark and A Dual Cascade Point Transformer Framework
di: Zhang, Fan, et al.
Pubblicazione: (2024)
di: Zhang, Fan, et al.
Pubblicazione: (2024)
Masked Point-Entity Contrast for Open-Vocabulary 3D Scene Understanding
di: Wang, Yan, et al.
Pubblicazione: (2025)
di: Wang, Yan, et al.
Pubblicazione: (2025)
Adaptively Enhancing Facial Expression Crucial Regions via Local Non-Local Joint Network
di: Shi, Guanghui, et al.
Pubblicazione: (2022)
di: Shi, Guanghui, et al.
Pubblicazione: (2022)
A Holistically Point-guided Text Framework for Weakly-Supervised Camouflaged Object Detection
di: Mok, Tsui Qin, et al.
Pubblicazione: (2025)
di: Mok, Tsui Qin, et al.
Pubblicazione: (2025)
TextCoT: Zoom In for Enhanced Multimodal Text-Rich Image Understanding
di: Luan, Bozhi, et al.
Pubblicazione: (2024)
di: Luan, Bozhi, et al.
Pubblicazione: (2024)
Multi-View Representation is What You Need for Point-Cloud Pre-Training
di: Yan, Siming, et al.
Pubblicazione: (2023)
di: Yan, Siming, et al.
Pubblicazione: (2023)
Encapsulated Composition of Text-to-Image and Text-to-Video Models for High-Quality Video Synthesis
di: Su, Tongtong, et al.
Pubblicazione: (2025)
di: Su, Tongtong, et al.
Pubblicazione: (2025)
Neural Point-based Volumetric Avatar: Surface-guided Neural Points for Efficient and Photorealistic Volumetric Head Avatar
di: Wang, Cong, et al.
Pubblicazione: (2023)
di: Wang, Cong, et al.
Pubblicazione: (2023)
Multi-level Mixture of Experts for Multimodal Entity Linking
di: Hu, Zhiwei, et al.
Pubblicazione: (2025)
di: Hu, Zhiwei, et al.
Pubblicazione: (2025)
Searching from Area to Point: A Hierarchical Framework for Semantic-Geometric Combined Feature Matching
di: Zhang, Yesheng, et al.
Pubblicazione: (2023)
di: Zhang, Yesheng, et al.
Pubblicazione: (2023)
Aligning Vision to Language: Annotation-Free Multimodal Knowledge Graph Construction for Enhanced LLMs Reasoning
di: Liu, Junming, et al.
Pubblicazione: (2025)
di: Liu, Junming, et al.
Pubblicazione: (2025)
SuPerPM: A Surgical Perception Framework Based on Deep Point Matching Learned from Physical Constrained Simulation Data
di: Lin, Shan, et al.
Pubblicazione: (2023)
di: Lin, Shan, et al.
Pubblicazione: (2023)
A Multimodal Cross-View Model for Predicting Postoperative Neck Pain in Cervical Spondylosis Patients
di: Shan, Jingyang, et al.
Pubblicazione: (2025)
di: Shan, Jingyang, et al.
Pubblicazione: (2025)
DamageArbiter: A CLIP-Enhanced Multimodal Arbitration Framework for Hurricane Damage Assessment from Street-View Imagery
di: Yang, Yifan, et al.
Pubblicazione: (2026)
di: Yang, Yifan, et al.
Pubblicazione: (2026)
Documenti analoghi
-
DWE+: Dual-Way Matching Enhanced Framework for Multimodal Entity Linking
di: Song, Shezheng, et al.
Pubblicazione: (2024) -
MOSABench: Multi-Object Sentiment Analysis Benchmark for Evaluating Multimodal Large Language Models Understanding of Complex Image
di: Song, Shezheng, et al.
Pubblicazione: (2024) -
How to Bridge the Gap between Modalities: Survey on Multimodal Large Language Model
di: Song, Shezheng, et al.
Pubblicazione: (2023) -
DcMatch: Unsupervised Multi-Shape Matching with Dual-Level Consistency
di: Ye, Tianwei, et al.
Pubblicazione: (2025) -
Where Does Vision Meet Language? Understanding and Refining Visual Fusion in MLLMs via Contrastive Attention
di: Song, Shezheng, et al.
Pubblicazione: (2026)