Leveraging Entity Information for Cross-Modality Correlation Learning: The Entity-Guided Multimodal Summarization
Fuente:
arXiv
Salvato in:
| Autori principali: | Zhang, Yanghai, Liu, Ye, Wu, Shiwei, Zhang, Kai, Liu, Xukai, Liu, Qi, Chen, Enhong |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Uncovering Entity Identity Confusion in Multimodal Knowledge Editing
di: Wu, Shu, et al.
Pubblicazione: (2026)
di: Wu, Shu, et al.
Pubblicazione: (2026)
Video Summarization: Towards Entity-Aware Captions
di: Ayyubi, Hammad A., et al.
Pubblicazione: (2023)
di: Ayyubi, Hammad A., et al.
Pubblicazione: (2023)
VP-MEL: Visual Prompts Guided Multimodal Entity Linking
di: Mi, Hongze, et al.
Pubblicazione: (2024)
di: Mi, Hongze, et al.
Pubblicazione: (2024)
E2E-GMNER: End-to-End Generative Grounded Multimodal Named Entity Recognition
di: Zhang, Meng, et al.
Pubblicazione: (2026)
di: Zhang, Meng, et al.
Pubblicazione: (2026)
EAMA : Entity-Aware Multimodal Alignment Based Approach for News Image Captioning
di: Zhang, Junzhe, et al.
Pubblicazione: (2024)
di: Zhang, Junzhe, et al.
Pubblicazione: (2024)
LLMs as Bridges: Reformulating Grounded Multimodal Named Entity Recognition
di: Li, Jinyuan, et al.
Pubblicazione: (2024)
di: Li, Jinyuan, et al.
Pubblicazione: (2024)
Grounding Language Models for Visual Entity Recognition
di: Xiao, Zilin, et al.
Pubblicazione: (2024)
di: Xiao, Zilin, et al.
Pubblicazione: (2024)
MOFI: Learning Image Representations from Noisy Entity Annotated Images
di: Wu, Wentao, et al.
Pubblicazione: (2023)
di: Wu, Wentao, et al.
Pubblicazione: (2023)
A Proposal-Free Query-Guided Network for Grounded Multimodal Named Entity Recognition
di: Li, Hongbing, et al.
Pubblicazione: (2026)
di: Li, Hongbing, et al.
Pubblicazione: (2026)
Light Up the Shadows: Enhance Long-Tailed Entity Grounding with Concept-Guided Vision-Language Models
di: Zhang, Yikai, et al.
Pubblicazione: (2024)
di: Zhang, Yikai, et al.
Pubblicazione: (2024)
OneNet: A Fine-Tuning Free Framework for Few-Shot Entity Linking via Large Language Model Prompting
di: Liu, Xukai, et al.
Pubblicazione: (2024)
di: Liu, Xukai, et al.
Pubblicazione: (2024)
Generalizable Entity Grounding via Assistance of Large Language Model
di: Qi, Lu, et al.
Pubblicazione: (2024)
di: Qi, Lu, et al.
Pubblicazione: (2024)
Mitigating Hallucinations in Large Vision-Language Models via Entity-Centric Multimodal Preference Optimization
di: Wu, Jiulong, et al.
Pubblicazione: (2025)
di: Wu, Jiulong, et al.
Pubblicazione: (2025)
DragEntity: Trajectory Guided Video Generation using Entity and Positional Relationships
di: Wan, Zhang, et al.
Pubblicazione: (2024)
di: Wan, Zhang, et al.
Pubblicazione: (2024)
Fine-Grained Zero-Shot Composed Image Retrieval with Complementary Visual-Semantic Integration
di: Ye, Yongcong, et al.
Pubblicazione: (2026)
di: Ye, Yongcong, et al.
Pubblicazione: (2026)
ECIS-VQG: Generation of Entity-centric Information-seeking Questions from Videos
di: Phukan, Arpan, et al.
Pubblicazione: (2024)
di: Phukan, Arpan, et al.
Pubblicazione: (2024)
EntityCLIP: Entity-Centric Image-Text Matching via Multimodal Attentive Contrastive Learning
di: Wang, Yaxiong, et al.
Pubblicazione: (2024)
di: Wang, Yaxiong, et al.
Pubblicazione: (2024)
Summarization of Multimodal Presentations with Vision-Language Models: Study of the Effect of Modalities and Structure
di: Gigant, Théo, et al.
Pubblicazione: (2025)
di: Gigant, Théo, et al.
Pubblicazione: (2025)
Reverse Region-to-Entity Annotation for Pixel-Level Visual Entity Linking
di: Xu, Zhengfei, et al.
Pubblicazione: (2024)
di: Xu, Zhengfei, et al.
Pubblicazione: (2024)
DWE+: Dual-Way Matching Enhanced Framework for Multimodal Entity Linking
di: Song, Shezheng, et al.
Pubblicazione: (2024)
di: Song, Shezheng, et al.
Pubblicazione: (2024)
Shared and Private Information Learning in Multimodal Sentiment Analysis with Deep Modal Alignment and Self-supervised Multi-Task Learning
di: Lai, Songning, et al.
Pubblicazione: (2023)
di: Lai, Songning, et al.
Pubblicazione: (2023)
Entity-Guided Multi-Task Learning for Infrared and Visible Image Fusion
di: Shao, Wenyu, et al.
Pubblicazione: (2026)
di: Shao, Wenyu, et al.
Pubblicazione: (2026)
Multi-Grained Query-Guided Set Prediction Network for Grounded Multimodal Named Entity Recognition
di: Tang, Jielong, et al.
Pubblicazione: (2024)
di: Tang, Jielong, et al.
Pubblicazione: (2024)
A Cross-Modal Rumor Detection Scheme via Contrastive Learning by Exploring Text and Image internal Correlations
di: Ma, Bin, et al.
Pubblicazione: (2025)
di: Ma, Bin, et al.
Pubblicazione: (2025)
Cross-Modal Retrieval for Motion and Text via DropTriple Loss
di: Yan, Sheng, et al.
Pubblicazione: (2023)
di: Yan, Sheng, et al.
Pubblicazione: (2023)
Advancing Grounded Multimodal Named Entity Recognition via LLM-Based Reformulation and Box-Based Segmentation
di: Li, Jinyuan, et al.
Pubblicazione: (2024)
di: Li, Jinyuan, et al.
Pubblicazione: (2024)
Generating Fine Details of Entity Interactions
di: Gu, Xinyi, et al.
Pubblicazione: (2025)
di: Gu, Xinyi, et al.
Pubblicazione: (2025)
PPTBench: Towards Holistic Evaluation of Large Language Models for PowerPoint Layout and Design Understanding
di: Huang, Zheng, et al.
Pubblicazione: (2025)
di: Huang, Zheng, et al.
Pubblicazione: (2025)
Multimodal Abstractive Summarization of Instructional Videos with Vision-Language Models
di: Nazir, Maham, et al.
Pubblicazione: (2026)
di: Nazir, Maham, et al.
Pubblicazione: (2026)
Multi-level Mixture of Experts for Multimodal Entity Linking
di: Hu, Zhiwei, et al.
Pubblicazione: (2025)
di: Hu, Zhiwei, et al.
Pubblicazione: (2025)
Masked Point-Entity Contrast for Open-Vocabulary 3D Scene Understanding
di: Wang, Yan, et al.
Pubblicazione: (2025)
di: Wang, Yan, et al.
Pubblicazione: (2025)
InfoMerge: Information-aware Token Compression for Efficient Video Large Language Models
di: Liu, Xinxin, et al.
Pubblicazione: (2026)
di: Liu, Xinxin, et al.
Pubblicazione: (2026)
Knowledge Guided Entity-aware Video Captioning and A Basketball Benchmark
di: Xi, Zeyu, et al.
Pubblicazione: (2024)
di: Xi, Zeyu, et al.
Pubblicazione: (2024)
CEIDM: A Controlled Entity and Interaction Diffusion Model for Enhanced Text-to-Image Generation
di: Yang, Mingyue, et al.
Pubblicazione: (2025)
di: Yang, Mingyue, et al.
Pubblicazione: (2025)
ViCA: Efficient Multimodal LLMs with Vision-Only Cross-Attention
di: Liu, Wenjie, et al.
Pubblicazione: (2026)
di: Liu, Wenjie, et al.
Pubblicazione: (2026)
Graph-Driven Multimodal Feature Learning Framework for Apparent Personality Assessment
di: Wang, Kangsheng, et al.
Pubblicazione: (2025)
di: Wang, Kangsheng, et al.
Pubblicazione: (2025)
Reading, Not Thinking: Understanding and Bridging the Modality Gap When Text Becomes Pixels in Multimodal LLMs
di: Sun, Kaiser, et al.
Pubblicazione: (2026)
di: Sun, Kaiser, et al.
Pubblicazione: (2026)
TextTIGER: Text-based Intelligent Generation with Entity Prompt Refinement for Text-to-Image Generation
di: Ozaki, Shintaro, et al.
Pubblicazione: (2025)
di: Ozaki, Shintaro, et al.
Pubblicazione: (2025)
Is Extending Modality The Right Path Towards Omni-Modality?
di: Zhu, Tinghui, et al.
Pubblicazione: (2025)
di: Zhu, Tinghui, et al.
Pubblicazione: (2025)
Unraveling Cross-Modality Knowledge Conflicts in Large Vision-Language Models
di: Zhu, Tinghui, et al.
Pubblicazione: (2024)
di: Zhu, Tinghui, et al.
Pubblicazione: (2024)
Documenti analoghi
-
Uncovering Entity Identity Confusion in Multimodal Knowledge Editing
di: Wu, Shu, et al.
Pubblicazione: (2026) -
Video Summarization: Towards Entity-Aware Captions
di: Ayyubi, Hammad A., et al.
Pubblicazione: (2023) -
VP-MEL: Visual Prompts Guided Multimodal Entity Linking
di: Mi, Hongze, et al.
Pubblicazione: (2024) -
E2E-GMNER: End-to-End Generative Grounded Multimodal Named Entity Recognition
di: Zhang, Meng, et al.
Pubblicazione: (2026) -
EAMA : Entity-Aware Multimodal Alignment Based Approach for News Image Captioning
di: Zhang, Junzhe, et al.
Pubblicazione: (2024)