Abstractive Visual Understanding of Multi-modal Structured Knowledge: A New Perspective for MLLM Evaluation
Fuente:
arXiv
Salvato in:
| Autori principali: | Zhang, Yichi, Chen, Zhuo, Guo, Lingbing, Xu, Yajing, Zhang, Min, Zhang, Wen, Chen, Huajun |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Structured and Abstractive Reasoning on Multi-modal Relational Knowledge Images
di: Zhang, Yichi, et al.
Pubblicazione: (2025)
di: Zhang, Yichi, et al.
Pubblicazione: (2025)
NativE: Multi-modal Knowledge Graph Completion in the Wild
di: Zhang, Yichi, et al.
Pubblicazione: (2024)
di: Zhang, Yichi, et al.
Pubblicazione: (2024)
Every Little Helps: Building Knowledge Graph Foundation Model with Fine-grained Transferable Multi-modal Tokens
di: Zhang, Yichi, et al.
Pubblicazione: (2026)
di: Zhang, Yichi, et al.
Pubblicazione: (2026)
Making Large Language Models Perform Better in Knowledge Graph Completion
di: Zhang, Yichi, et al.
Pubblicazione: (2023)
di: Zhang, Yichi, et al.
Pubblicazione: (2023)
Noise-powered Multi-modal Knowledge Graph Representation Framework
di: Chen, Zhuo, et al.
Pubblicazione: (2024)
di: Chen, Zhuo, et al.
Pubblicazione: (2024)
Have We Designed Generalizable Structural Knowledge Promptings? Systematic Evaluation and Rethinking
di: Zhang, Yichi, et al.
Pubblicazione: (2024)
di: Zhang, Yichi, et al.
Pubblicazione: (2024)
Multiple Heads are Better than One: Mixture of Modality Knowledge Experts for Entity Representation Learning
di: Zhang, Yichi, et al.
Pubblicazione: (2024)
di: Zhang, Yichi, et al.
Pubblicazione: (2024)
Unleashing the Power of Imbalanced Modality Information for Multi-modal Knowledge Graph Completion
di: Zhang, Yichi, et al.
Pubblicazione: (2024)
di: Zhang, Yichi, et al.
Pubblicazione: (2024)
K-ON: Stacking Knowledge On the Head Layer of Large Language Model
di: Guo, Lingbing, et al.
Pubblicazione: (2025)
di: Guo, Lingbing, et al.
Pubblicazione: (2025)
Multi-domain Knowledge Graph Collaborative Pre-training and Prompt Tuning for Diverse Downstream Tasks
di: Zhang, Yichi, et al.
Pubblicazione: (2024)
di: Zhang, Yichi, et al.
Pubblicazione: (2024)
Revisit and Outstrip Entity Alignment: A Perspective of Generative Models
di: Guo, Lingbing, et al.
Pubblicazione: (2023)
di: Guo, Lingbing, et al.
Pubblicazione: (2023)
Distributed Representations of Entities in Open-World Knowledge Graphs
di: Guo, Lingbing, et al.
Pubblicazione: (2020)
di: Guo, Lingbing, et al.
Pubblicazione: (2020)
Knowledgeable Preference Alignment for LLMs in Domain-specific Question Answering
di: Zhang, Yichi, et al.
Pubblicazione: (2023)
di: Zhang, Yichi, et al.
Pubblicazione: (2023)
Tokenization, Fusion, and Augmentation: Towards Fine-grained Multi-modal Entity Representation
di: Zhang, Yichi, et al.
Pubblicazione: (2024)
di: Zhang, Yichi, et al.
Pubblicazione: (2024)
MLLM can see? Dynamic Correction Decoding for Hallucination Mitigation
di: Wang, Chenxi, et al.
Pubblicazione: (2024)
di: Wang, Chenxi, et al.
Pubblicazione: (2024)
SciVerse: Unveiling the Knowledge Comprehension and Visual Reasoning of LMMs on Multi-modal Scientific Problems
di: Guo, Ziyu, et al.
Pubblicazione: (2025)
di: Guo, Ziyu, et al.
Pubblicazione: (2025)
Domain-Agnostic Molecular Generation with Chemical Feedback
di: Fang, Yin, et al.
Pubblicazione: (2023)
di: Fang, Yin, et al.
Pubblicazione: (2023)
Multi-modal Semantic Understanding with Contrastive Cross-modal Feature Alignment
di: Zhang, Ming, et al.
Pubblicazione: (2024)
di: Zhang, Ming, et al.
Pubblicazione: (2024)
Multi-SpatialMLLM: Multi-Frame Spatial Understanding with Multi-Modal Large Language Models
di: Xu, Runsen, et al.
Pubblicazione: (2025)
di: Xu, Runsen, et al.
Pubblicazione: (2025)
MKGL: Mastery of a Three-Word Language
di: Guo, Lingbing, et al.
Pubblicazione: (2024)
di: Guo, Lingbing, et al.
Pubblicazione: (2024)
How Do LLMs Acquire New Knowledge? A Knowledge Circuits Perspective on Continual Pre-Training
di: Ou, Yixin, et al.
Pubblicazione: (2025)
di: Ou, Yixin, et al.
Pubblicazione: (2025)
Natural Language Understanding and Inference with MLLM in Visual Question Answering: A Survey
di: Kuang, Jiayi, et al.
Pubblicazione: (2024)
di: Kuang, Jiayi, et al.
Pubblicazione: (2024)
Knowledge Graphs Meet Multi-Modal Learning: A Comprehensive Survey
di: Chen, Zhuo, et al.
Pubblicazione: (2024)
di: Chen, Zhuo, et al.
Pubblicazione: (2024)
MathVerse: Does Your Multi-modal LLM Truly See the Diagrams in Visual Math Problems?
di: Zhang, Renrui, et al.
Pubblicazione: (2024)
di: Zhang, Renrui, et al.
Pubblicazione: (2024)
Customizing Visual-Language Foundation Models for Multi-modal Anomaly Detection and Reasoning
di: Xu, Xiaohao, et al.
Pubblicazione: (2024)
di: Xu, Xiaohao, et al.
Pubblicazione: (2024)
SKA-Bench: A Fine-Grained Benchmark for Evaluating Structured Knowledge Understanding of LLMs
di: Liu, Zhiqiang, et al.
Pubblicazione: (2025)
di: Liu, Zhiqiang, et al.
Pubblicazione: (2025)
Cognitive Visual-Language Mapper: Advancing Multimodal Comprehension with Enhanced Visual Knowledge Alignment
di: Li, Yunxin, et al.
Pubblicazione: (2024)
di: Li, Yunxin, et al.
Pubblicazione: (2024)
Visual Language Tracking with Multi-modal Interaction: A Robust Benchmark
di: Li, Xuchen, et al.
Pubblicazione: (2024)
di: Li, Xuchen, et al.
Pubblicazione: (2024)
Large Knowledge Model: Perspectives and Challenges
di: Chen, Huajun
Pubblicazione: (2023)
di: Chen, Huajun
Pubblicazione: (2023)
Unveil: Unified Visual-Textual Integration and Distillation for Multi-modal Document Retrieval
di: Sun, Hao, et al.
Pubblicazione: (2026)
di: Sun, Hao, et al.
Pubblicazione: (2026)
Effective Training Data Synthesis for Improving MLLM Chart Understanding
di: Yang, Yuwei, et al.
Pubblicazione: (2025)
di: Yang, Yuwei, et al.
Pubblicazione: (2025)
Cantor: Inspiring Multimodal Chain-of-Thought of MLLM
di: Gao, Timin, et al.
Pubblicazione: (2024)
di: Gao, Timin, et al.
Pubblicazione: (2024)
Multimodal Abstractive Summarization of Instructional Videos with Vision-Language Models
di: Nazir, Maham, et al.
Pubblicazione: (2026)
di: Nazir, Maham, et al.
Pubblicazione: (2026)
Understanding LLM Reasoning for Abstractive Summarization
di: Yuan, Haohan, et al.
Pubblicazione: (2025)
di: Yuan, Haohan, et al.
Pubblicazione: (2025)
Visual Puns from Idioms: An Iterative LLM-T2IM-MLLM Framework
di: Xiao, Kelaiti, et al.
Pubblicazione: (2025)
di: Xiao, Kelaiti, et al.
Pubblicazione: (2025)
MLLM-Protector: Ensuring MLLM's Safety without Hurting Performance
di: Pi, Renjie, et al.
Pubblicazione: (2024)
di: Pi, Renjie, et al.
Pubblicazione: (2024)
Spatial Knowledge Graph-Guided Multimodal Synthesis
di: Xue, Yida, et al.
Pubblicazione: (2025)
di: Xue, Yida, et al.
Pubblicazione: (2025)
Fair Abstractive Summarization of Diverse Perspectives
di: Zhang, Yusen, et al.
Pubblicazione: (2023)
di: Zhang, Yusen, et al.
Pubblicazione: (2023)
Visual Thoughts: A Unified Perspective of Understanding Multimodal Chain-of-Thought
di: Cheng, Zihui, et al.
Pubblicazione: (2025)
di: Cheng, Zihui, et al.
Pubblicazione: (2025)
DTVLT: A Multi-modal Diverse Text Benchmark for Visual Language Tracking Based on LLM
di: Li, Xuchen, et al.
Pubblicazione: (2024)
di: Li, Xuchen, et al.
Pubblicazione: (2024)
Documenti analoghi
-
Structured and Abstractive Reasoning on Multi-modal Relational Knowledge Images
di: Zhang, Yichi, et al.
Pubblicazione: (2025) -
NativE: Multi-modal Knowledge Graph Completion in the Wild
di: Zhang, Yichi, et al.
Pubblicazione: (2024) -
Every Little Helps: Building Knowledge Graph Foundation Model with Fine-grained Transferable Multi-modal Tokens
di: Zhang, Yichi, et al.
Pubblicazione: (2026) -
Making Large Language Models Perform Better in Knowledge Graph Completion
di: Zhang, Yichi, et al.
Pubblicazione: (2023) -
Noise-powered Multi-modal Knowledge Graph Representation Framework
di: Chen, Zhuo, et al.
Pubblicazione: (2024)