Robust Symbolic Reasoning for Visual Narratives via Hierarchical and Semantically Normalized Knowledge Graphs
Fuente:
arXiv
Gespeichert in:
| 1. Verfasser: | Chen, Yi-Chun |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Hierarchical Knowledge Graphs for Story Understanding in Visual Narratives
von: Chen, Yi-Chun
Veröffentlicht: (2025)
von: Chen, Yi-Chun
Veröffentlicht: (2025)
Narrative-to-Scene Generation: An LLM-Driven Pipeline for 2D Game Environments
von: Chen, Yi-Chun, et al.
Veröffentlicht: (2025)
von: Chen, Yi-Chun, et al.
Veröffentlicht: (2025)
Contrast then Memorize: Semantic Neighbor Retrieval-Enhanced Inductive Multimodal Knowledge Graph Completion
von: Zhao, Yu, et al.
Veröffentlicht: (2024)
von: Zhao, Yu, et al.
Veröffentlicht: (2024)
MMPKUBase: A Comprehensive and High-quality Chinese Multi-modal Knowledge Graph
von: Yi, Xuan, et al.
Veröffentlicht: (2024)
von: Yi, Xuan, et al.
Veröffentlicht: (2024)
SarcasmMiner: A Dual-Track Post-Training Framework for Robust Audio-Visual Sarcasm Reasoning
von: Li, Zhu, et al.
Veröffentlicht: (2026)
von: Li, Zhu, et al.
Veröffentlicht: (2026)
Dual Knowledge-Enhanced Two-Stage Reasoner for Multimodal Dialog Systems
von: Chen, Xiaolin, et al.
Veröffentlicht: (2025)
von: Chen, Xiaolin, et al.
Veröffentlicht: (2025)
LaF-GRPO: In-Situ Navigation Instruction Generation for the Visually Impaired via GRPO with LLM-as-Follower Reward
von: Zhao, Yi, et al.
Veröffentlicht: (2025)
von: Zhao, Yi, et al.
Veröffentlicht: (2025)
KG-RAG: Enhancing GUI Agent Decision-Making via Knowledge Graph-Driven Retrieval-Augmented Generation
von: Guan, Ziyi, et al.
Veröffentlicht: (2025)
von: Guan, Ziyi, et al.
Veröffentlicht: (2025)
ShieldVLM: Safeguarding the Multimodal Implicit Toxicity via Deliberative Reasoning with LVLMs
von: Cui, Shiyao, et al.
Veröffentlicht: (2025)
von: Cui, Shiyao, et al.
Veröffentlicht: (2025)
An Evaluation of Interleaved Instruction Tuning on Semantic Reasoning Performance in an Audio MLLM
von: Liu, Jiawei, et al.
Veröffentlicht: (2025)
von: Liu, Jiawei, et al.
Veröffentlicht: (2025)
Hierarchical Vision-Language Reasoning for Multimodal Multiple-Choice Question Answering
von: Zhou, Ao, et al.
Veröffentlicht: (2025)
von: Zhou, Ao, et al.
Veröffentlicht: (2025)
GameTileNet: A Semantic Dataset for Low-Resolution Game Art in Procedural Content Generation
von: Chen, Yi-Chun, et al.
Veröffentlicht: (2025)
von: Chen, Yi-Chun, et al.
Veröffentlicht: (2025)
Integrating Fine-Grained Audio-Visual Evidence for Robust Multimodal Emotion Reasoning
von: Zhao, Zhixian, et al.
Veröffentlicht: (2026)
von: Zhao, Zhixian, et al.
Veröffentlicht: (2026)
Modeling Layout Reading Order as Ordering Relations for Visually-rich Document Understanding
von: Zhang, Chong, et al.
Veröffentlicht: (2024)
von: Zhang, Chong, et al.
Veröffentlicht: (2024)
Dental-TriageBench: Benchmarking Multimodal Reasoning for Hierarchical Dental Triage
von: He, Ziyi, et al.
Veröffentlicht: (2026)
von: He, Ziyi, et al.
Veröffentlicht: (2026)
MMESGBench: Pioneering Multimodal Understanding and Complex Reasoning Benchmark for ESG Tasks
von: Zhang, Lei, et al.
Veröffentlicht: (2025)
von: Zhang, Lei, et al.
Veröffentlicht: (2025)
Multimodal Sentiment Analysis Based on Causal Reasoning
von: Chen, Fuhai, et al.
Veröffentlicht: (2024)
von: Chen, Fuhai, et al.
Veröffentlicht: (2024)
Hierarchical Textual Knowledge for Enhanced Image Clustering
von: Zhong, Yijie, et al.
Veröffentlicht: (2026)
von: Zhong, Yijie, et al.
Veröffentlicht: (2026)
CMATH: Cross-Modality Augmented Transformer with Hierarchical Variational Distillation for Multimodal Emotion Recognition in Conversation
von: Zhu, Xiaofei, et al.
Veröffentlicht: (2024)
von: Zhu, Xiaofei, et al.
Veröffentlicht: (2024)
LLM-Guided Semantic Relational Reasoning for Multimodal Intent Recognition
von: Zhou, Qianrui, et al.
Veröffentlicht: (2025)
von: Zhou, Qianrui, et al.
Veröffentlicht: (2025)
Knowledge Acquisition Disentanglement for Knowledge-based Visual Question Answering with Large Language Models
von: An, Wenbin, et al.
Veröffentlicht: (2024)
von: An, Wenbin, et al.
Veröffentlicht: (2024)
Hierarchical Aligned Multimodal Learning for NER on Tweet Posts
von: Liu, Peipei, et al.
Veröffentlicht: (2023)
von: Liu, Peipei, et al.
Veröffentlicht: (2023)
History-Guided Iterative Visual Reasoning with Self-Correction
von: Yang, Xinglong, et al.
Veröffentlicht: (2026)
von: Yang, Xinglong, et al.
Veröffentlicht: (2026)
Mixture-of-Prompt-Experts for Multi-modal Semantic Understanding
von: Wu, Zichen, et al.
Veröffentlicht: (2024)
von: Wu, Zichen, et al.
Veröffentlicht: (2024)
ChartEditor: A Reinforcement Learning Framework for Robust Chart Editing
von: Chen, Liangyu, et al.
Veröffentlicht: (2025)
von: Chen, Liangyu, et al.
Veröffentlicht: (2025)
MIND Your Reasoning: A Meta-Cognitive Intuitive-Reflective Network for Dual-Reasoning in Multimodal Stance Detection
von: Wang, Bingbing, et al.
Veröffentlicht: (2025)
von: Wang, Bingbing, et al.
Veröffentlicht: (2025)
The Interspeech 2026 Audio Reasoning Challenge: Evaluating Reasoning Process Quality for Audio Reasoning Models and Agents
von: Ma, Ziyang, et al.
Veröffentlicht: (2026)
von: Ma, Ziyang, et al.
Veröffentlicht: (2026)
ReAG: Reasoning-Augmented Generation for Knowledge-based Visual Question Answering
von: Compagnoni, Alberto, et al.
Veröffentlicht: (2025)
von: Compagnoni, Alberto, et al.
Veröffentlicht: (2025)
GeoGuess: Multimodal Reasoning based on Hierarchy of Visual Information in Street View
von: Cheng, Fenghua, et al.
Veröffentlicht: (2025)
von: Cheng, Fenghua, et al.
Veröffentlicht: (2025)
Multi-MLLM Knowledge Distillation for Out-of-Context News Detection
von: Gu, Yimeng, et al.
Veröffentlicht: (2025)
von: Gu, Yimeng, et al.
Veröffentlicht: (2025)
Sentiment-enhanced Graph-based Sarcasm Explanation in Dialogue
von: Ouyang, Kun, et al.
Veröffentlicht: (2024)
von: Ouyang, Kun, et al.
Veröffentlicht: (2024)
Distilling Implicit Multimodal Knowledge into Large Language Models for Zero-Resource Dialogue Generation
von: Zhang, Bo, et al.
Veröffentlicht: (2024)
von: Zhang, Bo, et al.
Veröffentlicht: (2024)
MuPHI: Learning Implicit Multimodal Harm Reasoning via Semantically Grounded Reward Optimization
von: Saha, Anisha, et al.
Veröffentlicht: (2026)
von: Saha, Anisha, et al.
Veröffentlicht: (2026)
Towards Pretraining Robust ASR Foundation Model with Acoustic-Aware Data Augmentation
von: Liu, Dancheng, et al.
Veröffentlicht: (2025)
von: Liu, Dancheng, et al.
Veröffentlicht: (2025)
Continual Multimodal Knowledge Graph Construction
von: Chen, Xiang, et al.
Veröffentlicht: (2023)
von: Chen, Xiang, et al.
Veröffentlicht: (2023)
TARQ: Tail-Aware Reconstruction Quantization for Rare-Word Robust Automatic Speech Recognition
von: Wang, Xinyu, et al.
Veröffentlicht: (2026)
von: Wang, Xinyu, et al.
Veröffentlicht: (2026)
A Benchmark and Robustness Study of In-Context-Learning with Large Language Models in Music Entity Detection
von: Hachmeier, Simon, et al.
Veröffentlicht: (2024)
von: Hachmeier, Simon, et al.
Veröffentlicht: (2024)
Unleashing the Power of Imbalanced Modality Information for Multi-modal Knowledge Graph Completion
von: Zhang, Yichi, et al.
Veröffentlicht: (2024)
von: Zhang, Yichi, et al.
Veröffentlicht: (2024)
Tailored Teaching with Balanced Difficulty: Elevating Reasoning in Multimodal Chain-of-Thought via Prompt Curriculum
von: Yang, Xinglong, et al.
Veröffentlicht: (2025)
von: Yang, Xinglong, et al.
Veröffentlicht: (2025)
MolCA: Molecular Graph-Language Modeling with Cross-Modal Projector and Uni-Modal Adapter
von: Liu, Zhiyuan, et al.
Veröffentlicht: (2023)
von: Liu, Zhiyuan, et al.
Veröffentlicht: (2023)
Ähnliche Einträge
-
Hierarchical Knowledge Graphs for Story Understanding in Visual Narratives
von: Chen, Yi-Chun
Veröffentlicht: (2025) -
Narrative-to-Scene Generation: An LLM-Driven Pipeline for 2D Game Environments
von: Chen, Yi-Chun, et al.
Veröffentlicht: (2025) -
Contrast then Memorize: Semantic Neighbor Retrieval-Enhanced Inductive Multimodal Knowledge Graph Completion
von: Zhao, Yu, et al.
Veröffentlicht: (2024) -
MMPKUBase: A Comprehensive and High-quality Chinese Multi-modal Knowledge Graph
von: Yi, Xuan, et al.
Veröffentlicht: (2024) -
SarcasmMiner: A Dual-Track Post-Training Framework for Robust Audio-Visual Sarcasm Reasoning
von: Li, Zhu, et al.
Veröffentlicht: (2026)