SCO-VIST: Social Interaction Commonsense Knowledge-based Visual Storytelling
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Wang, Eileen, Han, Soyeon Caren, Poon, Josiah |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
VRD-IU: Lessons from Visually Rich Document Intelligence and Understanding
von: Ding, Yihao, et al.
Veröffentlicht: (2025)
von: Ding, Yihao, et al.
Veröffentlicht: (2025)
VIST-GPT: Ushering in the Era of Visual Storytelling with LLMs?
von: Gado, Mohamed, et al.
Veröffentlicht: (2025)
von: Gado, Mohamed, et al.
Veröffentlicht: (2025)
GEM-VPC: A dual Graph-Enhanced Multimodal integration for Video Paragraph Captioning
von: Wang, Eileen, et al.
Veröffentlicht: (2024)
von: Wang, Eileen, et al.
Veröffentlicht: (2024)
TARN-VIST: Topic Aware Reinforcement Network for Visual Storytelling
von: Chen, Weiran, et al.
Veröffentlicht: (2024)
von: Chen, Weiran, et al.
Veröffentlicht: (2024)
Multimodal Commonsense Knowledge Distillation for Visual Question Answering
von: Yang, Shuo, et al.
Veröffentlicht: (2024)
von: Yang, Shuo, et al.
Veröffentlicht: (2024)
Location-Aware Pretraining for Medical Difference Visual Question Answering
von: Musinguzi, Denis, et al.
Veröffentlicht: (2026)
von: Musinguzi, Denis, et al.
Veröffentlicht: (2026)
VCD: A Dataset for Visual Commonsense Discovery in Images
von: Shen, Xiangqing, et al.
Veröffentlicht: (2024)
von: Shen, Xiangqing, et al.
Veröffentlicht: (2024)
Augmented Commonsense Knowledge for Remote Object Grounding
von: Mohammadi, Bahram, et al.
Veröffentlicht: (2024)
von: Mohammadi, Bahram, et al.
Veröffentlicht: (2024)
Activating Visual Context and Commonsense Reasoning through Masked Prediction in VLMs
von: Yu, Jiaao, et al.
Veröffentlicht: (2025)
von: Yu, Jiaao, et al.
Veröffentlicht: (2025)
DIVE: Towards Descriptive and Diverse Visual Commonsense Generation
von: Park, Jun-Hyung, et al.
Veröffentlicht: (2024)
von: Park, Jun-Hyung, et al.
Veröffentlicht: (2024)
A Study of Commonsense Reasoning over Visual Object Properties
von: Kolari, Abhishek, et al.
Veröffentlicht: (2025)
von: Kolari, Abhishek, et al.
Veröffentlicht: (2025)
SPoRC-VIST: A Benchmark for Evaluating Generative Natural Narrative in Vision-Language Models
von: Zeng, Yunlin
Veröffentlicht: (2026)
von: Zeng, Yunlin
Veröffentlicht: (2026)
Towards Commonsense Knowledge based Fuzzy Systems for Supporting Size-Related Fine-Grained Object Detection
von: Zhang, Pu, et al.
Veröffentlicht: (2023)
von: Zhang, Pu, et al.
Veröffentlicht: (2023)
ViCor: Bridging Visual Understanding and Commonsense Reasoning with Large Language Models
von: Zhou, Kaiwen, et al.
Veröffentlicht: (2023)
von: Zhou, Kaiwen, et al.
Veröffentlicht: (2023)
Commonsense-T2I Challenge: Can Text-to-Image Generation Models Understand Commonsense?
von: Fu, Xingyu, et al.
Veröffentlicht: (2024)
von: Fu, Xingyu, et al.
Veröffentlicht: (2024)
Learning to Correction: Explainable Feedback Generation for Visual Commonsense Reasoning Distractor
von: Chen, Jiali, et al.
Veröffentlicht: (2024)
von: Chen, Jiali, et al.
Veröffentlicht: (2024)
Deep Learning based Visually Rich Document Content Understanding: A Survey
von: Ding, Yihao, et al.
Veröffentlicht: (2024)
von: Ding, Yihao, et al.
Veröffentlicht: (2024)
Enhancing Scene Graph Generation with Hierarchical Relationships and Commonsense Knowledge
von: Jiang, Bowen, et al.
Veröffentlicht: (2023)
von: Jiang, Bowen, et al.
Veröffentlicht: (2023)
Text-guided 3D Human Motion Generation with Keyframe-based Parallel Skip Transformer
von: Geng, Zichen, et al.
Veröffentlicht: (2024)
von: Geng, Zichen, et al.
Veröffentlicht: (2024)
ContextualStory: Consistent Visual Storytelling with Spatially-Enhanced and Storyline Context
von: Zheng, Sixiao, et al.
Veröffentlicht: (2024)
von: Zheng, Sixiao, et al.
Veröffentlicht: (2024)
CK-Transformer: Commonsense Knowledge Enhanced Transformers for Referring Expression Comprehension
von: Zhang, Zhi, et al.
Veröffentlicht: (2023)
von: Zhang, Zhi, et al.
Veröffentlicht: (2023)
3MVRD: Multimodal Multi-task Multi-teacher Visually-Rich Form Document Understanding
von: Ding, Yihao, et al.
Veröffentlicht: (2024)
von: Ding, Yihao, et al.
Veröffentlicht: (2024)
A Knowledge Noise Mitigation Framework for Knowledge-based Visual Question Answering
von: Liu, Zhiyue, et al.
Veröffentlicht: (2025)
von: Liu, Zhiyue, et al.
Veröffentlicht: (2025)
Gemini in Reasoning: Unveiling Commonsense in Multimodal Large Language Models
von: Wang, Yuqing, et al.
Veröffentlicht: (2023)
von: Wang, Yuqing, et al.
Veröffentlicht: (2023)
CDH-Bench: A Commonsense-Driven Hallucination Benchmark for Evaluating Visual Fidelity in Vision-Language Models
von: Chen, Kesheng, et al.
Veröffentlicht: (2026)
von: Chen, Kesheng, et al.
Veröffentlicht: (2026)
Knowledge-based Visual Question Answer with Multimodal Processing, Retrieval and Filtering
von: Hong, Yuyang, et al.
Veröffentlicht: (2025)
von: Hong, Yuyang, et al.
Veröffentlicht: (2025)
MMCOMET: A Large-Scale Multimodal Commonsense Knowledge Graph for Contextual Reasoning
von: Wang, Eileen, et al.
Veröffentlicht: (2026)
von: Wang, Eileen, et al.
Veröffentlicht: (2026)
PDF-MVQA: A Dataset for Multimodal Information Retrieval in PDF-based Visual Question Answering
von: Ding, Yihao, et al.
Veröffentlicht: (2024)
von: Ding, Yihao, et al.
Veröffentlicht: (2024)
Commonsense Video Question Answering through Video-Grounded Entailment Tree Reasoning
von: Liu, Huabin, et al.
Veröffentlicht: (2025)
von: Liu, Huabin, et al.
Veröffentlicht: (2025)
IIU: Independent Inference Units for Knowledge-based Visual Question Answering
von: Li, Yili, et al.
Veröffentlicht: (2024)
von: Li, Yili, et al.
Veröffentlicht: (2024)
Map-Free Visual Relocalization Enhanced by Instance Knowledge and Depth Knowledge
von: Xiao, Mingyu, et al.
Veröffentlicht: (2024)
von: Xiao, Mingyu, et al.
Veröffentlicht: (2024)
Is Your Image a Good Storyteller?
von: Song, Xiujie, et al.
Veröffentlicht: (2024)
von: Song, Xiujie, et al.
Veröffentlicht: (2024)
Robust MLLM Unlearning via Visual Knowledge Distillation
von: Wang, Yuhang, et al.
Veröffentlicht: (2025)
von: Wang, Yuhang, et al.
Veröffentlicht: (2025)
MSG-Chart: Multimodal Scene Graph for ChartQA
von: Dai, Yue, et al.
Veröffentlicht: (2024)
von: Dai, Yue, et al.
Veröffentlicht: (2024)
An Efficient System for Automatic Map Storytelling -- A Case Study on Historical Maps
von: Liu, Ziyi, et al.
Veröffentlicht: (2024)
von: Liu, Ziyi, et al.
Veröffentlicht: (2024)
VIRTUE: Visual-Interactive Text-Image Universal Embedder
von: Wang, Wei-Yao, et al.
Veröffentlicht: (2025)
von: Wang, Wei-Yao, et al.
Veröffentlicht: (2025)
VIKSER: Visual Knowledge-Driven Self-Reinforcing Reasoning Framework
von: Wang, Chao, et al.
Veröffentlicht: (2025)
von: Wang, Chao, et al.
Veröffentlicht: (2025)
DrivingDojo Dataset: Advancing Interactive and Knowledge-Enriched Driving World Model
von: Wang, Yuqi, et al.
Veröffentlicht: (2024)
von: Wang, Yuqi, et al.
Veröffentlicht: (2024)
A Knowledge-guided Adversarial Defense for Resisting Malicious Visual Manipulation
von: Zhou, Dawei, et al.
Veröffentlicht: (2025)
von: Zhou, Dawei, et al.
Veröffentlicht: (2025)
VideoPhy: Evaluating Physical Commonsense for Video Generation
von: Bansal, Hritik, et al.
Veröffentlicht: (2024)
von: Bansal, Hritik, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
VRD-IU: Lessons from Visually Rich Document Intelligence and Understanding
von: Ding, Yihao, et al.
Veröffentlicht: (2025) -
VIST-GPT: Ushering in the Era of Visual Storytelling with LLMs?
von: Gado, Mohamed, et al.
Veröffentlicht: (2025) -
GEM-VPC: A dual Graph-Enhanced Multimodal integration for Video Paragraph Captioning
von: Wang, Eileen, et al.
Veröffentlicht: (2024) -
TARN-VIST: Topic Aware Reinforcement Network for Visual Storytelling
von: Chen, Weiran, et al.
Veröffentlicht: (2024) -
Multimodal Commonsense Knowledge Distillation for Visual Question Answering
von: Yang, Shuo, et al.
Veröffentlicht: (2024)