Not (yet) the whole story: Evaluating Visual Storytelling Requires More than Measuring Coherence, Grounding, and Repetition
Fuente:
arXiv
Saved in:
| Main Authors: | Surikuchi, Aditya K, Fernández, Raquel, Pezzelle, Sandro |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Natural Language Generation from Visual Events: State-of-the-Art and Key Open Questions
by: Surikuchi, Aditya K, et al.
Published: (2025)
by: Surikuchi, Aditya K, et al.
Published: (2025)
Where is the multimodal goal post? On the Ability of Foundation Models to Recognize Contextually Important Moments
by: Surikuchi, Aditya K, et al.
Published: (2026)
by: Surikuchi, Aditya K, et al.
Published: (2026)
Describing Images $\textit{Fast and Slow}$: Quantifying and Predicting the Variation in Human Signals during Visuo-Linguistic Processes
by: Takmaz, Ece, et al.
Published: (2024)
by: Takmaz, Ece, et al.
Published: (2024)
The BLA Benchmark: Investigating Basic Language Abilities of Pre-Trained Multimodal Models
by: Chen, Xinyi, et al.
Published: (2023)
by: Chen, Xinyi, et al.
Published: (2023)
More than a Moment: Towards Coherent Sequences of Audio Descriptions
by: Khandelwal, Eshika, et al.
Published: (2025)
by: Khandelwal, Eshika, et al.
Published: (2025)
From Image Captioning to Visual Storytelling
by: Passadakis, Admitos, et al.
Published: (2025)
by: Passadakis, Admitos, et al.
Published: (2025)
Context-aware Visual Storytelling with Visual Prefix Tuning and Contrastive Learning
by: Song, Yingjin, et al.
Published: (2024)
by: Song, Yingjin, et al.
Published: (2024)
ReCap: Lightweight Referential Grounding for Coherent Story Visualization
by: Arora, Aditya, et al.
Published: (2026)
by: Arora, Aditya, et al.
Published: (2026)
Generating Storytelling Images with Rich Chains-of-Reasoning
by: Song, Xiujie, et al.
Published: (2025)
by: Song, Xiujie, et al.
Published: (2025)
Evaluation of Audio-Visual Alignments in Visually Grounded Speech Models
by: Khorrami, Khazar, et al.
Published: (2021)
by: Khorrami, Khazar, et al.
Published: (2021)
Beyond the Textual: Generating Coherent Visual Options for MCQs
by: Wang, Wanqiang, et al.
Published: (2025)
by: Wang, Wanqiang, et al.
Published: (2025)
Grounding Language Models for Visual Entity Recognition
by: Xiao, Zilin, et al.
Published: (2024)
by: Xiao, Zilin, et al.
Published: (2024)
Opinion: Learning Intuitive Physics May Require More than Visual Data
by: Su, Ellen, et al.
Published: (2025)
by: Su, Ellen, et al.
Published: (2025)
Weakly-Supervised 3D Visual Grounding based on Visual Language Alignment
by: Xu, Xiaoxu, et al.
Published: (2023)
by: Xu, Xiaoxu, et al.
Published: (2023)
Who Evaluates the Evaluations? Objectively Scoring Text-to-Image Prompt Coherence Metrics with T2IScoreScore (TS2)
by: Saxon, Michael, et al.
Published: (2024)
by: Saxon, Michael, et al.
Published: (2024)
Less is More: High-value Data Selection for Visual Instruction Tuning
by: Liu, Zikang, et al.
Published: (2024)
by: Liu, Zikang, et al.
Published: (2024)
Traceable Evidence Enhanced Visual Grounded Reasoning: Evaluation and Methodology
by: Wang, Haochen, et al.
Published: (2025)
by: Wang, Haochen, et al.
Published: (2025)
VIST-GPT: Ushering in the Era of Visual Storytelling with LLMs?
by: Gado, Mohamed, et al.
Published: (2025)
by: Gado, Mohamed, et al.
Published: (2025)
Less is More: A Simple yet Effective Token Reduction Method for Efficient Multi-modal LLMs
by: Song, Dingjie, et al.
Published: (2024)
by: Song, Dingjie, et al.
Published: (2024)
Learning GUI Grounding with Spatial Reasoning from Visual Feedback
by: Zhao, Yu, et al.
Published: (2025)
by: Zhao, Yu, et al.
Published: (2025)
Referencing Where to Focus: Improving VisualGrounding with Referential Query
by: Wang, Yabing, et al.
Published: (2024)
by: Wang, Yabing, et al.
Published: (2024)
Movie Facts and Fibs (MF$^2$): A Benchmark for Long Movie Understanding
by: Zaranis, Emmanouil, et al.
Published: (2025)
by: Zaranis, Emmanouil, et al.
Published: (2025)
Less Is More? Selective Visual Attention to High-Importance Regions for Multimodal Radiology Summarization
by: Naznin, Mst. Fahmida Sultana, et al.
Published: (2026)
by: Naznin, Mst. Fahmida Sultana, et al.
Published: (2026)
Visually Grounded Speech Models for Low-resource Languages and Cognitive Modelling
by: Nortje, Leanne
Published: (2024)
by: Nortje, Leanne
Published: (2024)
Vision-Language Modeling in PET/CT for Visual Grounding of Positive Findings
by: Huemann, Zachary, et al.
Published: (2025)
by: Huemann, Zachary, et al.
Published: (2025)
Adaptive Markup Language Generation for Contextually-Grounded Visual Document Understanding
by: Xiao, Han, et al.
Published: (2025)
by: Xiao, Han, et al.
Published: (2025)
Discovering Meaningful Units with Visually Grounded Semantics from Image Captions
by: Behjati, Melika, et al.
Published: (2025)
by: Behjati, Melika, et al.
Published: (2025)
v1: Learning to Point Visual Tokens for Multimodal Grounded Reasoning
by: Chung, Jiwan, et al.
Published: (2025)
by: Chung, Jiwan, et al.
Published: (2025)
VividMed: Vision Language Model with Versatile Visual Grounding for Medicine
by: Luo, Lingxiao, et al.
Published: (2024)
by: Luo, Lingxiao, et al.
Published: (2024)
VL-GLUE: A Suite of Fundamental yet Challenging Visuo-Linguistic Reasoning Tasks
by: Sampat, Shailaja Keyur, et al.
Published: (2024)
by: Sampat, Shailaja Keyur, et al.
Published: (2024)
Is Your Image a Good Storyteller?
by: Song, Xiujie, et al.
Published: (2024)
by: Song, Xiujie, et al.
Published: (2024)
Decomposed On-Policy Distillation for Vision-Language Reasoning: Steering Gradients for Visual Grounding
by: Yoon, Hee Suk, et al.
Published: (2026)
by: Yoon, Hee Suk, et al.
Published: (2026)
Grounding Chest X-Ray Visual Question Answering with Generated Radiology Reports
by: Serra, Francesco Dalla, et al.
Published: (2025)
by: Serra, Francesco Dalla, et al.
Published: (2025)
Evaluation and Enhancement of Semantic Grounding in Large Vision-Language Models
by: Lu, Jiaying, et al.
Published: (2023)
by: Lu, Jiaying, et al.
Published: (2023)
VGR: Visual Grounded Reasoning
by: Wang, Jiacong, et al.
Published: (2025)
by: Wang, Jiacong, et al.
Published: (2025)
Learning To See But Forgetting To Follow: Visual Instruction Tuning Makes LLMs More Prone To Jailbreak Attacks
by: Pantazopoulos, Georgios, et al.
Published: (2024)
by: Pantazopoulos, Georgios, et al.
Published: (2024)
Grounded Visual Factualization: Factual Anchor-Based Finetuning for Enhancing MLLM Factual Consistency
by: Morbiato, Filippo, et al.
Published: (2025)
by: Morbiato, Filippo, et al.
Published: (2025)
MOTOR: Multimodal Optimal Transport via Grounded Retrieval in Medical Visual Question Answering
by: Shaaban, Mai A., et al.
Published: (2025)
by: Shaaban, Mai A., et al.
Published: (2025)
Four Ways to Improve Verbo-visual Fusion for Dense 3D Visual Grounding
by: Unal, Ozan, et al.
Published: (2023)
by: Unal, Ozan, et al.
Published: (2023)
Seeing is Believing: Rich-Context Hallucination Detection for MLLMs via Backward Visual Grounding
by: Guo, Pinxue, et al.
Published: (2025)
by: Guo, Pinxue, et al.
Published: (2025)
Similar Items
-
Natural Language Generation from Visual Events: State-of-the-Art and Key Open Questions
by: Surikuchi, Aditya K, et al.
Published: (2025) -
Where is the multimodal goal post? On the Ability of Foundation Models to Recognize Contextually Important Moments
by: Surikuchi, Aditya K, et al.
Published: (2026) -
Describing Images $\textit{Fast and Slow}$: Quantifying and Predicting the Variation in Human Signals during Visuo-Linguistic Processes
by: Takmaz, Ece, et al.
Published: (2024) -
The BLA Benchmark: Investigating Basic Language Abilities of Pre-Trained Multimodal Models
by: Chen, Xinyi, et al.
Published: (2023) -
More than a Moment: Towards Coherent Sequences of Audio Descriptions
by: Khandelwal, Eshika, et al.
Published: (2025)