From Image Captioning to Visual Storytelling
Fuente:
arXiv
Saved in:
| Main Authors: | Passadakis, Admitos, Song, Yingjin, Gatt, Albert |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Context-aware Visual Storytelling with Visual Prefix Tuning and Contrastive Learning
by: Song, Yingjin, et al.
Published: (2024)
by: Song, Yingjin, et al.
Published: (2024)
Burn After Reading: Do Multimodal Large Language Models Truly Capture Order of Events in Image Sequences?
by: Song, Yingjin, et al.
Published: (2025)
by: Song, Yingjin, et al.
Published: (2025)
Common Objects Out of Context (COOCo): Investigating Multimodal Context and Semantic Scene Violations in Referential Communication
by: Merlo, Filippo, et al.
Published: (2025)
by: Merlo, Filippo, et al.
Published: (2025)
Generating Storytelling Images with Rich Chains-of-Reasoning
by: Song, Xiujie, et al.
Published: (2025)
by: Song, Xiujie, et al.
Published: (2025)
VIXEN: Visual Text Comparison Network for Image Difference Captioning
by: Black, Alexander, et al.
Published: (2024)
by: Black, Alexander, et al.
Published: (2024)
Discovering Meaningful Units with Visually Grounded Semantics from Image Captions
by: Behjati, Melika, et al.
Published: (2025)
by: Behjati, Melika, et al.
Published: (2025)
Is Your Image a Good Storyteller?
by: Song, Xiujie, et al.
Published: (2024)
by: Song, Xiujie, et al.
Published: (2024)
Image Captioning via Compact Bidirectional Architecture
by: Song, Zijie, et al.
Published: (2022)
by: Song, Zijie, et al.
Published: (2022)
See or Guess: Counterfactually Regularized Image Captioning
by: Cao, Qian, et al.
Published: (2024)
by: Cao, Qian, et al.
Published: (2024)
Unifying Vision-Language Latents for Zero-label Image Caption Enhancement
by: Byun, Sanghyun, et al.
Published: (2025)
by: Byun, Sanghyun, et al.
Published: (2025)
Distinctive Image Captioning: Leveraging Ground Truth Captions in CLIP Guided Reinforcement Learning
by: Chaffin, Antoine, et al.
Published: (2024)
by: Chaffin, Antoine, et al.
Published: (2024)
CapArena: Benchmarking and Analyzing Detailed Image Captioning in the LLM Era
by: Cheng, Kanzhi, et al.
Published: (2025)
by: Cheng, Kanzhi, et al.
Published: (2025)
MUNIChus: Multilingual News Image Captioning Benchmark
by: Chen, Yuji, et al.
Published: (2026)
by: Chen, Yuji, et al.
Published: (2026)
The Devil is in the EOS: Sequence Training for Detailed Image Captioning
by: Mohamed, Abdelrahman, et al.
Published: (2025)
by: Mohamed, Abdelrahman, et al.
Published: (2025)
Understanding Retrieval Robustness for Retrieval-Augmented Image Captioning
by: Li, Wenyan, et al.
Published: (2024)
by: Li, Wenyan, et al.
Published: (2024)
Temporal Image Caption Retrieval Competition -- Description and Results
by: Pokrywka, Jakub, et al.
Published: (2024)
by: Pokrywka, Jakub, et al.
Published: (2024)
Correlates of Image Memorability in Vision Encoders: Activations, Attention Entropy, Patch Uniformity and Autoencoder Losses
by: Takmaz, Ece, et al.
Published: (2025)
by: Takmaz, Ece, et al.
Published: (2025)
Precision or Recall? An Analysis of Image Captions for Training Text-to-Image Generation Model
by: Cheng, Sheng, et al.
Published: (2024)
by: Cheng, Sheng, et al.
Published: (2024)
Transformer based Multitask Learning for Image Captioning and Object Detection
by: Basak, Debolena, et al.
Published: (2024)
by: Basak, Debolena, et al.
Published: (2024)
Altogether: Image Captioning via Re-aligning Alt-text
by: Xu, Hu, et al.
Published: (2024)
by: Xu, Hu, et al.
Published: (2024)
OmniCaptioner: One Captioner to Rule Them All
by: Lu, Yiting, et al.
Published: (2025)
by: Lu, Yiting, et al.
Published: (2025)
From Pixels to Prose: A Large Dataset of Dense Image Captions
by: Singla, Vasu, et al.
Published: (2024)
by: Singla, Vasu, et al.
Published: (2024)
Improving Image Captioning by Mimicking Human Reformulation Feedback at Inference-time
by: Berger, Uri, et al.
Published: (2025)
by: Berger, Uri, et al.
Published: (2025)
SPECS: Specificity-Enhanced CLIP-Score for Long Image Caption Evaluation
by: Chen, Xiaofu, et al.
Published: (2025)
by: Chen, Xiaofu, et al.
Published: (2025)
#PraCegoVer: A Large Dataset for Image Captioning in Portuguese
by: Santos, Gabriel Oliveira dos, et al.
Published: (2021)
by: Santos, Gabriel Oliveira dos, et al.
Published: (2021)
Towards Adaptable and Interactive Image Captioning with Data Augmentation and Episodic Memory
by: Anagnostopoulou, Aliki, et al.
Published: (2023)
by: Anagnostopoulou, Aliki, et al.
Published: (2023)
BRIDGE: Bridging Gaps in Image Captioning Evaluation with Stronger Visual Cues
by: Sarto, Sara, et al.
Published: (2024)
by: Sarto, Sara, et al.
Published: (2024)
GeReA: Question-Aware Prompt Captions for Knowledge-based Visual Question Answering
by: Ma, Ziyu, et al.
Published: (2024)
by: Ma, Ziyu, et al.
Published: (2024)
Regional Attention-Enhanced Swin Transformer for Clinically Relevant Medical Image Captioning
by: Naz, Zubia, et al.
Published: (2025)
by: Naz, Zubia, et al.
Published: (2025)
Chain-of-Talkers (CoTalk): Fast Human Annotation of Dense Image Captions
by: Shen, Yijun, et al.
Published: (2025)
by: Shen, Yijun, et al.
Published: (2025)
Text-Only Training for Image Captioning with Retrieval Augmentation and Modality Gap Correction
by: Fonseca, Rui, et al.
Published: (2025)
by: Fonseca, Rui, et al.
Published: (2025)
EAMA : Entity-Aware Multimodal Alignment Based Approach for News Image Captioning
by: Zhang, Junzhe, et al.
Published: (2024)
by: Zhang, Junzhe, et al.
Published: (2024)
Diffusion-RSCC: Diffusion Probabilistic Model for Change Captioning in Remote Sensing Images
by: Yu, Xiaofei, et al.
Published: (2024)
by: Yu, Xiaofei, et al.
Published: (2024)
LAViTeR: Learning Aligned Visual and Textual Representations Assisted by Image and Caption Generation
by: Hashemi, Mohammad Abuzar, et al.
Published: (2021)
by: Hashemi, Mohammad Abuzar, et al.
Published: (2021)
FigEx2: Visual-Conditioned Panel Detection and Captioning for Scientific Compound Figures
by: Song, Jifeng, et al.
Published: (2026)
by: Song, Jifeng, et al.
Published: (2026)
Text-only Synthesis for Image Captioning
by: Zhou, Qing, et al.
Published: (2024)
by: Zhou, Qing, et al.
Published: (2024)
The Role of Data Curation in Image Captioning
by: Li, Wenyan, et al.
Published: (2023)
by: Li, Wenyan, et al.
Published: (2023)
BIOMEDICA: An Open Biomedical Image-Caption Archive, Dataset, and Vision-Language Models Derived from Scientific Literature
by: Lozano, Alejandro, et al.
Published: (2025)
by: Lozano, Alejandro, et al.
Published: (2025)
CultureCLIP: Empowering CLIP with Cultural Awareness through Synthetic Images and Contextualized Captions
by: Huang, Yuchen, et al.
Published: (2025)
by: Huang, Yuchen, et al.
Published: (2025)
ScaleCap: Inference-Time Scalable Image Captioning via Dual-Modality Debiasing
by: Xing, Long, et al.
Published: (2025)
by: Xing, Long, et al.
Published: (2025)
Similar Items
-
Context-aware Visual Storytelling with Visual Prefix Tuning and Contrastive Learning
by: Song, Yingjin, et al.
Published: (2024) -
Burn After Reading: Do Multimodal Large Language Models Truly Capture Order of Events in Image Sequences?
by: Song, Yingjin, et al.
Published: (2025) -
Common Objects Out of Context (COOCo): Investigating Multimodal Context and Semantic Scene Violations in Referential Communication
by: Merlo, Filippo, et al.
Published: (2025) -
Generating Storytelling Images with Rich Chains-of-Reasoning
by: Song, Xiujie, et al.
Published: (2025) -
VIXEN: Visual Text Comparison Network for Image Difference Captioning
by: Black, Alexander, et al.
Published: (2024)