VIST-GPT: Ushering in the Era of Visual Storytelling with LLMs?
Fuente:
arXiv
Saved in:
| Main Authors: | Gado, Mohamed, Taliee, Towhid, Memon, Muhammad, Ignatov, Dmitry, Timofte, Radu |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
LLM as a Neural Architect: Controlled Generation of Image Captioning Models Under Strict API Contracts
by: Jesani, Krunal, et al.
Published: (2025)
by: Jesani, Krunal, et al.
Published: (2025)
Delta-Based Neural Architecture Search: LLM Fine-Tuning via Code Diffs
by: Adhikari, Santosh Premi, et al.
Published: (2026)
by: Adhikari, Santosh Premi, et al.
Published: (2026)
SCO-VIST: Social Interaction Commonsense Knowledge-based Visual Storytelling
by: Wang, Eileen, et al.
Published: (2024)
by: Wang, Eileen, et al.
Published: (2024)
Virtually Enriched NYU Depth V2 Dataset for Monocular Depth Estimation: Do We Need Artificial Augmentation?
by: Ignatov, Dmitry, et al.
Published: (2024)
by: Ignatov, Dmitry, et al.
Published: (2024)
MobileAgeNet: Lightweight Facial Age Estimation for Mobile Deployment
by: Kumar, Arun, et al.
Published: (2026)
by: Kumar, Arun, et al.
Published: (2026)
From Brute Force to Semantic Insight: Performance-Guided Data Transformation Design with LLMs
by: Shrestha, Usha, et al.
Published: (2026)
by: Shrestha, Usha, et al.
Published: (2026)
From Code to Prediction: Fine-Tuning LLMs for Neural Network Performance Classification in NNGPT
by: Hanouneh, Mahmoud, et al.
Published: (2026)
by: Hanouneh, Mahmoud, et al.
Published: (2026)
Enhancing LLM-Based Neural Network Generation: Few-Shot Prompting and Efficient Validation for Automated Architecture Design
by: Duvvuri, Raghuvir, et al.
Published: (2025)
by: Duvvuri, Raghuvir, et al.
Published: (2025)
Preparation of Fractal-Inspired Computational Architectures for Advanced Large Language Model Analysis
by: Mittal, Yash, et al.
Published: (2025)
by: Mittal, Yash, et al.
Published: (2025)
From Memorization to Creativity: LLM as a Designer of Novel Neural Architectures
by: Khalid, Waleed, et al.
Published: (2026)
by: Khalid, Waleed, et al.
Published: (2026)
Learning Transformer-based World Models with Contrastive Predictive Coding
by: Burchi, Maxime, et al.
Published: (2025)
by: Burchi, Maxime, et al.
Published: (2025)
Accurate and Efficient World Modeling with Masked Latent Transformers
by: Burchi, Maxime, et al.
Published: (2025)
by: Burchi, Maxime, et al.
Published: (2025)
HuatuoGPT-Vision, Towards Injecting Medical Visual Knowledge into Multimodal LLMs at Scale
by: Chen, Junying, et al.
Published: (2024)
by: Chen, Junying, et al.
Published: (2024)
AugmentGest: Can Random Data Cropping Augmentation Boost Gesture Recognition Performance?
by: Aboudeshish, Nada, et al.
Published: (2025)
by: Aboudeshish, Nada, et al.
Published: (2025)
e5-omni: Explicit Cross-modal Alignment for Omni-modal Embeddings
by: Chen, Haonan, et al.
Published: (2026)
by: Chen, Haonan, et al.
Published: (2026)
Not (yet) the whole story: Evaluating Visual Storytelling Requires More than Measuring Coherence, Grounding, and Repetition
by: Surikuchi, Aditya K, et al.
Published: (2024)
by: Surikuchi, Aditya K, et al.
Published: (2024)
TARN-VIST: Topic Aware Reinforcement Network for Visual Storytelling
by: Chen, Weiran, et al.
Published: (2024)
by: Chen, Weiran, et al.
Published: (2024)
Closed-Loop LLM Discovery of Non-Standard Channel Priors in Vision Models
by: Uzun, Tolgay Atinc, et al.
Published: (2026)
by: Uzun, Tolgay Atinc, et al.
Published: (2026)
Medical Reasoning in the Era of LLMs: A Systematic Review of Enhancement Techniques and Applications
by: Wang, Wenxuan, et al.
Published: (2025)
by: Wang, Wenxuan, et al.
Published: (2025)
Real Image Denoising with Knowledge Distillation for High-Performance Mobile NPUs
by: Kayani, Faraz, et al.
Published: (2026)
by: Kayani, Faraz, et al.
Published: (2026)
A Retrieval-Augmented Generation Approach to Extracting Algorithmic Logic from Neural Networks
by: Khalid, Waleed, et al.
Published: (2025)
by: Khalid, Waleed, et al.
Published: (2025)
SPoRC-VIST: A Benchmark for Evaluating Generative Natural Narrative in Vision-Language Models
by: Zeng, Yunlin
Published: (2026)
by: Zeng, Yunlin
Published: (2026)
MuDreamer: Learning Predictive World Models without Reconstruction
by: Burchi, Maxime, et al.
Published: (2024)
by: Burchi, Maxime, et al.
Published: (2024)
Learned Lightweight Smartphone ISP with Unpaired Data
by: Arhire, Andrei, et al.
Published: (2025)
by: Arhire, Andrei, et al.
Published: (2025)
LLMs Can Compensate for Deficiencies in Visual Representations
by: Takishita, Sho, et al.
Published: (2025)
by: Takishita, Sho, et al.
Published: (2025)
ShizhenGPT: Towards Multimodal LLMs for Traditional Chinese Medicine
by: Chen, Junying, et al.
Published: (2025)
by: Chen, Junying, et al.
Published: (2025)
WaveHiT-SR: Hierarchical Wavelet Network for Efficient Image Super-Resolution
by: Ali, Fayaz, et al.
Published: (2025)
by: Ali, Fayaz, et al.
Published: (2025)
Is Your Image a Good Storyteller?
by: Song, Xiujie, et al.
Published: (2024)
by: Song, Xiujie, et al.
Published: (2024)
DefAn: Definitive Answer Dataset for LLMs Hallucination Evaluation
by: Rahman, A B M Ashikur, et al.
Published: (2024)
by: Rahman, A B M Ashikur, et al.
Published: (2024)
The LLM Bottleneck: Why Open-Source Vision LLMs Struggle with Hierarchical Visual Recognition
by: Tan, Yuwen, et al.
Published: (2025)
by: Tan, Yuwen, et al.
Published: (2025)
Promptception: How Sensitive Are Large Multimodal Models to Prompts?
by: Ismithdeen, Mohamed Insaf, et al.
Published: (2025)
by: Ismithdeen, Mohamed Insaf, et al.
Published: (2025)
Machine Unlearning in the Era of Quantum Machine Learning: An Empirical Study
by: Crivoi, Carla, et al.
Published: (2025)
by: Crivoi, Carla, et al.
Published: (2025)
mBLIP: Efficient Bootstrapping of Multilingual Vision-LLMs
by: Geigle, Gregor, et al.
Published: (2023)
by: Geigle, Gregor, et al.
Published: (2023)
Unification of Balti and trans-border sister dialects in the essence of LLMs and AI Technology
by: Sharif, Muhammad, et al.
Published: (2024)
by: Sharif, Muhammad, et al.
Published: (2024)
AnyGPT: Unified Multimodal LLM with Discrete Sequence Modeling
by: Zhan, Jun, et al.
Published: (2024)
by: Zhan, Jun, et al.
Published: (2024)
From Prompts to Pavement Through Time: Temporal Grounding in Agentic Scene-to-Plan Reasoning
by: Gado, Ahmed Y., et al.
Published: (2026)
by: Gado, Ahmed Y., et al.
Published: (2026)
Harnessing GPT-4V(ision) for Insurance: A Preliminary Exploration
by: Lin, Chenwei, et al.
Published: (2024)
by: Lin, Chenwei, et al.
Published: (2024)
DiagrammerGPT: Generating Open-Domain, Open-Platform Diagrams via LLM Planning
by: Zala, Abhay, et al.
Published: (2023)
by: Zala, Abhay, et al.
Published: (2023)
VideoDirectorGPT: Consistent Multi-scene Video Generation via LLM-Guided Planning
by: Lin, Han, et al.
Published: (2023)
by: Lin, Han, et al.
Published: (2023)
PQPP: A Joint Benchmark for Text-to-Image Prompt and Query Performance Prediction
by: Poesina, Eduard, et al.
Published: (2024)
by: Poesina, Eduard, et al.
Published: (2024)
Similar Items
-
LLM as a Neural Architect: Controlled Generation of Image Captioning Models Under Strict API Contracts
by: Jesani, Krunal, et al.
Published: (2025) -
Delta-Based Neural Architecture Search: LLM Fine-Tuning via Code Diffs
by: Adhikari, Santosh Premi, et al.
Published: (2026) -
SCO-VIST: Social Interaction Commonsense Knowledge-based Visual Storytelling
by: Wang, Eileen, et al.
Published: (2024) -
Virtually Enriched NYU Depth V2 Dataset for Monocular Depth Estimation: Do We Need Artificial Augmentation?
by: Ignatov, Dmitry, et al.
Published: (2024) -
MobileAgeNet: Lightweight Facial Age Estimation for Mobile Deployment
by: Kumar, Arun, et al.
Published: (2026)