Audio Description Generation in the Era of LLMs and VLMs: A Review of Transferable Generative AI Technologies
Fuente:
arXiv
Saved in:
| Main Authors: | Gao, Yingqiang, Fischer, Lukas, Lintner, Alexa, Ebling, Sarah |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
SwissADT: An Audio Description Translation System for Swiss Languages
by: Fischer, Lukas, et al.
Published: (2024)
by: Fischer, Lukas, et al.
Published: (2024)
Global-Local Tree Search in VLMs for 3D Indoor Scene Generation
by: Deng, Wei, et al.
Published: (2025)
by: Deng, Wei, et al.
Published: (2025)
Shakti-VLMs: Scalable Vision-Language Models for Enterprise AI
by: Shakhadri, Syed Abdul Gaffar, et al.
Published: (2025)
by: Shakhadri, Syed Abdul Gaffar, et al.
Published: (2025)
Multimodal Event Detection: Current Approaches and Defining the New Playground through LLMs and VLMs
by: Dey, Abhishek, et al.
Published: (2025)
by: Dey, Abhishek, et al.
Published: (2025)
Sensitivity of Generative VLMs to Semantically and Lexically Altered Prompts
by: Dumpala, Sri Harsha, et al.
Published: (2024)
by: Dumpala, Sri Harsha, et al.
Published: (2024)
Are VLMs Really Blind
by: Singh, Ayush, et al.
Published: (2024)
by: Singh, Ayush, et al.
Published: (2024)
More than a Moment: Towards Coherent Sequences of Audio Descriptions
by: Khandelwal, Eshika, et al.
Published: (2025)
by: Khandelwal, Eshika, et al.
Published: (2025)
AI-Generated Fall Data: Assessing LLMs and Diffusion Model for Wearable Fall Detection
by: Alamgeer, Sana, et al.
Published: (2025)
by: Alamgeer, Sana, et al.
Published: (2025)
Generative Technology for Human Emotion Recognition: A Scope Review
by: Ma, Fei, et al.
Published: (2024)
by: Ma, Fei, et al.
Published: (2024)
On the Perception Bottleneck of VLMs for Chart Understanding
by: Liu, Junteng, et al.
Published: (2025)
by: Liu, Junteng, et al.
Published: (2025)
CIVET: Systematic Evaluation of Understanding in VLMs
by: Rizzoli, Massimo, et al.
Published: (2025)
by: Rizzoli, Massimo, et al.
Published: (2025)
Prism: A Framework for Decoupling and Assessing the Capabilities of VLMs
by: Qiao, Yuxuan, et al.
Published: (2024)
by: Qiao, Yuxuan, et al.
Published: (2024)
A Multimodal In-Context Tuning Approach for E-Commerce Product Description Generation
by: Li, Yunxin, et al.
Published: (2024)
by: Li, Yunxin, et al.
Published: (2024)
Investigating Disability Representations in Text-to-Image Models
by: Tian, Yang, et al.
Published: (2026)
by: Tian, Yang, et al.
Published: (2026)
[De|Re]constructing VLMs' Reasoning in Counting
by: Alghisi, Simone, et al.
Published: (2025)
by: Alghisi, Simone, et al.
Published: (2025)
AnnoTheia: A Semi-Automatic Annotation Toolkit for Audio-Visual Speech Technologies
by: Acosta-Triana, José-M., et al.
Published: (2024)
by: Acosta-Triana, José-M., et al.
Published: (2024)
GrAInS: Gradient-based Attribution for Inference-Time Steering of LLMs and VLMs
by: Nguyen, Duy, et al.
Published: (2025)
by: Nguyen, Duy, et al.
Published: (2025)
Generalizing from SIMPLE to HARD Visual Reasoning: Can We Mitigate Modality Imbalance in VLMs?
by: Park, Simon, et al.
Published: (2025)
by: Park, Simon, et al.
Published: (2025)
Fine-Grained Preference Optimization Improves Spatial Reasoning in VLMs
by: Shen, Yifan, et al.
Published: (2025)
by: Shen, Yifan, et al.
Published: (2025)
Iteratively Prompting Multimodal LLMs to Reproduce Natural and AI-Generated Images
by: Naseh, Ali, et al.
Published: (2024)
by: Naseh, Ali, et al.
Published: (2024)
LLMPopcorn: Exploring LLMs as Assistants for Popular Micro-video Generation
by: Fu, Junchen, et al.
Published: (2025)
by: Fu, Junchen, et al.
Published: (2025)
Medical Reasoning in the Era of LLMs: A Systematic Review of Enhancement Techniques and Applications
by: Wang, Wenxuan, et al.
Published: (2025)
by: Wang, Wenxuan, et al.
Published: (2025)
Level Up Your Tutorials: VLMs for Game Tutorials Quality Assessment
by: Cambrin, Daniele Rege, et al.
Published: (2024)
by: Cambrin, Daniele Rege, et al.
Published: (2024)
Smart Eyes for Silent Threats: VLMs and In-Context Learning for THz Imaging
by: Poggi, Nicolas, et al.
Published: (2025)
by: Poggi, Nicolas, et al.
Published: (2025)
VLMs-in-the-Wild: Bridging the Gap Between Academic Benchmarks and Enterprise Reality
by: Bandraupalli, Srihari, et al.
Published: (2025)
by: Bandraupalli, Srihari, et al.
Published: (2025)
Are VLMs Seeing or Just Saying? Uncovering the Illusion of Visual Re-examination
by: Shi, Chufan, et al.
Published: (2026)
by: Shi, Chufan, et al.
Published: (2026)
Reasmory: 3D Reconstruction as Explicit Memory for VLMs Spatial Reasoning
by: He, Jixuan, et al.
Published: (2026)
by: He, Jixuan, et al.
Published: (2026)
Audio-visual training for improved grounding in video-text LLMs
by: Sagare, Shivprasad, et al.
Published: (2024)
by: Sagare, Shivprasad, et al.
Published: (2024)
VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs
by: Cheng, Zesen, et al.
Published: (2024)
by: Cheng, Zesen, et al.
Published: (2024)
Learning Human-Perceived Fakeness in AI-Generated Videos via Multimodal LLMs
by: Fu, Xingyu, et al.
Published: (2025)
by: Fu, Xingyu, et al.
Published: (2025)
Unification of Balti and trans-border sister dialects in the essence of LLMs and AI Technology
by: Sharif, Muhammad, et al.
Published: (2024)
by: Sharif, Muhammad, et al.
Published: (2024)
Template-Based Text-to-Image Alignment for Language Accessibility: A Study on Visualizing Text Simplifications
by: Souayed, Belkiss, et al.
Published: (2025)
by: Souayed, Belkiss, et al.
Published: (2025)
What You See is What You Ask: Evaluating Audio Descriptions
by: Kala, Divy, et al.
Published: (2025)
by: Kala, Divy, et al.
Published: (2025)
A Systematic Review of Deep Learning-based Research on Radiology Report Generation
by: Liu, Chang, et al.
Published: (2023)
by: Liu, Chang, et al.
Published: (2023)
Beyond End-to-End VLMs: Leveraging Intermediate Text Representations for Superior Flowchart Understanding
by: Ye, Junyi, et al.
Published: (2024)
by: Ye, Junyi, et al.
Published: (2024)
Do Text Edits Generalize to Visual Generation? Benchmarking Cross-Modal Knowledge Editing in UMMs
by: Gao, Xin, et al.
Published: (2026)
by: Gao, Xin, et al.
Published: (2026)
SK-VQA: Synthetic Knowledge Generation at Scale for Training Context-Augmented Multimodal LLMs
by: Su, Xin, et al.
Published: (2024)
by: Su, Xin, et al.
Published: (2024)
Personalized Scientific Figure Caption Generation: An Empirical Study on Author-Specific Writing Style Transfer
by: Kim, Jaeyoung, et al.
Published: (2025)
by: Kim, Jaeyoung, et al.
Published: (2025)
Decompose and Compare Consistency: Measuring VLMs' Answer Reliability via Task-Decomposition Consistency Comparison
by: Yang, Qian, et al.
Published: (2024)
by: Yang, Qian, et al.
Published: (2024)
VLMs Need Words: Vision Language Models Ignore Visual Detail In Favor of Semantic Anchors
by: Shahgir, Haz Sameen, et al.
Published: (2026)
by: Shahgir, Haz Sameen, et al.
Published: (2026)
Similar Items
-
SwissADT: An Audio Description Translation System for Swiss Languages
by: Fischer, Lukas, et al.
Published: (2024) -
Global-Local Tree Search in VLMs for 3D Indoor Scene Generation
by: Deng, Wei, et al.
Published: (2025) -
Shakti-VLMs: Scalable Vision-Language Models for Enterprise AI
by: Shakhadri, Syed Abdul Gaffar, et al.
Published: (2025) -
Multimodal Event Detection: Current Approaches and Defining the New Playground through LLMs and VLMs
by: Dey, Abhishek, et al.
Published: (2025) -
Sensitivity of Generative VLMs to Semantically and Lexically Altered Prompts
by: Dumpala, Sri Harsha, et al.
Published: (2024)