From Pixels to Posts: Retrieval-Augmented Fashion Captioning and Hashtag Generation
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Gondal, Moazzam Umer, Qudous, Hamad Ul, Siddiqui, Daniya, Farhan, Asma Ahmad |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Beyond the Hype: Comparing Lightweight and Deep Learning Models for Air Quality Forecasting
von: Gondal, Moazzam Umer, et al.
Veröffentlicht: (2025)
von: Gondal, Moazzam Umer, et al.
Veröffentlicht: (2025)
Interpretable PM2.5 Forecasting for Urban Air Quality: A Comparative Study of Operational Time-Series Models
von: Gondal, Moazzam Umer, et al.
Veröffentlicht: (2026)
von: Gondal, Moazzam Umer, et al.
Veröffentlicht: (2026)
Understanding Retrieval Robustness for Retrieval-Augmented Image Captioning
von: Li, Wenyan, et al.
Veröffentlicht: (2024)
von: Li, Wenyan, et al.
Veröffentlicht: (2024)
Text-Only Training for Image Captioning with Retrieval Augmentation and Modality Gap Correction
von: Fonseca, Rui, et al.
Veröffentlicht: (2025)
von: Fonseca, Rui, et al.
Veröffentlicht: (2025)
From Pixels to Prose: A Large Dataset of Dense Image Captions
von: Singla, Vasu, et al.
Veröffentlicht: (2024)
von: Singla, Vasu, et al.
Veröffentlicht: (2024)
Towards Retrieval-Augmented Architectures for Image Captioning
von: Sarto, Sara, et al.
Veröffentlicht: (2024)
von: Sarto, Sara, et al.
Veröffentlicht: (2024)
Temporal Image Caption Retrieval Competition -- Description and Results
von: Pokrywka, Jakub, et al.
Veröffentlicht: (2024)
von: Pokrywka, Jakub, et al.
Veröffentlicht: (2024)
Regional Attention-Enhanced Swin Transformer for Clinically Relevant Medical Image Captioning
von: Naz, Zubia, et al.
Veröffentlicht: (2025)
von: Naz, Zubia, et al.
Veröffentlicht: (2025)
Chain of Evidence: Pixel-Level Visual Attribution for Iterative Retrieval-Augmented Generation
von: Liu, Peiyang, et al.
Veröffentlicht: (2026)
von: Liu, Peiyang, et al.
Veröffentlicht: (2026)
From Image Captioning to Visual Storytelling
von: Passadakis, Admitos, et al.
Veröffentlicht: (2025)
von: Passadakis, Admitos, et al.
Veröffentlicht: (2025)
Towards Adaptable and Interactive Image Captioning with Data Augmentation and Episodic Memory
von: Anagnostopoulou, Aliki, et al.
Veröffentlicht: (2023)
von: Anagnostopoulou, Aliki, et al.
Veröffentlicht: (2023)
MobileRAG: Enhancing Mobile Agent with Retrieval-Augmented Generation
von: Loo, Gowen, et al.
Veröffentlicht: (2025)
von: Loo, Gowen, et al.
Veröffentlicht: (2025)
The Devil is in the Prompts: Retrieval-Augmented Prompt Optimization for Text-to-Video Generation
von: Gao, Bingjie, et al.
Veröffentlicht: (2025)
von: Gao, Bingjie, et al.
Veröffentlicht: (2025)
AIP: Subverting Retrieval-Augmented Generation via Adversarial Instructional Prompt
von: Chaturvedi, Saket S., et al.
Veröffentlicht: (2025)
von: Chaturvedi, Saket S., et al.
Veröffentlicht: (2025)
LaB-RAG: Label Boosted Retrieval Augmented Generation for Radiology Report Generation
von: Song, Steven, et al.
Veröffentlicht: (2024)
von: Song, Steven, et al.
Veröffentlicht: (2024)
OmniCaptioner: One Captioner to Rule Them All
von: Lu, Yiting, et al.
Veröffentlicht: (2025)
von: Lu, Yiting, et al.
Veröffentlicht: (2025)
Multi-LLM Collaborative Caption Generation in Scientific Documents
von: Kim, Jaeyoung, et al.
Veröffentlicht: (2025)
von: Kim, Jaeyoung, et al.
Veröffentlicht: (2025)
Retrieval-Augmented Long-Context Translation for Cultural Image Captioning: Gators submission for AmericasNLP 2026 shared task
von: Dhawan, Aashish, et al.
Veröffentlicht: (2026)
von: Dhawan, Aashish, et al.
Veröffentlicht: (2026)
Retrieval-Augmented Egocentric Video Captioning
von: Xu, Jilan, et al.
Veröffentlicht: (2024)
von: Xu, Jilan, et al.
Veröffentlicht: (2024)
From Text to Pixel: Advancing Long-Context Understanding in MLLMs
von: Lu, Yujie, et al.
Veröffentlicht: (2024)
von: Lu, Yujie, et al.
Veröffentlicht: (2024)
TurboRAG: Accelerating Retrieval-Augmented Generation with Precomputed KV Caches for Chunked Text
von: Lu, Songshuo, et al.
Veröffentlicht: (2024)
von: Lu, Songshuo, et al.
Veröffentlicht: (2024)
AgenticOCR: Parsing Only What You Need for Efficient Retrieval-Augmented Generation
von: Wang, Zhengren, et al.
Veröffentlicht: (2026)
von: Wang, Zhengren, et al.
Veröffentlicht: (2026)
Scaling Beyond Context: A Survey of Multimodal Retrieval-Augmented Generation for Document Understanding
von: Gao, Sensen, et al.
Veröffentlicht: (2025)
von: Gao, Sensen, et al.
Veröffentlicht: (2025)
A Video is Worth 10,000 Words: Training and Benchmarking with Diverse Captions for Better Long Video Retrieval
von: Gwilliam, Matthew, et al.
Veröffentlicht: (2023)
von: Gwilliam, Matthew, et al.
Veröffentlicht: (2023)
From Reasoning to Pixels: Benchmarking the Alignment Gap in Unified Multimodal Models
von: Yang, Cheng, et al.
Veröffentlicht: (2026)
von: Yang, Cheng, et al.
Veröffentlicht: (2026)
Open Vocabulary Panoptic Segmentation With Retrieval Augmentation
von: Sadeq, Nafis, et al.
Veröffentlicht: (2026)
von: Sadeq, Nafis, et al.
Veröffentlicht: (2026)
Fashion-RAG: Multimodal Fashion Image Editing via Retrieval-Augmented Generation
von: Sanguigni, Fulvio, et al.
Veröffentlicht: (2025)
von: Sanguigni, Fulvio, et al.
Veröffentlicht: (2025)
Visual-RAG: Benchmarking Text-to-Image Retrieval Augmented Generation for Visual Knowledge Intensive Queries
von: Wu, Yin, et al.
Veröffentlicht: (2025)
von: Wu, Yin, et al.
Veröffentlicht: (2025)
Distinctive Image Captioning: Leveraging Ground Truth Captions in CLIP Guided Reinforcement Learning
von: Chaffin, Antoine, et al.
Veröffentlicht: (2024)
von: Chaffin, Antoine, et al.
Veröffentlicht: (2024)
From Pixels to Tokens: Revisiting Object Hallucinations in Large Vision-Language Models
von: Shang, Yuying, et al.
Veröffentlicht: (2024)
von: Shang, Yuying, et al.
Veröffentlicht: (2024)
PixelWorld: How Far Are We from Perceiving Everything as Pixels?
von: Lyu, Zhiheng, et al.
Veröffentlicht: (2025)
von: Lyu, Zhiheng, et al.
Veröffentlicht: (2025)
Precision or Recall? An Analysis of Image Captions for Training Text-to-Image Generation Model
von: Cheng, Sheng, et al.
Veröffentlicht: (2024)
von: Cheng, Sheng, et al.
Veröffentlicht: (2024)
Culture-Aware Humorous Captioning: Multimodal Humor Generation across Cultural Contexts
von: Xu, Run, et al.
Veröffentlicht: (2026)
von: Xu, Run, et al.
Veröffentlicht: (2026)
CAIRe: Cultural Attribution of Images by Retrieval-Augmented Evaluation
von: Yayavaram, Arnav, et al.
Veröffentlicht: (2025)
von: Yayavaram, Arnav, et al.
Veröffentlicht: (2025)
VimRAG: Navigating Massive Visual Context in Retrieval-Augmented Generation via Multimodal Memory Graph
von: Wang, Qiuchen, et al.
Veröffentlicht: (2026)
von: Wang, Qiuchen, et al.
Veröffentlicht: (2026)
VisRAG 2.0: Evidence-Guided Multi-Image Reasoning in Visual Retrieval-Augmented Generation
von: Sun, Yubo, et al.
Veröffentlicht: (2025)
von: Sun, Yubo, et al.
Veröffentlicht: (2025)
Windsock is Dancing: Adaptive Multimodal Retrieval-Augmented Generation
von: Zhao, Shu, et al.
Veröffentlicht: (2025)
von: Zhao, Shu, et al.
Veröffentlicht: (2025)
Visual Lifelog Retrieval through Captioning-Enhanced Interpretation
von: Shih, Yu-Fei, et al.
Veröffentlicht: (2025)
von: Shih, Yu-Fei, et al.
Veröffentlicht: (2025)
LLMVA-GEBC: Large Language Model with Video Adapter for Generic Event Boundary Captioning
von: Tang, Yolo Yunlong, et al.
Veröffentlicht: (2023)
von: Tang, Yolo Yunlong, et al.
Veröffentlicht: (2023)
Diffusion Based Augmentation for Captioning and Retrieval in Cultural Heritage
von: Cioni, Dario, et al.
Veröffentlicht: (2023)
von: Cioni, Dario, et al.
Veröffentlicht: (2023)
Ähnliche Einträge
-
Beyond the Hype: Comparing Lightweight and Deep Learning Models for Air Quality Forecasting
von: Gondal, Moazzam Umer, et al.
Veröffentlicht: (2025) -
Interpretable PM2.5 Forecasting for Urban Air Quality: A Comparative Study of Operational Time-Series Models
von: Gondal, Moazzam Umer, et al.
Veröffentlicht: (2026) -
Understanding Retrieval Robustness for Retrieval-Augmented Image Captioning
von: Li, Wenyan, et al.
Veröffentlicht: (2024) -
Text-Only Training for Image Captioning with Retrieval Augmentation and Modality Gap Correction
von: Fonseca, Rui, et al.
Veröffentlicht: (2025) -
From Pixels to Prose: A Large Dataset of Dense Image Captions
von: Singla, Vasu, et al.
Veröffentlicht: (2024)