From Pixels to Posts: Retrieval-Augmented Fashion Captioning and Hashtag Generation
Fuente:
arXiv
Guardado en:
| Autores principales: | Gondal, Moazzam Umer, Qudous, Hamad Ul, Siddiqui, Daniya, Farhan, Asma Ahmad |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Beyond the Hype: Comparing Lightweight and Deep Learning Models for Air Quality Forecasting
por: Gondal, Moazzam Umer, et al.
Publicado: (2025)
por: Gondal, Moazzam Umer, et al.
Publicado: (2025)
Interpretable PM2.5 Forecasting for Urban Air Quality: A Comparative Study of Operational Time-Series Models
por: Gondal, Moazzam Umer, et al.
Publicado: (2026)
por: Gondal, Moazzam Umer, et al.
Publicado: (2026)
Understanding Retrieval Robustness for Retrieval-Augmented Image Captioning
por: Li, Wenyan, et al.
Publicado: (2024)
por: Li, Wenyan, et al.
Publicado: (2024)
Text-Only Training for Image Captioning with Retrieval Augmentation and Modality Gap Correction
por: Fonseca, Rui, et al.
Publicado: (2025)
por: Fonseca, Rui, et al.
Publicado: (2025)
From Pixels to Prose: A Large Dataset of Dense Image Captions
por: Singla, Vasu, et al.
Publicado: (2024)
por: Singla, Vasu, et al.
Publicado: (2024)
Towards Retrieval-Augmented Architectures for Image Captioning
por: Sarto, Sara, et al.
Publicado: (2024)
por: Sarto, Sara, et al.
Publicado: (2024)
Temporal Image Caption Retrieval Competition -- Description and Results
por: Pokrywka, Jakub, et al.
Publicado: (2024)
por: Pokrywka, Jakub, et al.
Publicado: (2024)
Regional Attention-Enhanced Swin Transformer for Clinically Relevant Medical Image Captioning
por: Naz, Zubia, et al.
Publicado: (2025)
por: Naz, Zubia, et al.
Publicado: (2025)
Chain of Evidence: Pixel-Level Visual Attribution for Iterative Retrieval-Augmented Generation
por: Liu, Peiyang, et al.
Publicado: (2026)
por: Liu, Peiyang, et al.
Publicado: (2026)
From Image Captioning to Visual Storytelling
por: Passadakis, Admitos, et al.
Publicado: (2025)
por: Passadakis, Admitos, et al.
Publicado: (2025)
Towards Adaptable and Interactive Image Captioning with Data Augmentation and Episodic Memory
por: Anagnostopoulou, Aliki, et al.
Publicado: (2023)
por: Anagnostopoulou, Aliki, et al.
Publicado: (2023)
MobileRAG: Enhancing Mobile Agent with Retrieval-Augmented Generation
por: Loo, Gowen, et al.
Publicado: (2025)
por: Loo, Gowen, et al.
Publicado: (2025)
The Devil is in the Prompts: Retrieval-Augmented Prompt Optimization for Text-to-Video Generation
por: Gao, Bingjie, et al.
Publicado: (2025)
por: Gao, Bingjie, et al.
Publicado: (2025)
AIP: Subverting Retrieval-Augmented Generation via Adversarial Instructional Prompt
por: Chaturvedi, Saket S., et al.
Publicado: (2025)
por: Chaturvedi, Saket S., et al.
Publicado: (2025)
LaB-RAG: Label Boosted Retrieval Augmented Generation for Radiology Report Generation
por: Song, Steven, et al.
Publicado: (2024)
por: Song, Steven, et al.
Publicado: (2024)
OmniCaptioner: One Captioner to Rule Them All
por: Lu, Yiting, et al.
Publicado: (2025)
por: Lu, Yiting, et al.
Publicado: (2025)
Multi-LLM Collaborative Caption Generation in Scientific Documents
por: Kim, Jaeyoung, et al.
Publicado: (2025)
por: Kim, Jaeyoung, et al.
Publicado: (2025)
Retrieval-Augmented Long-Context Translation for Cultural Image Captioning: Gators submission for AmericasNLP 2026 shared task
por: Dhawan, Aashish, et al.
Publicado: (2026)
por: Dhawan, Aashish, et al.
Publicado: (2026)
Retrieval-Augmented Egocentric Video Captioning
por: Xu, Jilan, et al.
Publicado: (2024)
por: Xu, Jilan, et al.
Publicado: (2024)
From Text to Pixel: Advancing Long-Context Understanding in MLLMs
por: Lu, Yujie, et al.
Publicado: (2024)
por: Lu, Yujie, et al.
Publicado: (2024)
TurboRAG: Accelerating Retrieval-Augmented Generation with Precomputed KV Caches for Chunked Text
por: Lu, Songshuo, et al.
Publicado: (2024)
por: Lu, Songshuo, et al.
Publicado: (2024)
AgenticOCR: Parsing Only What You Need for Efficient Retrieval-Augmented Generation
por: Wang, Zhengren, et al.
Publicado: (2026)
por: Wang, Zhengren, et al.
Publicado: (2026)
Scaling Beyond Context: A Survey of Multimodal Retrieval-Augmented Generation for Document Understanding
por: Gao, Sensen, et al.
Publicado: (2025)
por: Gao, Sensen, et al.
Publicado: (2025)
A Video is Worth 10,000 Words: Training and Benchmarking with Diverse Captions for Better Long Video Retrieval
por: Gwilliam, Matthew, et al.
Publicado: (2023)
por: Gwilliam, Matthew, et al.
Publicado: (2023)
From Reasoning to Pixels: Benchmarking the Alignment Gap in Unified Multimodal Models
por: Yang, Cheng, et al.
Publicado: (2026)
por: Yang, Cheng, et al.
Publicado: (2026)
Open Vocabulary Panoptic Segmentation With Retrieval Augmentation
por: Sadeq, Nafis, et al.
Publicado: (2026)
por: Sadeq, Nafis, et al.
Publicado: (2026)
Fashion-RAG: Multimodal Fashion Image Editing via Retrieval-Augmented Generation
por: Sanguigni, Fulvio, et al.
Publicado: (2025)
por: Sanguigni, Fulvio, et al.
Publicado: (2025)
Visual-RAG: Benchmarking Text-to-Image Retrieval Augmented Generation for Visual Knowledge Intensive Queries
por: Wu, Yin, et al.
Publicado: (2025)
por: Wu, Yin, et al.
Publicado: (2025)
Distinctive Image Captioning: Leveraging Ground Truth Captions in CLIP Guided Reinforcement Learning
por: Chaffin, Antoine, et al.
Publicado: (2024)
por: Chaffin, Antoine, et al.
Publicado: (2024)
From Pixels to Tokens: Revisiting Object Hallucinations in Large Vision-Language Models
por: Shang, Yuying, et al.
Publicado: (2024)
por: Shang, Yuying, et al.
Publicado: (2024)
PixelWorld: How Far Are We from Perceiving Everything as Pixels?
por: Lyu, Zhiheng, et al.
Publicado: (2025)
por: Lyu, Zhiheng, et al.
Publicado: (2025)
Precision or Recall? An Analysis of Image Captions for Training Text-to-Image Generation Model
por: Cheng, Sheng, et al.
Publicado: (2024)
por: Cheng, Sheng, et al.
Publicado: (2024)
Culture-Aware Humorous Captioning: Multimodal Humor Generation across Cultural Contexts
por: Xu, Run, et al.
Publicado: (2026)
por: Xu, Run, et al.
Publicado: (2026)
CAIRe: Cultural Attribution of Images by Retrieval-Augmented Evaluation
por: Yayavaram, Arnav, et al.
Publicado: (2025)
por: Yayavaram, Arnav, et al.
Publicado: (2025)
VimRAG: Navigating Massive Visual Context in Retrieval-Augmented Generation via Multimodal Memory Graph
por: Wang, Qiuchen, et al.
Publicado: (2026)
por: Wang, Qiuchen, et al.
Publicado: (2026)
VisRAG 2.0: Evidence-Guided Multi-Image Reasoning in Visual Retrieval-Augmented Generation
por: Sun, Yubo, et al.
Publicado: (2025)
por: Sun, Yubo, et al.
Publicado: (2025)
Windsock is Dancing: Adaptive Multimodal Retrieval-Augmented Generation
por: Zhao, Shu, et al.
Publicado: (2025)
por: Zhao, Shu, et al.
Publicado: (2025)
Visual Lifelog Retrieval through Captioning-Enhanced Interpretation
por: Shih, Yu-Fei, et al.
Publicado: (2025)
por: Shih, Yu-Fei, et al.
Publicado: (2025)
LLMVA-GEBC: Large Language Model with Video Adapter for Generic Event Boundary Captioning
por: Tang, Yolo Yunlong, et al.
Publicado: (2023)
por: Tang, Yolo Yunlong, et al.
Publicado: (2023)
Diffusion Based Augmentation for Captioning and Retrieval in Cultural Heritage
por: Cioni, Dario, et al.
Publicado: (2023)
por: Cioni, Dario, et al.
Publicado: (2023)
Ejemplares similares
-
Beyond the Hype: Comparing Lightweight and Deep Learning Models for Air Quality Forecasting
por: Gondal, Moazzam Umer, et al.
Publicado: (2025) -
Interpretable PM2.5 Forecasting for Urban Air Quality: A Comparative Study of Operational Time-Series Models
por: Gondal, Moazzam Umer, et al.
Publicado: (2026) -
Understanding Retrieval Robustness for Retrieval-Augmented Image Captioning
por: Li, Wenyan, et al.
Publicado: (2024) -
Text-Only Training for Image Captioning with Retrieval Augmentation and Modality Gap Correction
por: Fonseca, Rui, et al.
Publicado: (2025) -
From Pixels to Prose: A Large Dataset of Dense Image Captions
por: Singla, Vasu, et al.
Publicado: (2024)