CIC-BART-SSA: Controllable Image Captioning with Structured Semantic Augmentation
Fuente:
arXiv
Guardado en:
| Autores principales: | Basioti, Kalliopi, Abdelsalam, Mohamed A., Fancellu, Federico, Pavlovic, Vladimir, Fazly, Afsaneh |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
CIC: A Framework for Culturally-Aware Image Captioning
por: Yun, Youngsik, et al.
Publicado: (2024)
por: Yun, Youngsik, et al.
Publicado: (2024)
Box2Flow: Instance-based Action Flow Graphs from Videos
por: Li, Jiatong, et al.
Publicado: (2024)
por: Li, Jiatong, et al.
Publicado: (2024)
GenVP: Generating Visual Puzzles with Contrastive Hierarchical VAEs
por: Basioti, Kalliopi, et al.
Publicado: (2025)
por: Basioti, Kalliopi, et al.
Publicado: (2025)
Towards Retrieval-Augmented Architectures for Image Captioning
por: Sarto, Sara, et al.
Publicado: (2024)
por: Sarto, Sara, et al.
Publicado: (2024)
Augmenting Perceptual Super-Resolution via Image Quality Predictors
por: Zhang, Fengjia, et al.
Publicado: (2025)
por: Zhang, Fengjia, et al.
Publicado: (2025)
EXPERT: An Explainable Image Captioning Evaluation Metric with Structured Explanations
por: Kim, Hyunjong, et al.
Publicado: (2025)
por: Kim, Hyunjong, et al.
Publicado: (2025)
Text-only Synthesis for Image Captioning
por: Zhou, Qing, et al.
Publicado: (2024)
por: Zhou, Qing, et al.
Publicado: (2024)
The Role of Data Curation in Image Captioning
por: Li, Wenyan, et al.
Publicado: (2023)
por: Li, Wenyan, et al.
Publicado: (2023)
From Pixels to Posts: Retrieval-Augmented Fashion Captioning and Hashtag Generation
por: Gondal, Moazzam Umer, et al.
Publicado: (2025)
por: Gondal, Moazzam Umer, et al.
Publicado: (2025)
Retrieval-Augmented Long-Context Translation for Cultural Image Captioning: Gators submission for AmericasNLP 2026 shared task
por: Dhawan, Aashish, et al.
Publicado: (2026)
por: Dhawan, Aashish, et al.
Publicado: (2026)
Understanding Retrieval Robustness for Retrieval-Augmented Image Captioning
por: Li, Wenyan, et al.
Publicado: (2024)
por: Li, Wenyan, et al.
Publicado: (2024)
The Devil is in the EOS: Sequence Training for Detailed Image Captioning
por: Mohamed, Abdelrahman, et al.
Publicado: (2025)
por: Mohamed, Abdelrahman, et al.
Publicado: (2025)
Structured Captions Improve Prompt Adherence in Text-to-Image Models (Re-LAION-Caption 19M)
por: Merchant, Nicholas, et al.
Publicado: (2025)
por: Merchant, Nicholas, et al.
Publicado: (2025)
CAPEEN: Image Captioning with Early Exits and Knowledge Distillation
por: Bajpai, Divya Jyoti, et al.
Publicado: (2024)
por: Bajpai, Divya Jyoti, et al.
Publicado: (2024)
Towards Adaptable and Interactive Image Captioning with Data Augmentation and Episodic Memory
por: Anagnostopoulou, Aliki, et al.
Publicado: (2023)
por: Anagnostopoulou, Aliki, et al.
Publicado: (2023)
Discovering Meaningful Units with Visually Grounded Semantics from Image Captions
por: Behjati, Melika, et al.
Publicado: (2025)
por: Behjati, Melika, et al.
Publicado: (2025)
Explaining Caption-Image Interactions in CLIP Models with Second-Order Attributions
por: Möller, Lucas, et al.
Publicado: (2024)
por: Möller, Lucas, et al.
Publicado: (2024)
Image Captioning Evaluation in the Age of Multimodal LLMs: Challenges and Future Perspectives
por: Sarto, Sara, et al.
Publicado: (2025)
por: Sarto, Sara, et al.
Publicado: (2025)
The Power of Many: Multi-Agent Multimodal Models for Cultural Image Captioning
por: Bai, Longju, et al.
Publicado: (2024)
por: Bai, Longju, et al.
Publicado: (2024)
DENEB: A Hallucination-Robust Automatic Evaluation Metric for Image Captioning
por: Matsuda, Kazuki, et al.
Publicado: (2024)
por: Matsuda, Kazuki, et al.
Publicado: (2024)
Polos: Multimodal Metric Learning from Human Feedback for Image Captioning
por: Wada, Yuiga, et al.
Publicado: (2024)
por: Wada, Yuiga, et al.
Publicado: (2024)
Text-Only Training for Image Captioning with Retrieval Augmentation and Modality Gap Correction
por: Fonseca, Rui, et al.
Publicado: (2025)
por: Fonseca, Rui, et al.
Publicado: (2025)
CapRL: Stimulating Dense Image Caption Capabilities via Reinforcement Learning
por: Xing, Long, et al.
Publicado: (2025)
por: Xing, Long, et al.
Publicado: (2025)
VELA: An LLM-Hybrid-as-a-Judge Approach for Evaluating Long Image Captions
por: Matsuda, Kazuki, et al.
Publicado: (2025)
por: Matsuda, Kazuki, et al.
Publicado: (2025)
AC-Lite : A Lightweight Image Captioning Model for Low-Resource Assamese Language
por: Choudhury, Pankaj, et al.
Publicado: (2025)
por: Choudhury, Pankaj, et al.
Publicado: (2025)
WsiCaption: Multiple Instance Generation of Pathology Reports for Gigapixel Whole-Slide Images
por: Chen, Pingyi, et al.
Publicado: (2023)
por: Chen, Pingyi, et al.
Publicado: (2023)
SynC: Synthetic Image Caption Dataset Refinement with One-to-many Mapping for Zero-shot Image Captioning
por: Kim, Si-Woo, et al.
Publicado: (2025)
por: Kim, Si-Woo, et al.
Publicado: (2025)
LAViTeR: Learning Aligned Visual and Textual Representations Assisted by Image and Caption Generation
por: Hashemi, Mohammad Abuzar, et al.
Publicado: (2021)
por: Hashemi, Mohammad Abuzar, et al.
Publicado: (2021)
Unveiling the Invisible: Captioning Videos with Metaphors
por: Kalarani, Abisek Rajakumar, et al.
Publicado: (2024)
por: Kalarani, Abisek Rajakumar, et al.
Publicado: (2024)
Inserting Faces inside Captions: Image Captioning with Attention Guided Merging
por: Tevissen, Yannis, et al.
Publicado: (2024)
por: Tevissen, Yannis, et al.
Publicado: (2024)
Personalizing Multimodal Large Language Models for Image Captioning: An Experimental Analysis
por: Bucciarelli, Davide, et al.
Publicado: (2024)
por: Bucciarelli, Davide, et al.
Publicado: (2024)
BRIDGE: Bridging Gaps in Image Captioning Evaluation with Stronger Visual Cues
por: Sarto, Sara, et al.
Publicado: (2024)
por: Sarto, Sara, et al.
Publicado: (2024)
FLEUR: An Explainable Reference-Free Evaluation Metric for Image Captioning Using a Large Multimodal Model
por: Lee, Yebin, et al.
Publicado: (2024)
por: Lee, Yebin, et al.
Publicado: (2024)
G-VEval: A Versatile Metric for Evaluating Image and Video Captions Using GPT-4o
por: Tong, Tony Cheng, et al.
Publicado: (2024)
por: Tong, Tony Cheng, et al.
Publicado: (2024)
LLM as a Neural Architect: Controlled Generation of Image Captioning Models Under Strict API Contracts
por: Jesani, Krunal, et al.
Publicado: (2025)
por: Jesani, Krunal, et al.
Publicado: (2025)
Updating CLIP to Prefer Descriptions Over Captions
por: Zur, Amir, et al.
Publicado: (2024)
por: Zur, Amir, et al.
Publicado: (2024)
Revisiting Image Captioning Training Paradigm via Direct CLIP-based Optimization
por: Moratelli, Nicholas, et al.
Publicado: (2024)
por: Moratelli, Nicholas, et al.
Publicado: (2024)
ChartCap: Mitigating Hallucination of Dense Chart Captioning
por: Lim, Junyoung, et al.
Publicado: (2025)
por: Lim, Junyoung, et al.
Publicado: (2025)
Top-Down Semantic Refinement for Image Captioning
por: Zhang, Jusheng, et al.
Publicado: (2025)
por: Zhang, Jusheng, et al.
Publicado: (2025)
From Image Captioning to Visual Storytelling
por: Passadakis, Admitos, et al.
Publicado: (2025)
por: Passadakis, Admitos, et al.
Publicado: (2025)
Ejemplares similares
-
CIC: A Framework for Culturally-Aware Image Captioning
por: Yun, Youngsik, et al.
Publicado: (2024) -
Box2Flow: Instance-based Action Flow Graphs from Videos
por: Li, Jiatong, et al.
Publicado: (2024) -
GenVP: Generating Visual Puzzles with Contrastive Hierarchical VAEs
por: Basioti, Kalliopi, et al.
Publicado: (2025) -
Towards Retrieval-Augmented Architectures for Image Captioning
por: Sarto, Sara, et al.
Publicado: (2024) -
Augmenting Perceptual Super-Resolution via Image Quality Predictors
por: Zhang, Fengjia, et al.
Publicado: (2025)