IFCap: Image-like Retrieval and Frequency-based Entity Filtering for Zero-shot Captioning
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Lee, Soeun, Kim, Si-Woo, Kim, Taewhan, Kim, Dong-Jin |
|---|---|
| Format: | Preprint |
| Publié: |
2024
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
ViPCap: Retrieval Text-Based Visual Prompts for Lightweight Image Captioning
par: Kim, Taewhan, et autres
Publié: (2024)
par: Kim, Taewhan, et autres
Publié: (2024)
SynC: Synthetic Image Caption Dataset Refinement with One-to-many Mapping for Zero-shot Image Captioning
par: Kim, Si-Woo, et autres
Publié: (2025)
par: Kim, Si-Woo, et autres
Publié: (2025)
SIDA: Synthetic Image Driven Zero-shot Domain Adaptation
par: Kim, Ye-Chan, et autres
Publié: (2025)
par: Kim, Ye-Chan, et autres
Publié: (2025)
SAIL: Similarity-Aware Guidance and Inter-Caption Augmentation-based Learning for Weakly-Supervised Dense Video Captioning
par: Kim, Ye-Chan, et autres
Publié: (2026)
par: Kim, Ye-Chan, et autres
Publié: (2026)
Sali4Vid: Saliency-Aware Video Reweighting and Adaptive Caption Retrieval for Dense Video Captioning
par: Jeon, MinJu, et autres
Publié: (2025)
par: Jeon, MinJu, et autres
Publié: (2025)
CIC: A Framework for Culturally-Aware Image Captioning
par: Yun, Youngsik, et autres
Publié: (2024)
par: Yun, Youngsik, et autres
Publié: (2024)
EXPERT: An Explainable Image Captioning Evaluation Metric with Structured Explanations
par: Kim, Hyunjong, et autres
Publié: (2025)
par: Kim, Hyunjong, et autres
Publié: (2025)
An Image Grid Can Be Worth a Video: Zero-shot Video Question Answering Using a VLM
par: Kim, Wonkyun, et autres
Publié: (2024)
par: Kim, Wonkyun, et autres
Publié: (2024)
Modality-Aware Representation Learning for Zero-shot Sketch-based Image Retrieval
par: Lyou, Eunyi, et autres
Publié: (2024)
par: Lyou, Eunyi, et autres
Publié: (2024)
Unifying Vision-Language Latents for Zero-label Image Caption Enhancement
par: Byun, Sanghyun, et autres
Publié: (2025)
par: Byun, Sanghyun, et autres
Publié: (2025)
ChartCap: Mitigating Hallucination of Dense Chart Captioning
par: Lim, Junyoung, et autres
Publié: (2025)
par: Lim, Junyoung, et autres
Publié: (2025)
Zero-shot Text-guided Infinite Image Synthesis with LLM guidance
par: Kwon, Soyeong, et autres
Publié: (2024)
par: Kwon, Soyeong, et autres
Publié: (2024)
See It All: Contextualized Late Aggregation for 3D Dense Captioning
par: Kim, Minjung, et autres
Publié: (2024)
par: Kim, Minjung, et autres
Publié: (2024)
Towards Retrieval-Augmented Architectures for Image Captioning
par: Sarto, Sara, et autres
Publié: (2024)
par: Sarto, Sara, et autres
Publié: (2024)
Video Summarization: Towards Entity-Aware Captions
par: Ayyubi, Hammad A., et autres
Publié: (2023)
par: Ayyubi, Hammad A., et autres
Publié: (2023)
Decoding fMRI Data into Captions using Prefix Language Modeling
par: Shen, Vyacheslav, et autres
Publié: (2025)
par: Shen, Vyacheslav, et autres
Publié: (2025)
Completely Weakly Supervised Class-Incremental Learning for Semantic Segmentation
par: Kim, David Minkwan, et autres
Publié: (2025)
par: Kim, David Minkwan, et autres
Publié: (2025)
Knowledge Generation for Zero-shot Knowledge-based VQA
par: Cao, Rui, et autres
Publié: (2024)
par: Cao, Rui, et autres
Publié: (2024)
Understanding Retrieval Robustness for Retrieval-Augmented Image Captioning
par: Li, Wenyan, et autres
Publié: (2024)
par: Li, Wenyan, et autres
Publié: (2024)
Preserving Multi-Modal Capabilities of Pre-trained VLMs for Improving Vision-Linguistic Compositionality
par: Oh, Youngtaek, et autres
Publié: (2024)
par: Oh, Youngtaek, et autres
Publié: (2024)
EAMA : Entity-Aware Multimodal Alignment Based Approach for News Image Captioning
par: Zhang, Junzhe, et autres
Publié: (2024)
par: Zhang, Junzhe, et autres
Publié: (2024)
Text-only Synthesis for Image Captioning
par: Zhou, Qing, et autres
Publié: (2024)
par: Zhou, Qing, et autres
Publié: (2024)
The Role of Data Curation in Image Captioning
par: Li, Wenyan, et autres
Publié: (2023)
par: Li, Wenyan, et autres
Publié: (2023)
Personalized Scientific Figure Caption Generation: An Empirical Study on Author-Specific Writing Style Transfer
par: Kim, Jaeyoung, et autres
Publié: (2025)
par: Kim, Jaeyoung, et autres
Publié: (2025)
LITTA: Late-Interaction and Test-Time Alignment for Visually-Grounded Multimodal Retrieval
par: Kim, Seonok
Publié: (2026)
par: Kim, Seonok
Publié: (2026)
Contrastive Language Prompting to Ease False Positives in Medical Anomaly Detection
par: Park, YeongHyeon, et autres
Publié: (2024)
par: Park, YeongHyeon, et autres
Publié: (2024)
From Pixels to Posts: Retrieval-Augmented Fashion Captioning and Hashtag Generation
par: Gondal, Moazzam Umer, et autres
Publié: (2025)
par: Gondal, Moazzam Umer, et autres
Publié: (2025)
Retrieval-Augmented Long-Context Translation for Cultural Image Captioning: Gators submission for AmericasNLP 2026 shared task
par: Dhawan, Aashish, et autres
Publié: (2026)
par: Dhawan, Aashish, et autres
Publié: (2026)
Temporal Image Caption Retrieval Competition -- Description and Results
par: Pokrywka, Jakub, et autres
Publié: (2024)
par: Pokrywka, Jakub, et autres
Publié: (2024)
LaMP-Cap: Personalized Figure Caption Generation With Multimodal Figure Profiles
par: Ng, Ho Yin 'Sam', et autres
Publié: (2025)
par: Ng, Ho Yin 'Sam', et autres
Publié: (2025)
Speaking Beyond Language: A Large-Scale Multimodal Dataset for Learning Nonverbal Cues from Video-Grounded Dialogues
par: Kim, Youngmin, et autres
Publié: (2025)
par: Kim, Youngmin, et autres
Publié: (2025)
Generating Accurate and Detailed Captions for High-Resolution Images
par: Lee, Hankyeol, et autres
Publié: (2025)
par: Lee, Hankyeol, et autres
Publié: (2025)
CapRL: Stimulating Dense Image Caption Capabilities via Reinforcement Learning
par: Xing, Long, et autres
Publié: (2025)
par: Xing, Long, et autres
Publié: (2025)
Multi-LLM Collaborative Caption Generation in Scientific Documents
par: Kim, Jaeyoung, et autres
Publié: (2025)
par: Kim, Jaeyoung, et autres
Publié: (2025)
FLEUR: An Explainable Reference-Free Evaluation Metric for Image Captioning Using a Large Multimodal Model
par: Lee, Yebin, et autres
Publié: (2024)
par: Lee, Yebin, et autres
Publié: (2024)
MATE: Meet At The Embedding -- Connecting Images with Long Texts
par: Jang, Young Kyun, et autres
Publié: (2024)
par: Jang, Young Kyun, et autres
Publié: (2024)
InstructNav: Zero-shot System for Generic Instruction Navigation in Unexplored Environment
par: Long, Yuxing, et autres
Publié: (2024)
par: Long, Yuxing, et autres
Publié: (2024)
CAPEEN: Image Captioning with Early Exits and Knowledge Distillation
par: Bajpai, Divya Jyoti, et autres
Publié: (2024)
par: Bajpai, Divya Jyoti, et autres
Publié: (2024)
InstantFamily: Masked Attention for Zero-shot Multi-ID Image Generation
par: Kim, Chanran, et autres
Publié: (2024)
par: Kim, Chanran, et autres
Publié: (2024)
ReflectCAP: Detailed Image Captioning with Reflective Memory
par: Min, Kyungmin, et autres
Publié: (2026)
par: Min, Kyungmin, et autres
Publié: (2026)
Documents similaires
-
ViPCap: Retrieval Text-Based Visual Prompts for Lightweight Image Captioning
par: Kim, Taewhan, et autres
Publié: (2024) -
SynC: Synthetic Image Caption Dataset Refinement with One-to-many Mapping for Zero-shot Image Captioning
par: Kim, Si-Woo, et autres
Publié: (2025) -
SIDA: Synthetic Image Driven Zero-shot Domain Adaptation
par: Kim, Ye-Chan, et autres
Publié: (2025) -
SAIL: Similarity-Aware Guidance and Inter-Caption Augmentation-based Learning for Weakly-Supervised Dense Video Captioning
par: Kim, Ye-Chan, et autres
Publié: (2026) -
Sali4Vid: Saliency-Aware Video Reweighting and Adaptive Caption Retrieval for Dense Video Captioning
par: Jeon, MinJu, et autres
Publié: (2025)