TROPE: TRaining-Free Object-Part Enhancement for Seamlessly Improving Fine-Grained Zero-Shot Image Captioning
Fuente:
arXiv
Salvato in:
| Autori principali: | Feinglass, Joshua, Yang, Yezhou |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
`Eyes of a Hawk and Ears of a Fox': Part Prototype Network for Generalized Zero-Shot Learning
di: Feinglass, Joshua, et al.
Pubblicazione: (2024)
di: Feinglass, Joshua, et al.
Pubblicazione: (2024)
Precision or Recall? An Analysis of Image Captions for Training Text-to-Image Generation Model
di: Cheng, Sheng, et al.
Pubblicazione: (2024)
di: Cheng, Sheng, et al.
Pubblicazione: (2024)
Image-Caption Encoding for Improving Zero-Shot Generalization
di: Yu, Eric Yang, et al.
Pubblicazione: (2024)
di: Yu, Eric Yang, et al.
Pubblicazione: (2024)
Unifying Vision-Language Latents for Zero-label Image Caption Enhancement
di: Byun, Sanghyun, et al.
Pubblicazione: (2025)
di: Byun, Sanghyun, et al.
Pubblicazione: (2025)
Grounding Stylistic Domain Generalization with Quantitative Domain Shift Measures and Synthetic Scene Images
di: Luo, Yiran, et al.
Pubblicazione: (2024)
di: Luo, Yiran, et al.
Pubblicazione: (2024)
Fine-Grained Zero-Shot Object Detection
di: Ma, Hongxu, et al.
Pubblicazione: (2025)
di: Ma, Hongxu, et al.
Pubblicazione: (2025)
Unleashing Text-to-Image Diffusion Prior for Zero-Shot Image Captioning
di: Luo, Jianjie, et al.
Pubblicazione: (2024)
di: Luo, Jianjie, et al.
Pubblicazione: (2024)
Improving Brain-to-Image Reconstruction via Fine-Grained Text Bridging
di: Xia, Runze, et al.
Pubblicazione: (2025)
di: Xia, Runze, et al.
Pubblicazione: (2025)
Transformer based Multitask Learning for Image Captioning and Object Detection
di: Basak, Debolena, et al.
Pubblicazione: (2024)
di: Basak, Debolena, et al.
Pubblicazione: (2024)
Unblocking Fine-Grained Evaluation of Detailed Captions: An Explaining AutoRater and Critic-and-Revise Pipeline
di: Gordon, Brian, et al.
Pubblicazione: (2025)
di: Gordon, Brian, et al.
Pubblicazione: (2025)
$λ$-ECLIPSE: Multi-Concept Personalized Text-to-Image Diffusion Models by Leveraging CLIP Latent Space
di: Patel, Maitreya, et al.
Pubblicazione: (2024)
di: Patel, Maitreya, et al.
Pubblicazione: (2024)
Mining Fine-Grained Image-Text Alignment for Zero-Shot Captioning via Text-Only Training
di: Qiu, Longtian, et al.
Pubblicazione: (2024)
di: Qiu, Longtian, et al.
Pubblicazione: (2024)
Improving Image Captioning by Mimicking Human Reformulation Feedback at Inference-time
di: Berger, Uri, et al.
Pubblicazione: (2025)
di: Berger, Uri, et al.
Pubblicazione: (2025)
BACON: Improving Clarity of Image Captions via Bag-of-Concept Graphs
di: Yang, Zhantao, et al.
Pubblicazione: (2024)
di: Yang, Zhantao, et al.
Pubblicazione: (2024)
Zero-Shot Low Light Image Enhancement with Diffusion Prior
di: Cho, Joshua, et al.
Pubblicazione: (2024)
di: Cho, Joshua, et al.
Pubblicazione: (2024)
HNC: Leveraging Hard Negative Captions towards Models with Fine-Grained Visual-Linguistic Comprehension Capabilities
di: Dönmez, Esra, et al.
Pubblicazione: (2026)
di: Dönmez, Esra, et al.
Pubblicazione: (2026)
Fine-Grained Preference Optimization Improves Spatial Reasoning in VLMs
di: Shen, Yifan, et al.
Pubblicazione: (2025)
di: Shen, Yifan, et al.
Pubblicazione: (2025)
Open-Source Image Editing Models Are Zero-Shot Vision Learners
di: Liu, Wei, et al.
Pubblicazione: (2026)
di: Liu, Wei, et al.
Pubblicazione: (2026)
ZALM3: Zero-Shot Enhancement of Vision-Language Alignment via In-Context Information in Multi-Turn Multimodal Medical Dialogue
di: Li, Zhangpu, et al.
Pubblicazione: (2024)
di: Li, Zhangpu, et al.
Pubblicazione: (2024)
TripletCLIP: Improving Compositional Reasoning of CLIP via Synthetic Vision-Language Negatives
di: Patel, Maitreya, et al.
Pubblicazione: (2024)
di: Patel, Maitreya, et al.
Pubblicazione: (2024)
From Image Captioning to Visual Storytelling
di: Passadakis, Admitos, et al.
Pubblicazione: (2025)
di: Passadakis, Admitos, et al.
Pubblicazione: (2025)
Do Vision Encoders Truly Explain Object Hallucination?: Mitigating Object Hallucination via Simple Fine-Grained CLIPScore
di: Oh, Hongseok, et al.
Pubblicazione: (2025)
di: Oh, Hongseok, et al.
Pubblicazione: (2025)
Distinctive Image Captioning: Leveraging Ground Truth Captions in CLIP Guided Reinforcement Learning
di: Chaffin, Antoine, et al.
Pubblicazione: (2024)
di: Chaffin, Antoine, et al.
Pubblicazione: (2024)
Negative Entity Suppression for Zero-Shot Captioning with Synthetic Images
di: Lu, Zimao, et al.
Pubblicazione: (2025)
di: Lu, Zimao, et al.
Pubblicazione: (2025)
ConceptBed: Evaluating Concept Learning Abilities of Text-to-Image Diffusion Models
di: Patel, Maitreya, et al.
Pubblicazione: (2023)
di: Patel, Maitreya, et al.
Pubblicazione: (2023)
African or European Swallow? Benchmarking Large Vision-Language Models for Fine-Grained Object Classification
di: Geigle, Gregor, et al.
Pubblicazione: (2024)
di: Geigle, Gregor, et al.
Pubblicazione: (2024)
Fine-Grained Prototypes Distillation for Few-Shot Object Detection
di: Wang, Zichen, et al.
Pubblicazione: (2024)
di: Wang, Zichen, et al.
Pubblicazione: (2024)
Image Captioning via Compact Bidirectional Architecture
di: Song, Zijie, et al.
Pubblicazione: (2022)
di: Song, Zijie, et al.
Pubblicazione: (2022)
MUNIChus: Multilingual News Image Captioning Benchmark
di: Chen, Yuji, et al.
Pubblicazione: (2026)
di: Chen, Yuji, et al.
Pubblicazione: (2026)
TikZero: Zero-Shot Text-Guided Graphics Program Synthesis
di: Belouadi, Jonas, et al.
Pubblicazione: (2025)
di: Belouadi, Jonas, et al.
Pubblicazione: (2025)
LLM-based Hierarchical Concept Decomposition for Interpretable Fine-Grained Image Classification
di: Qu, Renyi, et al.
Pubblicazione: (2024)
di: Qu, Renyi, et al.
Pubblicazione: (2024)
Enhancing Fine-Grained Image Classifications via Cascaded Vision Language Models
di: Wei, Canshi
Pubblicazione: (2024)
di: Wei, Canshi
Pubblicazione: (2024)
SynC: Synthetic Image Caption Dataset Refinement with One-to-many Mapping for Zero-shot Image Captioning
di: Kim, Si-Woo, et al.
Pubblicazione: (2025)
di: Kim, Si-Woo, et al.
Pubblicazione: (2025)
Understanding Retrieval Robustness for Retrieval-Augmented Image Captioning
di: Li, Wenyan, et al.
Pubblicazione: (2024)
di: Li, Wenyan, et al.
Pubblicazione: (2024)
Temporal Image Caption Retrieval Competition -- Description and Results
di: Pokrywka, Jakub, et al.
Pubblicazione: (2024)
di: Pokrywka, Jakub, et al.
Pubblicazione: (2024)
The Devil is in the EOS: Sequence Training for Detailed Image Captioning
di: Mohamed, Abdelrahman, et al.
Pubblicazione: (2025)
di: Mohamed, Abdelrahman, et al.
Pubblicazione: (2025)
Zero-Shot Fine-Grained Image Classification Using Large Vision-Language Models
di: Atabuzzaman, Md., et al.
Pubblicazione: (2025)
di: Atabuzzaman, Md., et al.
Pubblicazione: (2025)
Fine-Grained Zero-Shot Composed Image Retrieval with Complementary Visual-Semantic Integration
di: Ye, Yongcong, et al.
Pubblicazione: (2026)
di: Ye, Yongcong, et al.
Pubblicazione: (2026)
Zero-Shot Action Recognition in Surveillance Videos
di: Pereira, Joao, et al.
Pubblicazione: (2024)
di: Pereira, Joao, et al.
Pubblicazione: (2024)
FineCops-Ref: A new Dataset and Task for Fine-Grained Compositional Referring Expression Comprehension
di: Liu, Junzhuo, et al.
Pubblicazione: (2024)
di: Liu, Junzhuo, et al.
Pubblicazione: (2024)
Documenti analoghi
-
`Eyes of a Hawk and Ears of a Fox': Part Prototype Network for Generalized Zero-Shot Learning
di: Feinglass, Joshua, et al.
Pubblicazione: (2024) -
Precision or Recall? An Analysis of Image Captions for Training Text-to-Image Generation Model
di: Cheng, Sheng, et al.
Pubblicazione: (2024) -
Image-Caption Encoding for Improving Zero-Shot Generalization
di: Yu, Eric Yang, et al.
Pubblicazione: (2024) -
Unifying Vision-Language Latents for Zero-label Image Caption Enhancement
di: Byun, Sanghyun, et al.
Pubblicazione: (2025) -
Grounding Stylistic Domain Generalization with Quantitative Domain Shift Measures and Synthetic Scene Images
di: Luo, Yiran, et al.
Pubblicazione: (2024)