Image-Caption Encoding for Improving Zero-Shot Generalization
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Yu, Eric Yang, Liao, Christopher, Ravi, Sathvik, Tsiligkaridis, Theodoros, Kulis, Brian |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Multimodal Unsupervised Domain Generalization by Retrieving Across the Modality Gap
von: Liao, Christopher, et al.
Veröffentlicht: (2024)
von: Liao, Christopher, et al.
Veröffentlicht: (2024)
Descriptor and Word Soups: Overcoming the Parameter Efficiency Accuracy Tradeoff for Out-of-Distribution Few-shot Learning
von: Liao, Christopher, et al.
Veröffentlicht: (2023)
von: Liao, Christopher, et al.
Veröffentlicht: (2023)
Unleashing Text-to-Image Diffusion Prior for Zero-Shot Image Captioning
von: Luo, Jianjie, et al.
Veröffentlicht: (2024)
von: Luo, Jianjie, et al.
Veröffentlicht: (2024)
ERM++: An Improved Baseline for Domain Generalization
von: Teterwak, Piotr, et al.
Veröffentlicht: (2023)
von: Teterwak, Piotr, et al.
Veröffentlicht: (2023)
FocalLens: Instruction Tuning Enables Zero-Shot Conditional Image Representations
von: Hsieh, Cheng-Yu, et al.
Veröffentlicht: (2025)
von: Hsieh, Cheng-Yu, et al.
Veröffentlicht: (2025)
SynC: Synthetic Image Caption Dataset Refinement with One-to-many Mapping for Zero-shot Image Captioning
von: Kim, Si-Woo, et al.
Veröffentlicht: (2025)
von: Kim, Si-Woo, et al.
Veröffentlicht: (2025)
Is Large-Scale Pretraining the Secret to Good Domain Generalization?
von: Teterwak, Piotr, et al.
Veröffentlicht: (2024)
von: Teterwak, Piotr, et al.
Veröffentlicht: (2024)
Data Alignment for Zero-Shot Concept Generation in Dermatology AI
von: Gadgil, Soham, et al.
Veröffentlicht: (2024)
von: Gadgil, Soham, et al.
Veröffentlicht: (2024)
TROPE: TRaining-Free Object-Part Enhancement for Seamlessly Improving Fine-Grained Zero-Shot Image Captioning
von: Feinglass, Joshua, et al.
Veröffentlicht: (2024)
von: Feinglass, Joshua, et al.
Veröffentlicht: (2024)
CLIP meets DINO for Tuning Zero-Shot Classifier using Unlabeled Image Collections
von: Imam, Mohamed Fazli, et al.
Veröffentlicht: (2024)
von: Imam, Mohamed Fazli, et al.
Veröffentlicht: (2024)
Structured Captions Improve Prompt Adherence in Text-to-Image Models (Re-LAION-Caption 19M)
von: Merchant, Nicholas, et al.
Veröffentlicht: (2025)
von: Merchant, Nicholas, et al.
Veröffentlicht: (2025)
Linear Alignment of Vision-language Models for Image Captioning
von: Paischer, Fabian, et al.
Veröffentlicht: (2023)
von: Paischer, Fabian, et al.
Veröffentlicht: (2023)
Optimizing GPT for Video Understanding: Zero-Shot Performance and Prompt Engineering
von: Beliaev, Mark, et al.
Veröffentlicht: (2025)
von: Beliaev, Mark, et al.
Veröffentlicht: (2025)
IFCap: Image-like Retrieval and Frequency-based Entity Filtering for Zero-shot Captioning
von: Lee, Soeun, et al.
Veröffentlicht: (2024)
von: Lee, Soeun, et al.
Veröffentlicht: (2024)
From Pixels to Prose: A Large Dataset of Dense Image Captions
von: Singla, Vasu, et al.
Veröffentlicht: (2024)
von: Singla, Vasu, et al.
Veröffentlicht: (2024)
Improving Image Captioning Descriptiveness by Ranking and LLM-based Fusion
von: Celona, Luigi, et al.
Veröffentlicht: (2023)
von: Celona, Luigi, et al.
Veröffentlicht: (2023)
FigCaps-HF: A Figure-to-Caption Generative Framework and Benchmark with Human Feedback
von: Singh, Ashish, et al.
Veröffentlicht: (2023)
von: Singh, Ashish, et al.
Veröffentlicht: (2023)
D3G: Diverse Demographic Data Generation Increases Zero-Shot Image Classification Accuracy within Multimodal Models
von: Hickmon, Javon
Veröffentlicht: (2025)
von: Hickmon, Javon
Veröffentlicht: (2025)
Brazilian Portuguese Image Captioning with Transformers: A Study on Cross-Native-Translated Dataset
von: Bromonschenkel, Gabriel, et al.
Veröffentlicht: (2026)
von: Bromonschenkel, Gabriel, et al.
Veröffentlicht: (2026)
Troika: Multi-Path Cross-Modal Traction for Compositional Zero-Shot Learning
von: Huang, Siteng, et al.
Veröffentlicht: (2023)
von: Huang, Siteng, et al.
Veröffentlicht: (2023)
Verbalized Representation Learning for Interpretable Few-Shot Generalization
von: Yang, Cheng-Fu, et al.
Veröffentlicht: (2024)
von: Yang, Cheng-Fu, et al.
Veröffentlicht: (2024)
CAPability: A Comprehensive Visual Caption Benchmark for Evaluating Both Correctness and Thoroughness
von: Liu, Zhihang, et al.
Veröffentlicht: (2025)
von: Liu, Zhihang, et al.
Veröffentlicht: (2025)
RadZero: Similarity-Based Cross-Attention for Explainable Vision-Language Alignment in Chest X-ray with Zero-Shot Multi-Task Capability
von: Park, Jonggwon, et al.
Veröffentlicht: (2025)
von: Park, Jonggwon, et al.
Veröffentlicht: (2025)
SketchDNN: Joint Continuous-Discrete Diffusion for CAD Sketch Generation
von: Chereddy, Sathvik, et al.
Veröffentlicht: (2025)
von: Chereddy, Sathvik, et al.
Veröffentlicht: (2025)
No "Zero-Shot" Without Exponential Data: Pretraining Concept Frequency Determines Multimodal Model Performance
von: Udandarao, Vishaal, et al.
Veröffentlicht: (2024)
von: Udandarao, Vishaal, et al.
Veröffentlicht: (2024)
Exploring the Limits of Zero Shot Vision Language Models for Hate Meme Detection: The Vulnerabilities and their Interpretations
von: Rizwan, Naquee, et al.
Veröffentlicht: (2024)
von: Rizwan, Naquee, et al.
Veröffentlicht: (2024)
Zero-Shot Defense Against Toxic Images via Inherent Multimodal Alignment in LVLMs
von: Zhao, Wei, et al.
Veröffentlicht: (2025)
von: Zhao, Wei, et al.
Veröffentlicht: (2025)
Transformer with Controlled Attention for Synchronous Motion Captioning
von: Radouane, Karim, et al.
Veröffentlicht: (2024)
von: Radouane, Karim, et al.
Veröffentlicht: (2024)
Wolf: Dense Video Captioning with a World Summarization Framework
von: Li, Boyi, et al.
Veröffentlicht: (2024)
von: Li, Boyi, et al.
Veröffentlicht: (2024)
Multimodal Arabic Captioning with Interpretable Visual Concept Integration
von: Elchafei, Passant, et al.
Veröffentlicht: (2025)
von: Elchafei, Passant, et al.
Veröffentlicht: (2025)
BioCAP: Exploiting Synthetic Captions Beyond Labels in Biological Foundation Models
von: Zhang, Ziheng, et al.
Veröffentlicht: (2025)
von: Zhang, Ziheng, et al.
Veröffentlicht: (2025)
Zero-Shot Vehicle Model Recognition via Text-Based Retrieval-Augmented Generation
von: Chang, Wei-Chia, et al.
Veröffentlicht: (2025)
von: Chang, Wei-Chia, et al.
Veröffentlicht: (2025)
FilterRAG: Zero-Shot Informed Retrieval-Augmented Generation to Mitigate Hallucinations in VQA
von: Sarwar, Nobin
Veröffentlicht: (2025)
von: Sarwar, Nobin
Veröffentlicht: (2025)
The Solution for the 5th GCAIAC Zero-shot Referring Expression Comprehension Challenge
von: Huang, Longfei, et al.
Veröffentlicht: (2024)
von: Huang, Longfei, et al.
Veröffentlicht: (2024)
LLM as a Neural Architect: Controlled Generation of Image Captioning Models Under Strict API Contracts
von: Jesani, Krunal, et al.
Veröffentlicht: (2025)
von: Jesani, Krunal, et al.
Veröffentlicht: (2025)
Energy-based Models are Zero-Shot Planners for Compositional Scene Rearrangement
von: Gkanatsios, Nikolaos, et al.
Veröffentlicht: (2023)
von: Gkanatsios, Nikolaos, et al.
Veröffentlicht: (2023)
An Examination of the Robustness of Reference-Free Image Captioning Evaluation Metrics
von: Ahmadi, Saba, et al.
Veröffentlicht: (2023)
von: Ahmadi, Saba, et al.
Veröffentlicht: (2023)
A Survey on Generative Modeling with Limited Data, Few Shots, and Zero Shot
von: Abdollahzadeh, Milad, et al.
Veröffentlicht: (2023)
von: Abdollahzadeh, Milad, et al.
Veröffentlicht: (2023)
Commonsense for Zero-Shot Natural Language Video Localization
von: Holla, Meghana, et al.
Veröffentlicht: (2023)
von: Holla, Meghana, et al.
Veröffentlicht: (2023)
CIC-BART-SSA: Controllable Image Captioning with Structured Semantic Augmentation
von: Basioti, Kalliopi, et al.
Veröffentlicht: (2024)
von: Basioti, Kalliopi, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Multimodal Unsupervised Domain Generalization by Retrieving Across the Modality Gap
von: Liao, Christopher, et al.
Veröffentlicht: (2024) -
Descriptor and Word Soups: Overcoming the Parameter Efficiency Accuracy Tradeoff for Out-of-Distribution Few-shot Learning
von: Liao, Christopher, et al.
Veröffentlicht: (2023) -
Unleashing Text-to-Image Diffusion Prior for Zero-Shot Image Captioning
von: Luo, Jianjie, et al.
Veröffentlicht: (2024) -
ERM++: An Improved Baseline for Domain Generalization
von: Teterwak, Piotr, et al.
Veröffentlicht: (2023) -
FocalLens: Instruction Tuning Enables Zero-Shot Conditional Image Representations
von: Hsieh, Cheng-Yu, et al.
Veröffentlicht: (2025)