Updating CLIP to Prefer Descriptions Over Captions
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Zur, Amir, Kreiss, Elisa, D'Oosterlinck, Karel, Potts, Christopher, Geiger, Atticus |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
HyperDAS: Towards Automating Mechanistic Interpretability with Hypernetworks
von: Sun, Jiuding, et al.
Veröffentlicht: (2025)
von: Sun, Jiuding, et al.
Veröffentlicht: (2025)
In-Context Learning for Extreme Multi-Label Classification
von: D'Oosterlinck, Karel, et al.
Veröffentlicht: (2024)
von: D'Oosterlinck, Karel, et al.
Veröffentlicht: (2024)
CommVQA: Situating Visual Question Answering in Communicative Contexts
von: Naik, Nandita Shankar, et al.
Veröffentlicht: (2024)
von: Naik, Nandita Shankar, et al.
Veröffentlicht: (2024)
Anchored Preference Optimization and Contrastive Revisions: Addressing Underspecification in Alignment
von: D'Oosterlinck, Karel, et al.
Veröffentlicht: (2024)
von: D'Oosterlinck, Karel, et al.
Veröffentlicht: (2024)
Explaining Caption-Image Interactions in CLIP Models with Second-Order Attributions
von: Möller, Lucas, et al.
Veröffentlicht: (2024)
von: Möller, Lucas, et al.
Veröffentlicht: (2024)
Revisiting Image Captioning Training Paradigm via Direct CLIP-based Optimization
von: Moratelli, Nicholas, et al.
Veröffentlicht: (2024)
von: Moratelli, Nicholas, et al.
Veröffentlicht: (2024)
Are language models aware of the road not taken? Token-level uncertainty and hidden state dynamics
von: Zur, Amir, et al.
Veröffentlicht: (2025)
von: Zur, Amir, et al.
Veröffentlicht: (2025)
Bifrost-1: Bridging Multimodal LLMs and Diffusion Models with Patch-level CLIP Latents
von: Lin, Han, et al.
Veröffentlicht: (2025)
von: Lin, Han, et al.
Veröffentlicht: (2025)
SynPO: Synergizing Descriptiveness and Preference Optimization for Video Detailed Captioning
von: Dang, Jisheng, et al.
Veröffentlicht: (2025)
von: Dang, Jisheng, et al.
Veröffentlicht: (2025)
LatteCLIP: Unsupervised CLIP Fine-Tuning via LMM-Synthetic Texts
von: Cao, Anh-Quan, et al.
Veröffentlicht: (2024)
von: Cao, Anh-Quan, et al.
Veröffentlicht: (2024)
Parrot Captions Teach CLIP to Spot Text
von: Lin, Yiqi, et al.
Veröffentlicht: (2023)
von: Lin, Yiqi, et al.
Veröffentlicht: (2023)
InterCLIP-MEP: Interactive CLIP and Memory-Enhanced Predictor for Multi-modal Sarcasm Detection
von: Chen, Junjie, et al.
Veröffentlicht: (2024)
von: Chen, Junjie, et al.
Veröffentlicht: (2024)
Unveiling the Invisible: Captioning Videos with Metaphors
von: Kalarani, Abisek Rajakumar, et al.
Veröffentlicht: (2024)
von: Kalarani, Abisek Rajakumar, et al.
Veröffentlicht: (2024)
Text-only Synthesis for Image Captioning
von: Zhou, Qing, et al.
Veröffentlicht: (2024)
von: Zhou, Qing, et al.
Veröffentlicht: (2024)
The Role of Data Curation in Image Captioning
von: Li, Wenyan, et al.
Veröffentlicht: (2023)
von: Li, Wenyan, et al.
Veröffentlicht: (2023)
CHART-6: Human-Centered Evaluation of Data Visualization Understanding in Vision-Language Models
von: Verma, Arnav, et al.
Veröffentlicht: (2025)
von: Verma, Arnav, et al.
Veröffentlicht: (2025)
Improving Image Captioning Descriptiveness by Ranking and LLM-based Fusion
von: Celona, Luigi, et al.
Veröffentlicht: (2023)
von: Celona, Luigi, et al.
Veröffentlicht: (2023)
Retrieval-Augmented Long-Context Translation for Cultural Image Captioning: Gators submission for AmericasNLP 2026 shared task
von: Dhawan, Aashish, et al.
Veröffentlicht: (2026)
von: Dhawan, Aashish, et al.
Veröffentlicht: (2026)
CAPEEN: Image Captioning with Early Exits and Knowledge Distillation
von: Bajpai, Divya Jyoti, et al.
Veröffentlicht: (2024)
von: Bajpai, Divya Jyoti, et al.
Veröffentlicht: (2024)
ChartCap: Mitigating Hallucination of Dense Chart Captioning
von: Lim, Junyoung, et al.
Veröffentlicht: (2025)
von: Lim, Junyoung, et al.
Veröffentlicht: (2025)
CIC: A Framework for Culturally-Aware Image Captioning
von: Yun, Youngsik, et al.
Veröffentlicht: (2024)
von: Yun, Youngsik, et al.
Veröffentlicht: (2024)
Imagine How To Change: Explicit Procedure Modeling for Change Captioning
von: Sun, Jiayang, et al.
Veröffentlicht: (2026)
von: Sun, Jiayang, et al.
Veröffentlicht: (2026)
CapGeo: A Caption-Assisted Approach to Geometric Reasoning
von: Li, Yuying, et al.
Veröffentlicht: (2025)
von: Li, Yuying, et al.
Veröffentlicht: (2025)
EXPERT: An Explainable Image Captioning Evaluation Metric with Structured Explanations
von: Kim, Hyunjong, et al.
Veröffentlicht: (2025)
von: Kim, Hyunjong, et al.
Veröffentlicht: (2025)
Diagnosing Bottlenecks in Data Visualization Understanding by Vision-Language Models
von: Tartaglini, Alexa R., et al.
Veröffentlicht: (2025)
von: Tartaglini, Alexa R., et al.
Veröffentlicht: (2025)
No Captions, No Problem: Captionless 3D-CLIP Alignment with Hard Negatives via CLIP Knowledge and LLMs
von: Sbrolli, Cristian, et al.
Veröffentlicht: (2024)
von: Sbrolli, Cristian, et al.
Veröffentlicht: (2024)
ComCLIP: Training-Free Compositional Image and Text Matching
von: Jiang, Kenan, et al.
Veröffentlicht: (2022)
von: Jiang, Kenan, et al.
Veröffentlicht: (2022)
MAMI: Multi-Attentional Mutual-Information for Long Sequence Neuron Captioning
von: Fauzulhaq, Alfirsa Damasyifa, et al.
Veröffentlicht: (2024)
von: Fauzulhaq, Alfirsa Damasyifa, et al.
Veröffentlicht: (2024)
Image Captioning Evaluation in the Age of Multimodal LLMs: Challenges and Future Perspectives
von: Sarto, Sara, et al.
Veröffentlicht: (2025)
von: Sarto, Sara, et al.
Veröffentlicht: (2025)
The Power of Many: Multi-Agent Multimodal Models for Cultural Image Captioning
von: Bai, Longju, et al.
Veröffentlicht: (2024)
von: Bai, Longju, et al.
Veröffentlicht: (2024)
DENEB: A Hallucination-Robust Automatic Evaluation Metric for Image Captioning
von: Matsuda, Kazuki, et al.
Veröffentlicht: (2024)
von: Matsuda, Kazuki, et al.
Veröffentlicht: (2024)
Decoding fMRI Data into Captions using Prefix Language Modeling
von: Shen, Vyacheslav, et al.
Veröffentlicht: (2025)
von: Shen, Vyacheslav, et al.
Veröffentlicht: (2025)
From Pixels to Posts: Retrieval-Augmented Fashion Captioning and Hashtag Generation
von: Gondal, Moazzam Umer, et al.
Veröffentlicht: (2025)
von: Gondal, Moazzam Umer, et al.
Veröffentlicht: (2025)
VC4VG: Optimizing Video Captions for Text-to-Video Generation
von: Du, Yang, et al.
Veröffentlicht: (2025)
von: Du, Yang, et al.
Veröffentlicht: (2025)
Polos: Multimodal Metric Learning from Human Feedback for Image Captioning
von: Wada, Yuiga, et al.
Veröffentlicht: (2024)
von: Wada, Yuiga, et al.
Veröffentlicht: (2024)
Figuring out Figures: Using Textual References to Caption Scientific Figures
von: Cao, Stanley, et al.
Veröffentlicht: (2024)
von: Cao, Stanley, et al.
Veröffentlicht: (2024)
Does CLIP Bind Concepts? Probing Compositionality in Large Image Models
von: Lewis, Martha, et al.
Veröffentlicht: (2022)
von: Lewis, Martha, et al.
Veröffentlicht: (2022)
TokLIP: Marry Visual Tokens to CLIP for Multimodal Comprehension and Generation
von: Lin, Haokun, et al.
Veröffentlicht: (2025)
von: Lin, Haokun, et al.
Veröffentlicht: (2025)
CapRL: Stimulating Dense Image Caption Capabilities via Reinforcement Learning
von: Xing, Long, et al.
Veröffentlicht: (2025)
von: Xing, Long, et al.
Veröffentlicht: (2025)
LaMP-Cap: Personalized Figure Caption Generation With Multimodal Figure Profiles
von: Ng, Ho Yin 'Sam', et al.
Veröffentlicht: (2025)
von: Ng, Ho Yin 'Sam', et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
HyperDAS: Towards Automating Mechanistic Interpretability with Hypernetworks
von: Sun, Jiuding, et al.
Veröffentlicht: (2025) -
In-Context Learning for Extreme Multi-Label Classification
von: D'Oosterlinck, Karel, et al.
Veröffentlicht: (2024) -
CommVQA: Situating Visual Question Answering in Communicative Contexts
von: Naik, Nandita Shankar, et al.
Veröffentlicht: (2024) -
Anchored Preference Optimization and Contrastive Revisions: Addressing Underspecification in Alignment
von: D'Oosterlinck, Karel, et al.
Veröffentlicht: (2024) -
Explaining Caption-Image Interactions in CLIP Models with Second-Order Attributions
von: Möller, Lucas, et al.
Veröffentlicht: (2024)