A Picture is Worth More Than 77 Text Tokens: Evaluating CLIP-Style Models on Dense Captions
Fuente:
arXiv
Guardado en:
| Autores principales: | Urbanek, Jack, Bordes, Florian, Astolfi, Pietro, Williamson, Mary, Sharma, Vasu, Romero-Soriano, Adriana |
|---|---|
| Formato: | Preprint |
| Publicado: |
2023
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Object-centric Binding in Contrastive Language-Image Pretraining
por: Assouel, Rim, et al.
Publicado: (2025)
por: Assouel, Rim, et al.
Publicado: (2025)
Improving Text-to-Image Consistency via Automatic Prompt Optimization
por: Mañas, Oscar, et al.
Publicado: (2024)
por: Mañas, Oscar, et al.
Publicado: (2024)
Improving the Scaling Laws of Synthetic Data with Deliberate Practice
por: Askari-Hemmat, Reyhane, et al.
Publicado: (2025)
por: Askari-Hemmat, Reyhane, et al.
Publicado: (2025)
CLIP with Quality Captions: A Strong Pretraining for Vision Tasks
por: Vasu, Pavan Kumar Anasosalu, et al.
Publicado: (2024)
por: Vasu, Pavan Kumar Anasosalu, et al.
Publicado: (2024)
Feedback-guided Data Synthesis for Imbalanced Classification
por: Hemmat, Reyhane Askari, et al.
Publicado: (2023)
por: Hemmat, Reyhane Askari, et al.
Publicado: (2023)
Tokenization Is More Than Compression
por: Schmidt, Craig W., et al.
Publicado: (2024)
por: Schmidt, Craig W., et al.
Publicado: (2024)
DP-RDM: Adapting Diffusion Models to Private Domains Without Fine-Tuning
por: Lebensold, Jonathan, et al.
Publicado: (2024)
por: Lebensold, Jonathan, et al.
Publicado: (2024)
Parrot Captions Teach CLIP to Spot Text
por: Lin, Yiqi, et al.
Publicado: (2023)
por: Lin, Yiqi, et al.
Publicado: (2023)
Scaling Laws in Patchification: An Image Is Worth 50,176 Tokens And More
por: Wang, Feng, et al.
Publicado: (2025)
por: Wang, Feng, et al.
Publicado: (2025)
An Image is Worth More Than 16x16 Patches: Exploring Transformers on Individual Pixels
por: Nguyen, Duy-Kien, et al.
Publicado: (2024)
por: Nguyen, Duy-Kien, et al.
Publicado: (2024)
Consistency-diversity-realism Pareto fronts of conditional image generative models
por: Astolfi, Pietro, et al.
Publicado: (2024)
por: Astolfi, Pietro, et al.
Publicado: (2024)
Discrete Audio Tokens: More Than a Survey!
por: Mousavi, Pooneh, et al.
Publicado: (2025)
por: Mousavi, Pooneh, et al.
Publicado: (2025)
From Pixels to Prose: A Large Dataset of Dense Image Captions
por: Singla, Vasu, et al.
Publicado: (2024)
por: Singla, Vasu, et al.
Publicado: (2024)
VoxelCodeBench: Benchmarking 3D World Modeling Through Code Generation
por: Zheng, Yan, et al.
Publicado: (2026)
por: Zheng, Yan, et al.
Publicado: (2026)
A Text is Worth Several Tokens: Text Embedding from LLMs Secretly Aligns Well with The Key Tokens
por: Nie, Zhijie, et al.
Publicado: (2024)
por: Nie, Zhijie, et al.
Publicado: (2024)
StyleHumanCLIP: Text-guided Garment Manipulation for StyleGAN-Human
por: Yoshikawa, Takato, et al.
Publicado: (2023)
por: Yoshikawa, Takato, et al.
Publicado: (2023)
Captions Are Worth a Thousand Words: Enhancing Product Retrieval with Pretrained Image-to-Text Models
por: Tang, Jason, et al.
Publicado: (2024)
por: Tang, Jason, et al.
Publicado: (2024)
A Picture is Worth a Thousand (Correct) Captions: A Vision-Guided Judge-Corrector System for Multimodal Machine Translation
por: Betala, Siddharth, et al.
Publicado: (2025)
por: Betala, Siddharth, et al.
Publicado: (2025)
A Pixel Is Worth More Than One 3D Gaussians in Single-View 3D Reconstruction
por: Shen, Jianghao, et al.
Publicado: (2024)
por: Shen, Jianghao, et al.
Publicado: (2024)
Distill CLIP (DCLIP): Enhancing Image-Text Retrieval via Cross-Modal Transformer Distillation
por: Csizmadia, Daniel, et al.
Publicado: (2025)
por: Csizmadia, Daniel, et al.
Publicado: (2025)
Exploring How Generative MLLMs Perceive More Than CLIP with the Same Vision Encoder
por: Li, Siting, et al.
Publicado: (2024)
por: Li, Siting, et al.
Publicado: (2024)
Boosting Latent Diffusion with Perceptual Objectives
por: Berrada, Tariq, et al.
Publicado: (2024)
por: Berrada, Tariq, et al.
Publicado: (2024)
iLLaVA: An Image is Worth Fewer Than 1/3 Input Tokens in Large Multimodal Models
por: Hu, Lianyu, et al.
Publicado: (2024)
por: Hu, Lianyu, et al.
Publicado: (2024)
MoreHopQA: More Than Multi-hop Reasoning
por: Schnitzler, Julian, et al.
Publicado: (2024)
por: Schnitzler, Julian, et al.
Publicado: (2024)
Ep. 305: Is Your Typing Style More Secure Than Your Password?
por: Rosehill, Daniel, et al.
Publicado: (2026)
por: Rosehill, Daniel, et al.
Publicado: (2026)
More Than Efficiency: Embedding Compression Improves Domain Adaptation in Dense Retrieval
por: Zuo, Chunsheng, et al.
Publicado: (2026)
por: Zuo, Chunsheng, et al.
Publicado: (2026)
Words Worth a Thousand Pictures: Measuring and Understanding Perceptual Variability in Text-to-Image Generation
por: Tang, Raphael, et al.
Publicado: (2024)
por: Tang, Raphael, et al.
Publicado: (2024)
Dense Motion Captioning
por: Xu, Shiyao, et al.
Publicado: (2025)
por: Xu, Shiyao, et al.
Publicado: (2025)
MobileCLIP: Fast Image-Text Models through Multi-Modal Reinforced Training
por: Vasu, Pavan Kumar Anasosalu, et al.
Publicado: (2023)
por: Vasu, Pavan Kumar Anasosalu, et al.
Publicado: (2023)
A LoRA is Worth a Thousand Pictures
por: Liu, Chenxi, et al.
Publicado: (2024)
por: Liu, Chenxi, et al.
Publicado: (2024)
CapSpeech: Enabling Downstream Applications in Style-Captioned Text-to-Speech
por: Wang, Helin, et al.
Publicado: (2025)
por: Wang, Helin, et al.
Publicado: (2025)
SPECS: Specificity-Enhanced CLIP-Score for Long Image Caption Evaluation
por: Chen, Xiaofu, et al.
Publicado: (2025)
por: Chen, Xiaofu, et al.
Publicado: (2025)
AVCap: Leveraging Audio-Visual Features as Text Tokens for Captioning
por: Kim, Jongsuk, et al.
Publicado: (2024)
por: Kim, Jongsuk, et al.
Publicado: (2024)
Towards Geographic Inclusion in the Evaluation of Text-to-Image Models
por: Hall, Melissa, et al.
Publicado: (2024)
por: Hall, Melissa, et al.
Publicado: (2024)
Film & Culture: Introduction. The Human/More‐Than‐Human Relationship
por: Laura Camille Tuley, et al.
Publicado: (2024)
por: Laura Camille Tuley, et al.
Publicado: (2024)
VeCLIP: Improving CLIP Training via Visual-enriched Captions
por: Lai, Zhengfeng, et al.
Publicado: (2023)
por: Lai, Zhengfeng, et al.
Publicado: (2023)
Semantic Token Reweighting for Interpretable and Controllable Text Embeddings in CLIP
por: Kim, Eunji, et al.
Publicado: (2024)
por: Kim, Eunji, et al.
Publicado: (2024)
MiSCHiEF: A Benchmark in Minimal-Pairs of Safety and Culture for Holistic Evaluation of Fine-Grained Image-Caption Alignment
por: Banerjee, Sagarika, et al.
Publicado: (2026)
por: Banerjee, Sagarika, et al.
Publicado: (2026)
Streaming Dense Video Captioning
por: Zhou, Xingyi, et al.
Publicado: (2024)
por: Zhou, Xingyi, et al.
Publicado: (2024)
A Style is Worth One Code: Unlocking Code-to-Style Image Generation with Discrete Style Space
por: Liu, Huijie, et al.
Publicado: (2025)
por: Liu, Huijie, et al.
Publicado: (2025)
Ejemplares similares
-
Object-centric Binding in Contrastive Language-Image Pretraining
por: Assouel, Rim, et al.
Publicado: (2025) -
Improving Text-to-Image Consistency via Automatic Prompt Optimization
por: Mañas, Oscar, et al.
Publicado: (2024) -
Improving the Scaling Laws of Synthetic Data with Deliberate Practice
por: Askari-Hemmat, Reyhane, et al.
Publicado: (2025) -
CLIP with Quality Captions: A Strong Pretraining for Vision Tasks
por: Vasu, Pavan Kumar Anasosalu, et al.
Publicado: (2024) -
Feedback-guided Data Synthesis for Imbalanced Classification
por: Hemmat, Reyhane Askari, et al.
Publicado: (2023)