Improving fine-grained understanding in image-text pre-training
Fuente:
arXiv
Saved in:
| Main Authors: | Bica, Ioana, Ilić, Anastasija, Bauer, Matthias, Erdogan, Goker, Bošnjak, Matko, Kaplanis, Christos, Gritsenko, Alexey A., Minderer, Matthias, Blundell, Charles, Pascanu, Razvan, Mitrović, Jovana |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Synth$^2$: Boosting Visual-Language Models with Synthetic Captions and Image Embeddings
by: Sharifzadeh, Sahand, et al.
Published: (2024)
by: Sharifzadeh, Sahand, et al.
Published: (2024)
Scaling Open-Vocabulary Object Detection
by: Minderer, Matthias, et al.
Published: (2023)
by: Minderer, Matthias, et al.
Published: (2023)
SemPPL: Predicting pseudo-labels for better contrastive representations
by: Bošnjak, Matko, et al.
Published: (2023)
by: Bošnjak, Matko, et al.
Published: (2023)
Softmax is not Enough (for Sharp Size Generalisation)
by: Veličković, Petar, et al.
Published: (2024)
by: Veličković, Petar, et al.
Published: (2024)
PaliGemma: A versatile 3B VLM for transfer
by: Beyer, Lucas, et al.
Published: (2024)
by: Beyer, Lucas, et al.
Published: (2024)
Round and Round We Go! What makes Rotary Positional Encodings useful?
by: Barbero, Federico, et al.
Published: (2024)
by: Barbero, Federico, et al.
Published: (2024)
Perplexity Cannot Always Tell Right from Wrong
by: Veličković, Petar, et al.
Published: (2026)
by: Veličković, Petar, et al.
Published: (2026)
Scene-Graph ViT: End-to-End Open-Vocabulary Visual Relationship Detection
by: Salzmann, Tim, et al.
Published: (2024)
by: Salzmann, Tim, et al.
Published: (2024)
Lattice: Learning to Efficiently Compress the Memory
by: Karami, Mahdi, et al.
Published: (2025)
by: Karami, Mahdi, et al.
Published: (2025)
Deep Grokking: Would Deep Neural Networks Generalize Better?
by: Fan, Simin, et al.
Published: (2024)
by: Fan, Simin, et al.
Published: (2024)
Filter Equivariant Functions: A symmetric account of length-general extrapolation on lists
by: Lewis, Owen, et al.
Published: (2025)
by: Lewis, Owen, et al.
Published: (2025)
Escuela de Administración, Pontificia Universidad Católica de Chile, 1994 2000
by: Matko Koljatic
Published: (2001)
by: Matko Koljatic
Published: (2001)
Vision-Language Model Dialog Games for Self-Improvement
by: Konyushkova, Ksenia, et al.
Published: (2025)
by: Konyushkova, Ksenia, et al.
Published: (2025)
NoProp: Training Neural Networks without Full Back-propagation or Full Forward-propagation
by: Li, Qinyu, et al.
Published: (2025)
by: Li, Qinyu, et al.
Published: (2025)
Meta-learning how to Share Credit among Macro-Actions
by: Hosu, Ionel-Alexandru, et al.
Published: (2025)
by: Hosu, Ionel-Alexandru, et al.
Published: (2025)
Latent Space Representations of Neural Algorithmic Reasoners
by: Mirjanić, Vladimir V., et al.
Published: (2023)
by: Mirjanić, Vladimir V., et al.
Published: (2023)
Revisiting Adam for Streaming Reinforcement Learning
by: Gogianu, Florin, et al.
Published: (2026)
by: Gogianu, Florin, et al.
Published: (2026)
The devil is in the fine-grained details: Evaluating open-vocabulary object detectors for fine-grained understanding
by: Bianchi, Lorenzo, et al.
Published: (2023)
by: Bianchi, Lorenzo, et al.
Published: (2023)
Computer Aided Assessment of Certain Neuromuscular Disorders Based on Surface Electromyography
by: Kaplanis, P. A.
Published: (2004)
by: Kaplanis, P. A.
Published: (2004)
Communicative patterns in Romanian workplace written texts
by: Razvan Saftoiu
Published: (2010)
by: Razvan Saftoiu
Published: (2010)
Building on Efficient Foundations: Effectively Training LLMs with Structured Feedforward Layers
by: Wei, Xiuying, et al.
Published: (2024)
by: Wei, Xiuying, et al.
Published: (2024)
What Can Grokking Teach Us About Learning Under Nonstationarity?
by: Lyle, Clare, et al.
Published: (2025)
by: Lyle, Clare, et al.
Published: (2025)
Investigating Low-Rank Training in Transformer Language Models: Efficiency and Scaling Analysis
by: Wei, Xiuying, et al.
Published: (2024)
by: Wei, Xiuying, et al.
Published: (2024)
RAT: Bridging RNN Efficiency and Attention Accuracy via Chunk-based Sequence Modeling
by: Wei, Xiuying, et al.
Published: (2025)
by: Wei, Xiuying, et al.
Published: (2025)
Why do LLMs attend to the first token?
by: Barbero, Federico, et al.
Published: (2025)
by: Barbero, Federico, et al.
Published: (2025)
Asynchronous Algorithmic Alignment with Cocycles
by: Dudzik, Andrew, et al.
Published: (2023)
by: Dudzik, Andrew, et al.
Published: (2023)
Time-, Memory- and Parameter-Efficient Visual Adaptation
by: Mercea, Otniel-Bogdan, et al.
Published: (2024)
by: Mercea, Otniel-Bogdan, et al.
Published: (2024)
Modular differential equations of minimal orders of the elliptic genus of Calabi--Yau varieties
by: Adler, Dmitrii, et al.
Published: (2025)
by: Adler, Dmitrii, et al.
Published: (2025)
Epidurinis skausmo malšinimas akušerijoje: gimdyvių savijauta, informacijos šaltiniai ir noras kito gimdymo metu keisti analgezijos būdą
by: Anastasija Bašarinaitė
Published: (2015)
by: Anastasija Bašarinaitė
Published: (2015)
Semilinear wave equations with time-dependent coefficients
by: Antonić, Nenad, et al.
Published: (2026)
by: Antonić, Nenad, et al.
Published: (2026)
Learning Successor Features the Simple Way
by: Chua, Raymond, et al.
Published: (2024)
by: Chua, Raymond, et al.
Published: (2024)
ACAVCaps: Enabling large-scale training for fine-grained and diverse audio understanding
by: Niu, Yadong, et al.
Published: (2026)
by: Niu, Yadong, et al.
Published: (2026)
MS-SSM: A Multi-Scale State Space Model for Efficient Sequence Modeling
by: Karami, Mahdi, et al.
Published: (2025)
by: Karami, Mahdi, et al.
Published: (2025)
No Representation, No Trust: Connecting Representation, Collapse, and Trust Issues in PPO
by: Moalla, Skander, et al.
Published: (2024)
by: Moalla, Skander, et al.
Published: (2024)
Attention as a Hypernetwork
by: Schug, Simon, et al.
Published: (2024)
by: Schug, Simon, et al.
Published: (2024)
Understanding Performance Gap Between Parallel and Sequential Sampling in Large Reasoning Models
by: Gu, Xiangming, et al.
Published: (2026)
by: Gu, Xiangming, et al.
Published: (2026)
FiCo-ITR: bridging fine-grained and coarse-grained image-text retrieval for comparative performance analysis
by: Williams-Lekuona, Mikel, et al.
Published: (2024)
by: Williams-Lekuona, Mikel, et al.
Published: (2024)
Economising Early Prints on Fight Books by Multiple Using Movable Half Page Woodcuts. Insights into the layout work on the illustrations of Andre Paurnfeindt’s Fight Book of 1516 published by Hieronymus Vietor
by: Matthias Johannes Bauer
Published: (2016)
by: Matthias Johannes Bauer
Published: (2016)
The Fight Book of Hugold Behr: A Late Sixteenth-Century Fight Book in Comparative Perspective
by: Matthias Johannes Bauer
Published: (2020)
by: Matthias Johannes Bauer
Published: (2020)
ТЕОРИЈА ДЕФИНИЦИОНОГ ОКВИРА НУЖНОСТИ (THEORY OF DEFINITIONAL FRAMEWORK OF NECESSITY)
by: Bošnjak, Ranko
Published: (2026)
by: Bošnjak, Ranko
Published: (2026)
Similar Items
-
Synth$^2$: Boosting Visual-Language Models with Synthetic Captions and Image Embeddings
by: Sharifzadeh, Sahand, et al.
Published: (2024) -
Scaling Open-Vocabulary Object Detection
by: Minderer, Matthias, et al.
Published: (2023) -
SemPPL: Predicting pseudo-labels for better contrastive representations
by: Bošnjak, Matko, et al.
Published: (2023) -
Softmax is not Enough (for Sharp Size Generalisation)
by: Veličković, Petar, et al.
Published: (2024) -
PaliGemma: A versatile 3B VLM for transfer
by: Beyer, Lucas, et al.
Published: (2024)