LocCa: Visual Pretraining with Location-aware Captioners
Fuente:
arXiv
Saved in:
| Main Authors: | Wan, Bo, Tschannen, Michael, Xian, Yongqin, Pavetic, Filip, Alabdulmohsin, Ibrahim, Wang, Xiao, Pinto, André Susano, Steiner, Andreas, Beyer, Lucas, Zhai, Xiaohua |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Getting ViT in Shape: Scaling Laws for Compute-Optimal Model Design
by: Alabdulmohsin, Ibrahim, et al.
Published: (2023)
by: Alabdulmohsin, Ibrahim, et al.
Published: (2023)
Recursive Inference Scaling: A Winning Path to Scalable Inference in Language and Multimodal Systems
by: Alabdulmohsin, Ibrahim, et al.
Published: (2025)
by: Alabdulmohsin, Ibrahim, et al.
Published: (2025)
JetFormer: An Autoregressive Generative Model of Raw Images and Text
by: Tschannen, Michael, et al.
Published: (2024)
by: Tschannen, Michael, et al.
Published: (2024)
Jet: A Modern Transformer-Based Normalizing Flow
by: Kolesnikov, Alexander, et al.
Published: (2024)
by: Kolesnikov, Alexander, et al.
Published: (2024)
No Filter: Cultural and Socioeconomic Diversity in Contrastive Vision-Language Models
by: Pouget, Angéline, et al.
Published: (2024)
by: Pouget, Angéline, et al.
Published: (2024)
PaliGemma 2: A Family of Versatile VLMs for Transfer
by: Steiner, Andreas, et al.
Published: (2024)
by: Steiner, Andreas, et al.
Published: (2024)
A Tale of Two Structures: Do LLMs Capture the Fractal Complexity of Language?
by: Alabdulmohsin, Ibrahim, et al.
Published: (2025)
by: Alabdulmohsin, Ibrahim, et al.
Published: (2025)
SigLIP 2: Multilingual Vision-Language Encoders with Improved Semantic Understanding, Localization, and Dense Features
by: Tschannen, Michael, et al.
Published: (2025)
by: Tschannen, Michael, et al.
Published: (2025)
CLIP the Bias: How Useful is Balancing Data in Multimodal Learning?
by: Alabdulmohsin, Ibrahim, et al.
Published: (2024)
by: Alabdulmohsin, Ibrahim, et al.
Published: (2024)
Scaling Pre-training to One Hundred Billion Data for Vision Language Models
by: Wang, Xiao, et al.
Published: (2025)
by: Wang, Xiao, et al.
Published: (2025)
Ask2Loc: Learning to Locate Instructional Visual Answers by Asking Questions
by: Zong, Chang, et al.
Published: (2025)
by: Zong, Chang, et al.
Published: (2025)
Fractal Patterns May Illuminate the Success of Next-Token Prediction
by: Alabdulmohsin, Ibrahim, et al.
Published: (2024)
by: Alabdulmohsin, Ibrahim, et al.
Published: (2024)
PaliGemma: A versatile 3B VLM for transfer
by: Beyer, Lucas, et al.
Published: (2024)
by: Beyer, Lucas, et al.
Published: (2024)
ProLoc: Robust Location Proofs in Hindsight
by: De Viti, Roberta, et al.
Published: (2024)
by: De Viti, Roberta, et al.
Published: (2024)
Towards Truly Zero-shot Compositional Visual Reasoning with LLMs as Programmers
by: Stanić, Aleksandar, et al.
Published: (2024)
by: Stanić, Aleksandar, et al.
Published: (2024)
Toward a Diffusion-Based Generalist for Dense Vision Tasks
by: Fan, Yue, et al.
Published: (2024)
by: Fan, Yue, et al.
Published: (2024)
Modeling Caption Diversity in Contrastive Vision-Language Pretraining
by: Lavoie, Samuel, et al.
Published: (2024)
by: Lavoie, Samuel, et al.
Published: (2024)
Campo Experimental Forestal "San Juan Tetla", Pue
by: Susano Hernández, Roberto
Published: (1963)
by: Susano Hernández, Roberto
Published: (1963)
Principales plantas potencialmente toxicas para la ganadería de la zona forestal de Zoquiapan, Edo. de México
by: Susano Hernández, Roberto
Published: (1970)
by: Susano Hernández, Roberto
Published: (1970)
Especies arbóreas forestales susceptibles de aprovecharse como forraje
by: Susano Hernández, Roberto
Published: (1981)
by: Susano Hernández, Roberto
Published: (1981)
LocDiff: Identifying Locations on Earth by Diffusing in the Hilbert Space
by: Wang, Zhangyu, et al.
Published: (2025)
by: Wang, Zhangyu, et al.
Published: (2025)
NextLocLLM: Location Semantics Modeling and Coordinate-Based Next Location Prediction with LLMs
by: Liu, Shuai, et al.
Published: (2024)
by: Liu, Shuai, et al.
Published: (2024)
Location-Aware Pretraining for Medical Difference Visual Question Answering
by: Musinguzi, Denis, et al.
Published: (2026)
by: Musinguzi, Denis, et al.
Published: (2026)
A case of Conradi‐Hünermann‐Happle syndrome treated with topical simvastatin‐cholesterol ointment
by: Jean Zevallos, et al.
Published: (2024)
by: Jean Zevallos, et al.
Published: (2024)
Training-free Uncertainty Guidance for Complex Visual Tasks with MLLMs
by: Kim, Sanghwan, et al.
Published: (2025)
by: Kim, Sanghwan, et al.
Published: (2025)
NeuraLoc: Visual Localization in Neural Implicit Map with Dual Complementary Features
by: Zhai, Hongjia, et al.
Published: (2025)
by: Zhai, Hongjia, et al.
Published: (2025)
LocPoseNet: Robust Location Prior for Unseen Object Pose Estimation
by: Zhao, Chen, et al.
Published: (2022)
by: Zhao, Chen, et al.
Published: (2022)
Loc4Plan: Locating Before Planning for Outdoor Vision and Language Navigation
by: Tian, Huilin, et al.
Published: (2024)
by: Tian, Huilin, et al.
Published: (2024)
Collateral‐Based Monetary Policy: Evidence From China
by: Hanming Fang, et al.
Published: (2025)
by: Hanming Fang, et al.
Published: (2025)
SplatLoc: 3D Gaussian Splatting-based Visual Localization for Augmented Reality
by: Zhai, Hongjia, et al.
Published: (2024)
by: Zhai, Hongjia, et al.
Published: (2024)
LocInv: Localization-aware Inversion for Text-Guided Image Editing
by: Tang, Chuanming, et al.
Published: (2024)
by: Tang, Chuanming, et al.
Published: (2024)
Eco-WakeLoc: An Energy-Neutral and Cooperative UWB Real-Time Locating System
by: Cortesi, Silvano, et al.
Published: (2026)
by: Cortesi, Silvano, et al.
Published: (2026)
GIVT: Generative Infinite-Vocabulary Transformers
by: Tschannen, Michael, et al.
Published: (2023)
by: Tschannen, Michael, et al.
Published: (2023)
LocBAM: Advancing 3D Patch-Based Image Segmentation by Integrating Location Contex
by: Hooft, Donnate, et al.
Published: (2026)
by: Hooft, Donnate, et al.
Published: (2026)
NewsCaption: Named-Entity aware Captioning for Out-of-Context Media
by: Singh, Anurag, et al.
Published: (2024)
by: Singh, Anurag, et al.
Published: (2024)
Pretrained Image-Text Models are Secretly Video Captioners
by: Zhang, Chunhui, et al.
Published: (2025)
by: Zhang, Chunhui, et al.
Published: (2025)
OPCap:Object-aware Prompting Captioning
by: Huang, Feiyang
Published: (2024)
by: Huang, Feiyang
Published: (2024)
Implicit Location-Caption Alignment via Complementary Masking for Weakly-Supervised Dense Video Captioning
by: Ge, Shiping, et al.
Published: (2024)
by: Ge, Shiping, et al.
Published: (2024)
PKDB161058
by: Yongqin, W
Published: (2000)
by: Yongqin, W
Published: (2000)
BFT-PoLoc: A Byzantine Fortified Trigonometric Proof of Location Protocol using Internet Delays
by: Sheng, Peiyao, et al.
Published: (2024)
by: Sheng, Peiyao, et al.
Published: (2024)
Similar Items
-
Getting ViT in Shape: Scaling Laws for Compute-Optimal Model Design
by: Alabdulmohsin, Ibrahim, et al.
Published: (2023) -
Recursive Inference Scaling: A Winning Path to Scalable Inference in Language and Multimodal Systems
by: Alabdulmohsin, Ibrahim, et al.
Published: (2025) -
JetFormer: An Autoregressive Generative Model of Raw Images and Text
by: Tschannen, Michael, et al.
Published: (2024) -
Jet: A Modern Transformer-Based Normalizing Flow
by: Kolesnikov, Alexander, et al.
Published: (2024) -
No Filter: Cultural and Socioeconomic Diversity in Contrastive Vision-Language Models
by: Pouget, Angéline, et al.
Published: (2024)