SigLIP 2: Multilingual Vision-Language Encoders with Improved Semantic Understanding, Localization, and Dense Features
Fuente:
arXiv
Saved in:
| Main Authors: | Tschannen, Michael, Gritsenko, Alexey, Wang, Xiao, Naeem, Muhammad Ferjad, Alabdulmohsin, Ibrahim, Parthasarathy, Nikhil, Evans, Talfan, Beyer, Lucas, Xia, Ye, Mustafa, Basil, Hénaff, Olivier, Harmsen, Jeremiah, Steiner, Andreas, Zhai, Xiaohua |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Data curation via joint example selection further accelerates multimodal learning
by: Evans, Talfan, et al.
Published: (2024)
by: Evans, Talfan, et al.
Published: (2024)
Active Data Curation Effectively Distills Large-Scale Multimodal Models
by: Udandarao, Vishaal, et al.
Published: (2024)
by: Udandarao, Vishaal, et al.
Published: (2024)
ClearVision: Leveraging CycleGAN and SigLIP-2 for Robust All-Weather Classification in Traffic Camera Imagery
by: Sivaraman, Anush Lakshman, et al.
Published: (2025)
by: Sivaraman, Anush Lakshman, et al.
Published: (2025)
Getting ViT in Shape: Scaling Laws for Compute-Optimal Model Design
by: Alabdulmohsin, Ibrahim, et al.
Published: (2023)
by: Alabdulmohsin, Ibrahim, et al.
Published: (2023)
Recursive Inference Scaling: A Winning Path to Scalable Inference in Language and Multimodal Systems
by: Alabdulmohsin, Ibrahim, et al.
Published: (2025)
by: Alabdulmohsin, Ibrahim, et al.
Published: (2025)
LocCa: Visual Pretraining with Location-aware Captioners
by: Wan, Bo, et al.
Published: (2024)
by: Wan, Bo, et al.
Published: (2024)
Toward a Diffusion-Based Generalist for Dense Vision Tasks
by: Fan, Yue, et al.
Published: (2024)
by: Fan, Yue, et al.
Published: (2024)
No Filter: Cultural and Socioeconomic Diversity in Contrastive Vision-Language Models
by: Pouget, Angéline, et al.
Published: (2024)
by: Pouget, Angéline, et al.
Published: (2024)
PaliGemma 2: A Family of Versatile VLMs for Transfer
by: Steiner, Andreas, et al.
Published: (2024)
by: Steiner, Andreas, et al.
Published: (2024)
A Tale of Two Structures: Do LLMs Capture the Fractal Complexity of Language?
by: Alabdulmohsin, Ibrahim, et al.
Published: (2025)
by: Alabdulmohsin, Ibrahim, et al.
Published: (2025)
Layerwise complexity-matched learning yields an improved model of cortical area V2
by: Parthasarathy, Nikhil, et al.
Published: (2023)
by: Parthasarathy, Nikhil, et al.
Published: (2023)
Bad Students Make Great Teachers: Active Learning Accelerates Large-Scale Visual Understanding
by: Evans, Talfan, et al.
Published: (2023)
by: Evans, Talfan, et al.
Published: (2023)
CLIP the Bias: How Useful is Balancing Data in Multimodal Learning?
by: Alabdulmohsin, Ibrahim, et al.
Published: (2024)
by: Alabdulmohsin, Ibrahim, et al.
Published: (2024)
PaliGemma: A versatile 3B VLM for transfer
by: Beyer, Lucas, et al.
Published: (2024)
by: Beyer, Lucas, et al.
Published: (2024)
Self-supervised video pretraining yields robust and more human-aligned visual representations
by: Parthasarathy, Nikhil, et al.
Published: (2022)
by: Parthasarathy, Nikhil, et al.
Published: (2022)
Scaling Open-Vocabulary Object Detection
by: Minderer, Matthias, et al.
Published: (2023)
by: Minderer, Matthias, et al.
Published: (2023)
Prompt-Conditioned FiLM and Multi-Scale Fusion on MedSigLIP for Low-Dose CT Quality Assessment
by: Demiroglu, Tolga, et al.
Published: (2025)
by: Demiroglu, Tolga, et al.
Published: (2025)
Scaling Pre-training to One Hundred Billion Data for Vision Language Models
by: Wang, Xiao, et al.
Published: (2025)
by: Wang, Xiao, et al.
Published: (2025)
Language-Unlocked ViT (LUViT): Empowering Self-Supervised Vision Transformers with LLMs
by: Kuzucu, Selim, et al.
Published: (2025)
by: Kuzucu, Selim, et al.
Published: (2025)
Islam, Civil Society and Social Work
by: Harmsen, Egbert
Published: (2010)
by: Harmsen, Egbert
Published: (2010)
Time-, Memory- and Parameter-Efficient Visual Adaptation
by: Mercea, Otniel-Bogdan, et al.
Published: (2024)
by: Mercea, Otniel-Bogdan, et al.
Published: (2024)
Modular differential equations of minimal orders of the elliptic genus of Calabi--Yau varieties
by: Adler, Dmitrii, et al.
Published: (2025)
by: Adler, Dmitrii, et al.
Published: (2025)
Learning to Prompt with Text Only Supervision for Vision-Language Models
by: Khattak, Muhammad Uzair, et al.
Published: (2024)
by: Khattak, Muhammad Uzair, et al.
Published: (2024)
PARCEL: Pool-Anchored Resampling with Conditioned Elastic Queries for Efficient Vision-Language Understanding
by: Kuzucu, Selim, et al.
Published: (2026)
by: Kuzucu, Selim, et al.
Published: (2026)
LiteFrame: Efficient Vision Encoders Unlock Frame Scaling in Video LLMs
by: Kim, Jihwan, et al.
Published: (2026)
by: Kim, Jihwan, et al.
Published: (2026)
DIFFERENCE IN PERCEPTION OF LIP LINE CANTING: A COMPARATIVE EVALUATION AMONG ORTHODONTISTS, GENERAL DENTISTS AND PATIENTS
by: Dr Lubna Amjad,Dr Abdullah Jan,Dr Erum Amin,Dr Sabeen Mustafa
Published: (2026)
by: Dr Lubna Amjad,Dr Abdullah Jan,Dr Erum Amin,Dr Sabeen Mustafa
Published: (2026)
SigXTalk software
by: Hou, Jiawen
Published: (2025)
by: Hou, Jiawen
Published: (2025)
GIVT: Generative Infinite-Vocabulary Transformers
by: Tschannen, Michael, et al.
Published: (2023)
by: Tschannen, Michael, et al.
Published: (2023)
Towards Truly Zero-shot Compositional Visual Reasoning with LLMs as Programmers
by: Stanić, Aleksandar, et al.
Published: (2024)
by: Stanić, Aleksandar, et al.
Published: (2024)
Development of 16S rRNA-Based Probes for theCoriobacterium Group and the Atopobium Cluster and Their Application for Enumeration of Coriobacteriaceaein Human Feces from Volunteers of Different Age Groups. / Hermie J. M. Hermsen
by: Harmsen, Hermie J. M
Published: (2000)
by: Harmsen, Hermie J. M
Published: (2000)
Rethinking Digitalization and Climate: Don't Predict, Mitigate
by: Gritsenko, Daria, et al.
Published: (2024)
by: Gritsenko, Daria, et al.
Published: (2024)
Multilingual Sentence-T5: Scalable Sentence Encoders for Multilingual Applications
by: Yano, Chihiro, et al.
Published: (2024)
by: Yano, Chihiro, et al.
Published: (2024)
RefAM: Attention Magnets for Zero-Shot Referral Segmentation
by: Kukleva, Anna, et al.
Published: (2025)
by: Kukleva, Anna, et al.
Published: (2025)
FACIAL EXTRACTION AND LIP TRACKING USING FACIAL POINTS
by: 'Adawiyyah, Rabi'atul
Published: (2026)
by: 'Adawiyyah, Rabi'atul
Published: (2026)
UNIFIED LATENT INTELLIGENCE PHILOSOPHY (LIP) FRAMEWORK (ULF)
by: Nalam, Allan John
Published: (2025)
by: Nalam, Allan John
Published: (2025)
Use the LIP Model to Identify Top‐Tier Prospects
Published: (2024)
Published: (2024)
Fractal Patterns May Illuminate the Success of Next-Token Prediction
by: Alabdulmohsin, Ibrahim, et al.
Published: (2024)
by: Alabdulmohsin, Ibrahim, et al.
Published: (2024)
JetFormer: An Autoregressive Generative Model of Raw Images and Text
by: Tschannen, Michael, et al.
Published: (2024)
by: Tschannen, Michael, et al.
Published: (2024)
Jet: A Modern Transformer-Based Normalizing Flow
by: Kolesnikov, Alexander, et al.
Published: (2024)
by: Kolesnikov, Alexander, et al.
Published: (2024)
Notes on forest succession in old fields in southeastern Ontario: the woody species
by: Crowder, A. A., et al.
Published: (1998)
by: Crowder, A. A., et al.
Published: (1998)
Similar Items
-
Data curation via joint example selection further accelerates multimodal learning
by: Evans, Talfan, et al.
Published: (2024) -
Active Data Curation Effectively Distills Large-Scale Multimodal Models
by: Udandarao, Vishaal, et al.
Published: (2024) -
ClearVision: Leveraging CycleGAN and SigLIP-2 for Robust All-Weather Classification in Traffic Camera Imagery
by: Sivaraman, Anush Lakshman, et al.
Published: (2025) -
Getting ViT in Shape: Scaling Laws for Compute-Optimal Model Design
by: Alabdulmohsin, Ibrahim, et al.
Published: (2023) -
Recursive Inference Scaling: A Winning Path to Scalable Inference in Language and Multimodal Systems
by: Alabdulmohsin, Ibrahim, et al.
Published: (2025)