TIPSv2: Advancing Vision-Language Pretraining with Enhanced Patch-Text Alignment
Fuente:
arXiv
Saved in:
| Main Authors: | Cao, Bingyi, Chen, Koert, Maninis, Kevis-Kokitsi, Chen, Kaifeng, Karpur, Arjun, Xia, Ye, Dua, Sahil, Dabral, Tanmaya, Han, Guangxing, Han, Bohyung, Ainslie, Joshua, Bewley, Alex, Jacob, Mithun, Wagner, René, Ramos, Washington, Choromanski, Krzysztof, Seyedhosseini, Mojtaba, Zhou, Howard, Araujo, André |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
TIPS: Text-Image Pretraining with Spatial awareness
by: Maninis, Kevis-Kokitsi, et al.
Published: (2024)
by: Maninis, Kevis-Kokitsi, et al.
Published: (2024)
OmniNOCS: A unified NOCS dataset and model for 3D lifting of 2D objects
by: Krishnan, Akshay, et al.
Published: (2024)
by: Krishnan, Akshay, et al.
Published: (2024)
EgoCast: Forecasting Egocentric Human Pose in the Wild
by: Escobar, Maria, et al.
Published: (2024)
by: Escobar, Maria, et al.
Published: (2024)
OmniGlue: Generalizable Feature Matching with Foundation Model Guidance
by: Jiang, Hanwen, et al.
Published: (2024)
by: Jiang, Hanwen, et al.
Published: (2024)
Probing the 3D Awareness of Visual Foundation Models
by: Banani, Mohamed El, et al.
Published: (2024)
by: Banani, Mohamed El, et al.
Published: (2024)
Re-evaluating Group Robustness via Adaptive Class-Specific Scaling
by: Seo, Seonguk, et al.
Published: (2024)
by: Seo, Seonguk, et al.
Published: (2024)
MCL-GAN: Generative Adversarial Networks with Multiple Specialized Discriminators
by: Choi, Jinyoung, et al.
Published: (2021)
by: Choi, Jinyoung, et al.
Published: (2021)
Global-to-Local or Local-to-Global? Enhancing Image Retrieval with Efficient Local Search and Effective Global Re-ranking
by: Aiger, Dror, et al.
Published: (2025)
by: Aiger, Dror, et al.
Published: (2025)
Linear Transformer Topological Masking with Graph Random Features
by: Reid, Isaac, et al.
Published: (2024)
by: Reid, Isaac, et al.
Published: (2024)
LFM-3D: Learnable Feature Matching Across Wide Baselines Using 3D Signals
by: Karpur, Arjun, et al.
Published: (2023)
by: Karpur, Arjun, et al.
Published: (2023)
Fine-Grained Captioning of Long Videos through Scene Graph Consolidation
by: Chu, Sanghyeok, et al.
Published: (2025)
by: Chu, Sanghyeok, et al.
Published: (2025)
STR-Match: Matching SpatioTemporal Relevance Score for Training-Free Video Editing
by: Lee, Junsung, et al.
Published: (2025)
by: Lee, Junsung, et al.
Published: (2025)
ICM-SR: Image-Conditioned Manifold Regularization for Image Super-Resolution
by: Kang, Junoh, et al.
Published: (2025)
by: Kang, Junoh, et al.
Published: (2025)
Merge and Bound: Direct Manipulations on Weights for Class Incremental Learning
by: Kim, Taehoon, et al.
Published: (2025)
by: Kim, Taehoon, et al.
Published: (2025)
Communication-Efficient Federated Learning with Accelerated Client Gradient
by: Kim, Geeho, et al.
Published: (2022)
by: Kim, Geeho, et al.
Published: (2022)
Diffusion-Based Conditional Image Editing through Optimized Inference with Guidance
by: Lee, Hyunsoo, et al.
Published: (2024)
by: Lee, Hyunsoo, et al.
Published: (2024)
Revisiting Machine Unlearning with Dimensional Alignment
by: Seo, Seonguk, et al.
Published: (2024)
by: Seo, Seonguk, et al.
Published: (2024)
Enhanced Diffusion Sampling via Extrapolation with Multiple ODE Solutions
by: Choi, Jinyoung, et al.
Published: (2025)
by: Choi, Jinyoung, et al.
Published: (2025)
A Training-Free Defense Framework for Robust Learned Image Compression
by: Song, Myungseo, et al.
Published: (2024)
by: Song, Myungseo, et al.
Published: (2024)
Diffusion-Based Image-to-Image Translation by Noise Correction via Prompt Interpolation
by: Lee, Junsung, et al.
Published: (2024)
by: Lee, Junsung, et al.
Published: (2024)
GP-4DGS: Probabilistic 4D Gaussian Splatting from Monocular Video via Variational Gaussian Processes
by: Kim, Mijeong, et al.
Published: (2026)
by: Kim, Mijeong, et al.
Published: (2026)
4D Gaussian Splatting in the Wild with Uncertainty-Aware Regularization
by: Kim, Mijeong, et al.
Published: (2024)
by: Kim, Mijeong, et al.
Published: (2024)
Cross-Class Feature Augmentation for Class Incremental Learning
by: Kim, Taehoon, et al.
Published: (2023)
by: Kim, Taehoon, et al.
Published: (2023)
Map the Flow: Revealing Hidden Pathways of Information in VideoLLMs
by: Kim, Minji, et al.
Published: (2025)
by: Kim, Minji, et al.
Published: (2025)
Leveraging Temporal Contextualization for Video Action Recognition
by: Kim, Minji, et al.
Published: (2024)
by: Kim, Minji, et al.
Published: (2024)
PIDM-DP: Physics-Informed Diffusion with Dormand-Prince Integration for Chaotic System Identification and State Reconstruction across Multiple Dynamical Regimes
by: Dabral, Shailendra
Published: (2026)
by: Dabral, Shailendra
Published: (2026)
Optimal bound for singularities on Fano type fibrations of relative dimension one
by: Chen, Bingyi
Published: (2022)
by: Chen, Bingyi
Published: (2022)
Boundedness of klt complements on Fano fibrations over surfaces
by: Chen, Bingyi
Published: (2024)
by: Chen, Bingyi
Published: (2024)
Effective bound for singularities on toric fibrations
by: Chen, Bingyi
Published: (2023)
by: Chen, Bingyi
Published: (2023)
A linear-time algorithm to compute the conjugate of nonconvex bivariate piecewise linear-quadratic functions
by: Karmarkar, Tanmaya, et al.
Published: (2025)
by: Karmarkar, Tanmaya, et al.
Published: (2025)
A Metal‐Free Eliminative Azide‐Olefinic Cycloaddition Route to Sulfonylated 1,2,3‐Triazoles From Vinyl Bissulfones: Impact of an Additional Sulfonyl Group on Regioselectivity
by: Amitabha Bose, et al.
Published: (2025)
by: Amitabha Bose, et al.
Published: (2025)
“Wet lung” of the newborn: Respiratory signs and symptoms caused by cardiac physiology?
by: Koert de Waal
Published: (2024)
by: Koert de Waal
Published: (2024)
kylieainslie/mitey: mitey 0.3.1
by: Kylie Ainslie
Published: (2026)
by: Kylie Ainslie
Published: (2026)
India : historia del subcontinente desde las culturas del indo hasta el comienzo del dominio inglés / Ainslie T. Embree y Friedrich Wilhelm; traductores Antón Dieterich, María Isabel Carrillo
by: Embree, Ainslie
by: Embree, Ainslie
Inadequate housing is not neglect: How the family regulation system punishes parents for a housing crisis out of their control
by: Ainslie Martin
Published: (2025)
by: Ainslie Martin
Published: (2025)
Autonomous Reshaping of Expression Landscapes by DNA Methylation
by: Wang, Kaifeng, et al.
Published: (2026)
by: Wang, Kaifeng, et al.
Published: (2026)
Commentary on Epidemiology of mental health comorbidity in patients with atopic dermatitis: An analysis of global trends from 1998 to 2022
by: Anthony Bewley
Published: (2024)
by: Anthony Bewley
Published: (2024)
FIFO-Diffusion: Generating Infinite Videos from Text without Training
by: Kim, Jihwan, et al.
Published: (2024)
by: Kim, Jihwan, et al.
Published: (2024)
Relaxed Contrastive Learning for Federated Learning
by: Seo, Seonguk, et al.
Published: (2024)
by: Seo, Seonguk, et al.
Published: (2024)
Beyond the Ground Truth: Enhanced Supervision for Image Restoration
by: Ryou, Donghun, et al.
Published: (2025)
by: Ryou, Donghun, et al.
Published: (2025)
Similar Items
-
TIPS: Text-Image Pretraining with Spatial awareness
by: Maninis, Kevis-Kokitsi, et al.
Published: (2024) -
OmniNOCS: A unified NOCS dataset and model for 3D lifting of 2D objects
by: Krishnan, Akshay, et al.
Published: (2024) -
EgoCast: Forecasting Egocentric Human Pose in the Wild
by: Escobar, Maria, et al.
Published: (2024) -
OmniGlue: Generalizable Feature Matching with Foundation Model Guidance
by: Jiang, Hanwen, et al.
Published: (2024) -
Probing the 3D Awareness of Visual Foundation Models
by: Banani, Mohamed El, et al.
Published: (2024)