AmorLIP: Efficient Language-Image Pretraining via Amortization
Fuente:
arXiv
Saved in:
| Main Authors: | Sun, Haotian, Li, Yitong, Zhuang, Yuchen, He, Niao, Dai, Hanjun, Dai, Bo |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Image Clustering via the Principle of Rate Reduction in the Age of Pretrained Models
by: Chu, Tianzhe, et al.
Published: (2023)
by: Chu, Tianzhe, et al.
Published: (2023)
FuseLIP: Multimodal Embeddings via Early Fusion of Discrete Tokens
by: Schlarmann, Christian, et al.
Published: (2025)
by: Schlarmann, Christian, et al.
Published: (2025)
LongProLIP: A Probabilistic Vision-Language Model with Long Context Text
by: Chun, Sanghyuk, et al.
Published: (2025)
by: Chun, Sanghyuk, et al.
Published: (2025)
Detecting Backdoor Samples in Contrastive Language Image Pretraining
by: Huang, Hanxun, et al.
Published: (2025)
by: Huang, Hanxun, et al.
Published: (2025)
Data-Efficient Contrastive Language-Image Pretraining: Prioritizing Data Quality over Quantity
by: Joshi, Siddharth, et al.
Published: (2024)
by: Joshi, Siddharth, et al.
Published: (2024)
RankSEG-RMA: An Efficient Segmentation Algorithm via Reciprocal Moment Approximation
by: Wang, Zixun, et al.
Published: (2025)
by: Wang, Zixun, et al.
Published: (2025)
Consistent Amortized Clustering via Generative Flow Networks
by: Chelly, Irit, et al.
Published: (2025)
by: Chelly, Irit, et al.
Published: (2025)
EC-DIT: Scaling Diffusion Transformers with Adaptive Expert-Choice Routing
by: Sun, Haotian, et al.
Published: (2024)
by: Sun, Haotian, et al.
Published: (2024)
FPPL: An Efficient and Non-IID Robust Federated Continual Learning Framework
by: He, Yuchen, et al.
Published: (2024)
by: He, Yuchen, et al.
Published: (2024)
Boosting 3D Object Generation through PBR Materials
by: Wang, Yitong, et al.
Published: (2024)
by: Wang, Yitong, et al.
Published: (2024)
Concept-Aware Batch Sampling Improves Language-Image Pretraining
by: Ghosh, Adhiraj, et al.
Published: (2025)
by: Ghosh, Adhiraj, et al.
Published: (2025)
RankCLIP: Ranking-Consistent Language-Image Pretraining
by: Zhang, Yiming, et al.
Published: (2024)
by: Zhang, Yiming, et al.
Published: (2024)
BIFRÖST: 3D-Aware Image compositing with Language Instructions
by: Li, Lingxiao, et al.
Published: (2024)
by: Li, Lingxiao, et al.
Published: (2024)
Amortized Posterior Sampling with Diffusion Prior Distillation
by: Mammadov, Abbas, et al.
Published: (2024)
by: Mammadov, Abbas, et al.
Published: (2024)
SoFlow: Solution Flow Models for One-Step Generative Modeling
by: Luo, Tianze, et al.
Published: (2025)
by: Luo, Tianze, et al.
Published: (2025)
VARD: Efficient and Dense Fine-Tuning for Diffusion Models with Value-based RL
by: Dai, Fengyuan, et al.
Published: (2025)
by: Dai, Fengyuan, et al.
Published: (2025)
ClearVision: Leveraging CycleGAN and SigLIP-2 for Robust All-Weather Classification in Traffic Camera Imagery
by: Sivaraman, Anush Lakshman, et al.
Published: (2025)
by: Sivaraman, Anush Lakshman, et al.
Published: (2025)
Functional Mean Flow in Hilbert Space
by: Li, Zhiqi, et al.
Published: (2025)
by: Li, Zhiqi, et al.
Published: (2025)
2.5 Years in Class: A Multimodal Textbook for Vision-Language Pretraining
by: Zhang, Wenqi, et al.
Published: (2025)
by: Zhang, Wenqi, et al.
Published: (2025)
Negative Label Guided OOD Detection with Pretrained Vision-Language Models
by: Jiang, Xue, et al.
Published: (2024)
by: Jiang, Xue, et al.
Published: (2024)
SeLIP: Similarity Enhanced Contrastive Language Image Pretraining for Multi-modal Head MRI
by: Liu, Zhiyang, et al.
Published: (2025)
by: Liu, Zhiyang, et al.
Published: (2025)
Amortizing intractable inference in diffusion models for vision, language, and control
by: Venkatraman, Siddarth, et al.
Published: (2024)
by: Venkatraman, Siddarth, et al.
Published: (2024)
Noise Hypernetworks: Amortizing Test-Time Compute in Diffusion Models
by: Eyring, Luca, et al.
Published: (2025)
by: Eyring, Luca, et al.
Published: (2025)
NeuroPictor: Refining fMRI-to-Image Reconstruction via Multi-individual Pretraining and Multi-level Modulation
by: Huo, Jingyang, et al.
Published: (2024)
by: Huo, Jingyang, et al.
Published: (2024)
Rethinking Amortized Neural Representations for High-Resolution Terrain Elevation Data
by: Feng, Haoan, et al.
Published: (2026)
by: Feng, Haoan, et al.
Published: (2026)
Reliable Thinking with Images
by: Li, Haobin, et al.
Published: (2026)
by: Li, Haobin, et al.
Published: (2026)
Federated EndoViT: Pretraining Vision Transformers via Federated Learning on Endoscopic Image Collections
by: Kirchner, Max, et al.
Published: (2025)
by: Kirchner, Max, et al.
Published: (2025)
Efficient Neural Network Training via Subset Pretraining
by: Spörer, Jan, et al.
Published: (2024)
by: Spörer, Jan, et al.
Published: (2024)
Stochastic Siamese MAE Pretraining for Longitudinal Medical Images
by: Emre, Taha, et al.
Published: (2025)
by: Emre, Taha, et al.
Published: (2025)
Time-to-Event Pretraining for 3D Medical Imaging
by: Huo, Zepeng, et al.
Published: (2024)
by: Huo, Zepeng, et al.
Published: (2024)
Pretrained Image-Text Models are Secretly Video Captioners
by: Zhang, Chunhui, et al.
Published: (2025)
by: Zhang, Chunhui, et al.
Published: (2025)
Trajectory Consistency for One-Step Generation on Euler Mean Flows
by: Li, Zhiqi, et al.
Published: (2026)
by: Li, Zhiqi, et al.
Published: (2026)
High-Fidelity Text-to-Image Generation from Pre-Trained Vision-Language Models via Distribution-Conditioned Diffusion Decoding
by: Hong, Ji Woo, et al.
Published: (2026)
by: Hong, Ji Woo, et al.
Published: (2026)
Foundation Model-oriented Robustness: Robust Image Model Evaluation with Pretrained Models
by: Zhang, Peiyan, et al.
Published: (2023)
by: Zhang, Peiyan, et al.
Published: (2023)
Region-centric Image-Language Pretraining for Open-Vocabulary Detection
by: Kim, Dahun, et al.
Published: (2023)
by: Kim, Dahun, et al.
Published: (2023)
Universal Pose Pretraining for Generalizable Vision-Language-Action Policies
by: Lin, Haitao, et al.
Published: (2026)
by: Lin, Haitao, et al.
Published: (2026)
Frequency-Aware Autoregressive Modeling for Efficient High-Resolution Image Synthesis
by: Chen, Zhuokun, et al.
Published: (2025)
by: Chen, Zhuokun, et al.
Published: (2025)
SmartFRZ: An Efficient Training Framework using Attention-Based Layer Freezing
by: Li, Sheng, et al.
Published: (2024)
by: Li, Sheng, et al.
Published: (2024)
Contrastive Localized Language-Image Pre-Training
by: Chen, Hong-You, et al.
Published: (2024)
by: Chen, Hong-You, et al.
Published: (2024)
E-CAR: Efficient Continuous Autoregressive Image Generation via Multistage Modeling
by: Yuan, Zhihang, et al.
Published: (2024)
by: Yuan, Zhihang, et al.
Published: (2024)
Similar Items
-
Image Clustering via the Principle of Rate Reduction in the Age of Pretrained Models
by: Chu, Tianzhe, et al.
Published: (2023) -
FuseLIP: Multimodal Embeddings via Early Fusion of Discrete Tokens
by: Schlarmann, Christian, et al.
Published: (2025) -
LongProLIP: A Probabilistic Vision-Language Model with Long Context Text
by: Chun, Sanghyuk, et al.
Published: (2025) -
Detecting Backdoor Samples in Contrastive Language Image Pretraining
by: Huang, Hanxun, et al.
Published: (2025) -
Data-Efficient Contrastive Language-Image Pretraining: Prioritizing Data Quality over Quantity
by: Joshi, Siddharth, et al.
Published: (2024)