Probabilistic Language-Image Pre-Training
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Chun, Sanghyuk, Kim, Wonjae, Park, Song, Yun, Sangdoo |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
LongProLIP: A Probabilistic Vision-Language Model with Long Context Text
von: Chun, Sanghyuk, et al.
Veröffentlicht: (2025)
von: Chun, Sanghyuk, et al.
Veröffentlicht: (2025)
Language-only Efficient Training of Zero-shot Composed Image Retrieval
von: Gu, Geonmo, et al.
Veröffentlicht: (2023)
von: Gu, Geonmo, et al.
Veröffentlicht: (2023)
Emergence of Text Readability in Vision Language Models
von: Park, Jaeyoo, et al.
Veröffentlicht: (2025)
von: Park, Jaeyoo, et al.
Veröffentlicht: (2025)
Improved Probabilistic Image-Text Representations
von: Chun, Sanghyuk
Veröffentlicht: (2023)
von: Chun, Sanghyuk
Veröffentlicht: (2023)
HYPE: Hyperbolic Entailment Filtering for Underspecified Images and Texts
von: Kim, Wonjae, et al.
Veröffentlicht: (2024)
von: Kim, Wonjae, et al.
Veröffentlicht: (2024)
Toward Interactive Regional Understanding in Vision-Large Language Models
von: Lee, Jungbeom, et al.
Veröffentlicht: (2024)
von: Lee, Jungbeom, et al.
Veröffentlicht: (2024)
CompoDiff: Versatile Composed Image Retrieval With Latent Diffusion
von: Gu, Geonmo, et al.
Veröffentlicht: (2023)
von: Gu, Geonmo, et al.
Veröffentlicht: (2023)
Multiplicity is an Inevitable and Inherent Challenge in Multimodal Learning
von: Chun, Sanghyuk
Veröffentlicht: (2025)
von: Chun, Sanghyuk
Veröffentlicht: (2025)
DNNs May Determine Major Properties of Their Outputs Early, with Timing Possibly Driven by Bias
von: Park, Song, et al.
Veröffentlicht: (2025)
von: Park, Song, et al.
Veröffentlicht: (2025)
Rotary Position Embedding for Vision Transformer
von: Heo, Byeongho, et al.
Veröffentlicht: (2024)
von: Heo, Byeongho, et al.
Veröffentlicht: (2024)
ECCV Caption: Correcting False Negatives by Collecting Machine-and-Human-verified Image-Caption Associations for MS-COCO
von: Chun, Sanghyuk, et al.
Veröffentlicht: (2022)
von: Chun, Sanghyuk, et al.
Veröffentlicht: (2022)
DaWin: Training-free Dynamic Weight Interpolation for Robust Adaptation
von: Oh, Changdae, et al.
Veröffentlicht: (2024)
von: Oh, Changdae, et al.
Veröffentlicht: (2024)
Match me if you can: Semi-Supervised Semantic Correspondence Learning with Unpaired Images
von: Kim, Jiwon, et al.
Veröffentlicht: (2023)
von: Kim, Jiwon, et al.
Veröffentlicht: (2023)
Masking meets Supervision: A Strong Learning Alliance
von: Heo, Byeongho, et al.
Veröffentlicht: (2023)
von: Heo, Byeongho, et al.
Veröffentlicht: (2023)
STELLA: Continual Audio-Video Pre-training with Spatio-Temporal Localized Alignment
von: Lee, Jaewoo, et al.
Veröffentlicht: (2023)
von: Lee, Jaewoo, et al.
Veröffentlicht: (2023)
Similarity of Neural Architectures using Adversarial Attack Transferability
von: Hwang, Jaehui, et al.
Veröffentlicht: (2022)
von: Hwang, Jaehui, et al.
Veröffentlicht: (2022)
RL makes MLLMs see better than SFT
von: Song, Junha, et al.
Veröffentlicht: (2025)
von: Song, Junha, et al.
Veröffentlicht: (2025)
Contrastive Localized Language-Image Pre-Training
von: Chen, Hong-You, et al.
Veröffentlicht: (2024)
von: Chen, Hong-You, et al.
Veröffentlicht: (2024)
Model Stock: All we need is just a few fine-tuned models
von: Jang, Dong-Hwan, et al.
Veröffentlicht: (2024)
von: Jang, Dong-Hwan, et al.
Veröffentlicht: (2024)
RoCOCO: Robustness Benchmark of MS-COCO to Stress-test Image-Text Matching Models
von: Park, Seulki, et al.
Veröffentlicht: (2023)
von: Park, Seulki, et al.
Veröffentlicht: (2023)
An Efficient Post-hoc Framework for Reducing Task Discrepancy of Text Encoders for Composed Image Retrieval
von: Byun, Jaeseok, et al.
Veröffentlicht: (2024)
von: Byun, Jaeseok, et al.
Veröffentlicht: (2024)
Centered Masking for Language-Image Pre-Training
von: Liang, Mingliang, et al.
Veröffentlicht: (2024)
von: Liang, Mingliang, et al.
Veröffentlicht: (2024)
Embedding Geometries of Contrastive Language-Image Pre-Training
von: Chou, Jason Chuan-Chih, et al.
Veröffentlicht: (2024)
von: Chou, Jason Chuan-Chih, et al.
Veröffentlicht: (2024)
A Closer Look at the Robustness of Contrastive Language-Image Pre-Training (CLIP)
von: Tu, Weijie, et al.
Veröffentlicht: (2024)
von: Tu, Weijie, et al.
Veröffentlicht: (2024)
Probabilistic Precision and Recall Towards Reliable Evaluation of Generative Models
von: Park, Dogyun, et al.
Veröffentlicht: (2023)
von: Park, Dogyun, et al.
Veröffentlicht: (2023)
Extract Free Dense Misalignment from CLIP
von: Nam, JeongYeon, et al.
Veröffentlicht: (2024)
von: Nam, JeongYeon, et al.
Veröffentlicht: (2024)
Improving Generative Pre-Training: An In-depth Study of Masked Image Modeling and Denoising Models
von: Choi, Hyesong, et al.
Veröffentlicht: (2024)
von: Choi, Hyesong, et al.
Veröffentlicht: (2024)
Prioritizing Image-Related Tokens Enhances Vision-Language Pre-Training
von: Chen, Yangyi, et al.
Veröffentlicht: (2025)
von: Chen, Yangyi, et al.
Veröffentlicht: (2025)
Memory-Efficient Personalization of Text-to-Image Diffusion Models via Selective Optimization Strategies
von: Choi, Seokeon, et al.
Veröffentlicht: (2025)
von: Choi, Seokeon, et al.
Veröffentlicht: (2025)
Visual Pre-Training on Unlabeled Images using Reinforcement Learning
von: Ghosh, Dibya, et al.
Veröffentlicht: (2025)
von: Ghosh, Dibya, et al.
Veröffentlicht: (2025)
Test-Time Training for Visual Foresight Vision-Language-Action Models
von: Park, Sangwu, et al.
Veröffentlicht: (2026)
von: Park, Sangwu, et al.
Veröffentlicht: (2026)
PARIC: Probabilistic Attention Regularization for Language Guided Image Classification from Pre-trained Vison Language Models
von: Nautiyal, Mayank, et al.
Veröffentlicht: (2025)
von: Nautiyal, Mayank, et al.
Veröffentlicht: (2025)
Steering Guidance for Personalized Text-to-Image Diffusion Models
von: Park, Sunghyun, et al.
Veröffentlicht: (2025)
von: Park, Sunghyun, et al.
Veröffentlicht: (2025)
High-Fidelity Text-to-Image Generation from Pre-Trained Vision-Language Models via Distribution-Conditioned Diffusion Decoding
von: Hong, Ji Woo, et al.
Veröffentlicht: (2026)
von: Hong, Ji Woo, et al.
Veröffentlicht: (2026)
Stabilizing Consistency Training: A Flow Map Analysis and Self-Distillation
von: Kim, Youngjoong, et al.
Veröffentlicht: (2026)
von: Kim, Youngjoong, et al.
Veröffentlicht: (2026)
Source-Free Domain Adaptation Guided by Vision and Vision-Language Pre-Training
von: Zhang, Wenyu, et al.
Veröffentlicht: (2024)
von: Zhang, Wenyu, et al.
Veröffentlicht: (2024)
Zero-Shot Video Restoration and Enhancement Using Pre-Trained Image Diffusion Model
von: Cao, Cong, et al.
Veröffentlicht: (2024)
von: Cao, Cong, et al.
Veröffentlicht: (2024)
TimePre: Bridging Accuracy, Efficiency, and Stability in Probabilistic Time-Series Forecasting
von: Jiang, Lingyu, et al.
Veröffentlicht: (2025)
von: Jiang, Lingyu, et al.
Veröffentlicht: (2025)
Unsupervised Video Domain Adaptation with Masked Pre-Training and Collaborative Self-Training
von: Reddy, Arun, et al.
Veröffentlicht: (2023)
von: Reddy, Arun, et al.
Veröffentlicht: (2023)
PTQ4VM: Post-Training Quantization for Visual Mamba
von: Cho, Younghyun, et al.
Veröffentlicht: (2024)
von: Cho, Younghyun, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
LongProLIP: A Probabilistic Vision-Language Model with Long Context Text
von: Chun, Sanghyuk, et al.
Veröffentlicht: (2025) -
Language-only Efficient Training of Zero-shot Composed Image Retrieval
von: Gu, Geonmo, et al.
Veröffentlicht: (2023) -
Emergence of Text Readability in Vision Language Models
von: Park, Jaeyoo, et al.
Veröffentlicht: (2025) -
Improved Probabilistic Image-Text Representations
von: Chun, Sanghyuk
Veröffentlicht: (2023) -
HYPE: Hyperbolic Entailment Filtering for Underspecified Images and Texts
von: Kim, Wonjae, et al.
Veröffentlicht: (2024)