LongProLIP: A Probabilistic Vision-Language Model with Long Context Text
Fuente:
arXiv
Saved in:
| Main Authors: | Chun, Sanghyuk, Yun, Sangdoo |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Probabilistic Language-Image Pre-Training
by: Chun, Sanghyuk, et al.
Published: (2024)
by: Chun, Sanghyuk, et al.
Published: (2024)
Improved Probabilistic Image-Text Representations
by: Chun, Sanghyuk
Published: (2023)
by: Chun, Sanghyuk
Published: (2023)
Toward Interactive Regional Understanding in Vision-Large Language Models
by: Lee, Jungbeom, et al.
Published: (2024)
by: Lee, Jungbeom, et al.
Published: (2024)
Emergence of Text Readability in Vision Language Models
by: Park, Jaeyoo, et al.
Published: (2025)
by: Park, Jaeyoo, et al.
Published: (2025)
Multiplicity is an Inevitable and Inherent Challenge in Multimodal Learning
by: Chun, Sanghyuk
Published: (2025)
by: Chun, Sanghyuk
Published: (2025)
HYPE: Hyperbolic Entailment Filtering for Underspecified Images and Texts
by: Kim, Wonjae, et al.
Published: (2024)
by: Kim, Wonjae, et al.
Published: (2024)
Rotary Position Embedding for Vision Transformer
by: Heo, Byeongho, et al.
Published: (2024)
by: Heo, Byeongho, et al.
Published: (2024)
HoPE: Hybrid of Position Embedding for Long Context Vision-Language Models
by: Li, Haoran, et al.
Published: (2025)
by: Li, Haoran, et al.
Published: (2025)
SimPro: A Simple Probabilistic Framework Towards Realistic Long-Tailed Semi-Supervised Learning
by: Du, Chaoqun, et al.
Published: (2024)
by: Du, Chaoqun, et al.
Published: (2024)
AmorLIP: Efficient Language-Image Pretraining via Amortization
by: Sun, Haotian, et al.
Published: (2025)
by: Sun, Haotian, et al.
Published: (2025)
Model Stock: All we need is just a few fine-tuned models
by: Jang, Dong-Hwan, et al.
Published: (2024)
by: Jang, Dong-Hwan, et al.
Published: (2024)
ProTIP: Probabilistic Robustness Verification on Text-to-Image Diffusion Models against Stochastic Perturbation
by: Zhang, Yi, et al.
Published: (2024)
by: Zhang, Yi, et al.
Published: (2024)
Language-only Efficient Training of Zero-shot Composed Image Retrieval
by: Gu, Geonmo, et al.
Published: (2023)
by: Gu, Geonmo, et al.
Published: (2023)
Masking meets Supervision: A Strong Learning Alliance
by: Heo, Byeongho, et al.
Published: (2023)
by: Heo, Byeongho, et al.
Published: (2023)
Post-hoc Probabilistic Vision-Language Models
by: Baumann, Anton, et al.
Published: (2024)
by: Baumann, Anton, et al.
Published: (2024)
Efficient and Long-Tailed Generalization for Pre-trained Vision-Language Model
by: Shi, Jiang-Xin, et al.
Published: (2024)
by: Shi, Jiang-Xin, et al.
Published: (2024)
RoCOCO: Robustness Benchmark of MS-COCO to Stress-test Image-Text Matching Models
by: Park, Seulki, et al.
Published: (2023)
by: Park, Seulki, et al.
Published: (2023)
DNNs May Determine Major Properties of Their Outputs Early, with Timing Possibly Driven by Bias
by: Park, Song, et al.
Published: (2025)
by: Park, Song, et al.
Published: (2025)
ClearVision: Leveraging CycleGAN and SigLIP-2 for Robust All-Weather Classification in Traffic Camera Imagery
by: Sivaraman, Anush Lakshman, et al.
Published: (2025)
by: Sivaraman, Anush Lakshman, et al.
Published: (2025)
Probabilistic Contrastive Learning for Long-Tailed Visual Recognition
by: Du, Chaoqun, et al.
Published: (2024)
by: Du, Chaoqun, et al.
Published: (2024)
Unleashing the Power of Vision-Language Models for Long-Tailed Multi-Label Visual Recognition
by: Tang, Wei, et al.
Published: (2025)
by: Tang, Wei, et al.
Published: (2025)
FuseLIP: Multimodal Embeddings via Early Fusion of Discrete Tokens
by: Schlarmann, Christian, et al.
Published: (2025)
by: Schlarmann, Christian, et al.
Published: (2025)
Fair Context Learning for Evidence-Balanced Test-Time Adaptation in Vision-Language Models
by: Yun, Sanggeon, et al.
Published: (2026)
by: Yun, Sanggeon, et al.
Published: (2026)
DreamLIP: Language-Image Pre-training with Long Captions
by: Zheng, Kecheng, et al.
Published: (2024)
by: Zheng, Kecheng, et al.
Published: (2024)
CAPT: Class-Aware Prompt Tuning for Federated Long-Tailed Learning with Vision-Language Model
by: Hou, Shihao, et al.
Published: (2025)
by: Hou, Shihao, et al.
Published: (2025)
Improving Long-Text Alignment for Text-to-Image Diffusion Models
by: Liu, Luping, et al.
Published: (2024)
by: Liu, Luping, et al.
Published: (2024)
Atlas: Multi-Scale Attention Improves Long Context Image Modeling
by: Agrawal, Kumar Krishna, et al.
Published: (2025)
by: Agrawal, Kumar Krishna, et al.
Published: (2025)
VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling
by: Li, Xinhao, et al.
Published: (2024)
by: Li, Xinhao, et al.
Published: (2024)
RL makes MLLMs see better than SFT
by: Song, Junha, et al.
Published: (2025)
by: Song, Junha, et al.
Published: (2025)
Match me if you can: Semi-Supervised Semantic Correspondence Learning with Unpaired Images
by: Kim, Jiwon, et al.
Published: (2023)
by: Kim, Jiwon, et al.
Published: (2023)
DaWin: Training-free Dynamic Weight Interpolation for Robust Adaptation
by: Oh, Changdae, et al.
Published: (2024)
by: Oh, Changdae, et al.
Published: (2024)
Towards Long-window Anchoring in Vision-Language Model Distillation
by: Zhou, Haoyi, et al.
Published: (2025)
by: Zhou, Haoyi, et al.
Published: (2025)
Scaling Vision Language Models for Pharmaceutical Long Form Video Reasoning on Industrial GenAI Platform
by: Mishra, Suyash, et al.
Published: (2026)
by: Mishra, Suyash, et al.
Published: (2026)
Probabilistic Embeddings for Frozen Vision-Language Models: Uncertainty Quantification with Gaussian Process Latent Variable Models
by: Venkataramanan, Aishwarya, et al.
Published: (2025)
by: Venkataramanan, Aishwarya, et al.
Published: (2025)
Not Only Text: Exploring Compositionality of Visual Representations in Vision-Language Models
by: Berasi, Davide, et al.
Published: (2025)
by: Berasi, Davide, et al.
Published: (2025)
LoGeR: Long-Context Geometric Reconstruction with Hybrid Memory
by: Zhang, Junyi, et al.
Published: (2026)
by: Zhang, Junyi, et al.
Published: (2026)
QLAM: A Quantum Long-Attention Memory Approach to Long-Sequence Token Modeling
by: Nguyen, Hoang-Quan, et al.
Published: (2026)
by: Nguyen, Hoang-Quan, et al.
Published: (2026)
Similarity of Neural Architectures using Adversarial Attack Transferability
by: Hwang, Jaehui, et al.
Published: (2022)
by: Hwang, Jaehui, et al.
Published: (2022)
Small Vision-Language Models are Smart Compressors for Long Video Understanding
by: Fei, Junjie, et al.
Published: (2026)
by: Fei, Junjie, et al.
Published: (2026)
ProHOC: Probabilistic Hierarchical Out-of-Distribution Classification via Multi-Depth Networks
by: Wallin, Erik, et al.
Published: (2025)
by: Wallin, Erik, et al.
Published: (2025)
Similar Items
-
Probabilistic Language-Image Pre-Training
by: Chun, Sanghyuk, et al.
Published: (2024) -
Improved Probabilistic Image-Text Representations
by: Chun, Sanghyuk
Published: (2023) -
Toward Interactive Regional Understanding in Vision-Large Language Models
by: Lee, Jungbeom, et al.
Published: (2024) -
Emergence of Text Readability in Vision Language Models
by: Park, Jaeyoo, et al.
Published: (2025) -
Multiplicity is an Inevitable and Inherent Challenge in Multimodal Learning
by: Chun, Sanghyuk
Published: (2025)