FineLIP: Extending CLIP's Reach via Fine-Grained Alignment with Longer Text Inputs
Fuente:
arXiv
Saved in:
| Main Authors: | Asokan, Mothilal, Wu, Kebin, Albreiki, Fatima |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Positional Bias in Multimodal Embedding Models: Do They Favor the Beginning, the Middle, or the End?
by: Wu, Kebin, et al.
Published: (2025)
by: Wu, Kebin, et al.
Published: (2025)
LatteCLIP: Unsupervised CLIP Fine-Tuning via LMM-Synthetic Texts
by: Cao, Anh-Quan, et al.
Published: (2024)
by: Cao, Anh-Quan, et al.
Published: (2024)
Medical Image Synthesis via Fine-Grained Image-Text Alignment and Anatomy-Pathology Prompting
by: Chen, Wenting, et al.
Published: (2024)
by: Chen, Wenting, et al.
Published: (2024)
TokLIP: Marry Visual Tokens to CLIP for Multimodal Comprehension and Generation
by: Lin, Haokun, et al.
Published: (2025)
by: Lin, Haokun, et al.
Published: (2025)
FG-CLIP: Fine-Grained Visual and Textual Alignment
by: Xie, Chunyu, et al.
Published: (2025)
by: Xie, Chunyu, et al.
Published: (2025)
Fine-Grained Image-Text Alignment in Medical Imaging Enables Explainable Cyclic Image-Report Generation
by: Chen, Wenting, et al.
Published: (2023)
by: Chen, Wenting, et al.
Published: (2023)
VideoComp: Advancing Fine-Grained Compositional and Temporal Alignment in Video-Text Models
by: Kim, Dahun, et al.
Published: (2025)
by: Kim, Dahun, et al.
Published: (2025)
microCLIP: Unsupervised CLIP Adaptation via Coarse-Fine Token Fusion for Fine-Grained Image Classification
by: Silva, Sathira, et al.
Published: (2025)
by: Silva, Sathira, et al.
Published: (2025)
A Federated Learning-Friendly Approach for Parameter-Efficient Fine-Tuning of SAM in 3D Segmentation
by: Asokan, Mothilal, et al.
Published: (2024)
by: Asokan, Mothilal, et al.
Published: (2024)
CatLIP: CLIP-level Visual Recognition Accuracy with 2.7x Faster Pre-training on Web-scale Image-Text Data
by: Mehta, Sachin, et al.
Published: (2024)
by: Mehta, Sachin, et al.
Published: (2024)
CompAlign: Improving Compositional Text-to-Image Generation with a Complex Benchmark and Fine-Grained Feedback
by: Wan, Yixin, et al.
Published: (2025)
by: Wan, Yixin, et al.
Published: (2025)
Mining Fine-Grained Image-Text Alignment for Zero-Shot Captioning via Text-Only Training
by: Qiu, Longtian, et al.
Published: (2024)
by: Qiu, Longtian, et al.
Published: (2024)
Analyzing Fine-Grained Alignment and Enhancing Vision Understanding in Multimodal Language Models
by: Jiang, Jiachen, et al.
Published: (2025)
by: Jiang, Jiachen, et al.
Published: (2025)
TOPA: Extending Large Language Models for Video Understanding via Text-Only Pre-Alignment
by: Li, Wei, et al.
Published: (2024)
by: Li, Wei, et al.
Published: (2024)
VISCO: Benchmarking Fine-Grained Critique and Correction Towards Self-Improvement in Visual Reasoning
by: Wu, Xueqing, et al.
Published: (2024)
by: Wu, Xueqing, et al.
Published: (2024)
COHERENCE: Benchmarking Fine-Grained Image-Text Alignment in Interleaved Multimodal Contexts
by: Wang, Bingli, et al.
Published: (2026)
by: Wang, Bingli, et al.
Published: (2026)
Evaluating Automated Radiology Report Quality through Fine-Grained Phrasal Grounding of Clinical Findings
by: Mahmood, Razi, et al.
Published: (2024)
by: Mahmood, Razi, et al.
Published: (2024)
ReasonMap: Towards Fine-Grained Visual Reasoning from Transit Maps
by: Feng, Sicheng, et al.
Published: (2025)
by: Feng, Sicheng, et al.
Published: (2025)
VideoEspresso: A Large-Scale Chain-of-Thought Dataset for Fine-Grained Video Reasoning via Core Frame Selection
by: Han, Songhao, et al.
Published: (2024)
by: Han, Songhao, et al.
Published: (2024)
DreamRunner: Fine-Grained Compositional Story-to-Video Generation with Retrieval-Augmented Motion Adaptation
by: Wang, Zun, et al.
Published: (2024)
by: Wang, Zun, et al.
Published: (2024)
ComCLIP: Training-Free Compositional Image and Text Matching
by: Jiang, Kenan, et al.
Published: (2022)
by: Jiang, Kenan, et al.
Published: (2022)
MMGeoLM: Hard Negative Contrastive Learning for Fine-Grained Geometric Understanding in Large Multimodal Models
by: Sun, Kai, et al.
Published: (2025)
by: Sun, Kai, et al.
Published: (2025)
HNC: Leveraging Hard Negative Captions towards Models with Fine-Grained Visual-Linguistic Comprehension Capabilities
by: Dönmez, Esra, et al.
Published: (2026)
by: Dönmez, Esra, et al.
Published: (2026)
LEGO Co-builder: Exploring Fine-Grained Vision-Language Modeling for Multimodal LEGO Assembly Assistants
by: Huang, Haochen, et al.
Published: (2025)
by: Huang, Haochen, et al.
Published: (2025)
BYO-Eval: Build Your Own Dataset for Fine-Grained Visual Assessment of Multimodal Language Models
by: Arnould, Ludovic, et al.
Published: (2025)
by: Arnould, Ludovic, et al.
Published: (2025)
FB-CLIP: Fine-Grained Zero-Shot Anomaly Detection with Foreground-Background Disentanglement
by: Hu, Ming, et al.
Published: (2026)
by: Hu, Ming, et al.
Published: (2026)
Pushing the Frontier of Black-Box LVLM Attacks via Fine-Grained Detail Targeting
by: Zhao, Xiaohan, et al.
Published: (2026)
by: Zhao, Xiaohan, et al.
Published: (2026)
KH-FUNSD: A Hierarchical and Fine-Grained Layout Analysis Dataset for Low-Resource Khmer Business Document
by: Thuon, Nimol, et al.
Published: (2025)
by: Thuon, Nimol, et al.
Published: (2025)
Unconstrained Open Vocabulary Image Classification: Zero-Shot Transfer from Text to Image via CLIP Inversion
by: Allgeuer, Philipp, et al.
Published: (2024)
by: Allgeuer, Philipp, et al.
Published: (2024)
FineBench: Benchmarking and Enhancing Vision-Language Models for Fine-grained Human Activity Understanding
by: Faure, Gueter Josmy, et al.
Published: (2026)
by: Faure, Gueter Josmy, et al.
Published: (2026)
Sim-CLIP: Unsupervised Siamese Adversarial Fine-Tuning for Robust and Semantically-Rich Vision-Language Models
by: Hossain, Md Zarif, et al.
Published: (2024)
by: Hossain, Md Zarif, et al.
Published: (2024)
Detecting and Mitigating Hallucination in Large Vision Language Models via Fine-Grained AI Feedback
by: Xiao, Wenyi, et al.
Published: (2024)
by: Xiao, Wenyi, et al.
Published: (2024)
Why LVLMs Are More Prone to Hallucinations in Longer Responses: The Role of Context
by: Zheng, Ge, et al.
Published: (2025)
by: Zheng, Ge, et al.
Published: (2025)
Uni-Mlip: Unified Self-supervision for Medical Vision Language Pre-training
by: Bawazir, Ameera, et al.
Published: (2024)
by: Bawazir, Ameera, et al.
Published: (2024)
Navigating the Nuances: A Fine-grained Evaluation of Vision-Language Navigation
by: Wang, Zehao, et al.
Published: (2024)
by: Wang, Zehao, et al.
Published: (2024)
Decomposed Vision-Language Alignment for Fine-Grained Open-Vocabulary Segmentation
by: Wang, Chenhao, et al.
Published: (2026)
by: Wang, Chenhao, et al.
Published: (2026)
Capturing Fine-Grained Alignments Improves 3D Affordance Detection
by: Tokumitsu, Junsei, et al.
Published: (2025)
by: Tokumitsu, Junsei, et al.
Published: (2025)
Multimodal Alignment with Cross-Attentive GRUs for Fine-Grained Video Understanding
by: Kim, Namho, et al.
Published: (2025)
by: Kim, Namho, et al.
Published: (2025)
CropVLM: Learning to Zoom for Fine-Grained Vision-Language Perception
by: Carvalho, Miguel, et al.
Published: (2025)
by: Carvalho, Miguel, et al.
Published: (2025)
Fine-tuning CLIP Text Encoders with Two-step Paraphrasing
by: Kim, Hyunjae, et al.
Published: (2024)
by: Kim, Hyunjae, et al.
Published: (2024)
Similar Items
-
Positional Bias in Multimodal Embedding Models: Do They Favor the Beginning, the Middle, or the End?
by: Wu, Kebin, et al.
Published: (2025) -
LatteCLIP: Unsupervised CLIP Fine-Tuning via LMM-Synthetic Texts
by: Cao, Anh-Quan, et al.
Published: (2024) -
Medical Image Synthesis via Fine-Grained Image-Text Alignment and Anatomy-Pathology Prompting
by: Chen, Wenting, et al.
Published: (2024) -
TokLIP: Marry Visual Tokens to CLIP for Multimodal Comprehension and Generation
by: Lin, Haokun, et al.
Published: (2025) -
FG-CLIP: Fine-Grained Visual and Textual Alignment
by: Xie, Chunyu, et al.
Published: (2025)