Long Story Short: Disentangling Compositionality and Long-Caption Understanding in Contrastive VLMs
Fuente:
arXiv
Saved in:
| Main Authors: | Salazar, Israfel, Elliott, Desmond, Kementchedjhieva, Yova |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
SPECS: Specificity-Enhanced CLIP-Score for Long Image Caption Evaluation
by: Chen, Xiaofu, et al.
Published: (2025)
by: Chen, Xiaofu, et al.
Published: (2025)
The Devil is in the EOS: Sequence Training for Detailed Image Captioning
by: Mohamed, Abdelrahman, et al.
Published: (2025)
by: Mohamed, Abdelrahman, et al.
Published: (2025)
A Simple Data Augmentation Strategy for Text-in-Image Scientific VQA
by: Shoer, Belal, et al.
Published: (2025)
by: Shoer, Belal, et al.
Published: (2025)
EFSA: Episodic Few-Shot Adaptation for Text-to-Image Retrieval
by: Huzaifa, Muhammad, et al.
Published: (2024)
by: Huzaifa, Muhammad, et al.
Published: (2024)
VLMs Need Words: Vision Language Models Ignore Visual Detail In Favor of Semantic Anchors
by: Shahgir, Haz Sameen, et al.
Published: (2026)
by: Shahgir, Haz Sameen, et al.
Published: (2026)
Understanding Retrieval Robustness for Retrieval-Augmented Image Captioning
by: Li, Wenyan, et al.
Published: (2024)
by: Li, Wenyan, et al.
Published: (2024)
Long Story Short: Story-level Video Understanding from 20K Short Films
by: Ghermi, Ridouane, et al.
Published: (2024)
by: Ghermi, Ridouane, et al.
Published: (2024)
Do Vision and Language Models Share Concepts? A Vector Space Alignment Study
by: Li, Jiaang, et al.
Published: (2023)
by: Li, Jiaang, et al.
Published: (2023)
Linear Scaling Video VLMs for Long Video Understanding
by: Eyzaguirre, Cristobal, et al.
Published: (2026)
by: Eyzaguirre, Cristobal, et al.
Published: (2026)
LLMs Can Compensate for Deficiencies in Visual Representations
by: Takishita, Sho, et al.
Published: (2025)
by: Takishita, Sho, et al.
Published: (2025)
Noise is an Efficient Learner for Zero-Shot Vision-Language Models
by: Imam, Raza, et al.
Published: (2025)
by: Imam, Raza, et al.
Published: (2025)
LinguDistill: Recovering Linguistic Ability in Vision-Language Models via Selective Cross-Modal Distillation
by: Irawan, Patrick Amadeus, et al.
Published: (2026)
by: Irawan, Patrick Amadeus, et al.
Published: (2026)
FIX-CLIP: Dual-Branch Hierarchical Contrastive Learning via Synthetic Captions for Better Understanding of Long Text
by: Wang, Bingchao, et al.
Published: (2025)
by: Wang, Bingchao, et al.
Published: (2025)
The Role of Data Curation in Image Captioning
by: Li, Wenyan, et al.
Published: (2023)
by: Li, Wenyan, et al.
Published: (2023)
LVC: A Lightweight Compression Framework for Enhancing VLMs in Long Video Understanding
by: Wang, Ziyi, et al.
Published: (2025)
by: Wang, Ziyi, et al.
Published: (2025)
LongCaptioning: Unlocking the Power of Long Video Caption Generation in Large Multimodal Models
by: Wei, Hongchen, et al.
Published: (2025)
by: Wei, Hongchen, et al.
Published: (2025)
Encoder-Decoder Based Long Short-Term Memory (LSTM) Model for Video Captioning
by: Adewale, Sikiru, et al.
Published: (2023)
by: Adewale, Sikiru, et al.
Published: (2023)
Controllable Hybrid Captioner for Improved Long-form Video Understanding
by: Sasse, Kuleen, et al.
Published: (2025)
by: Sasse, Kuleen, et al.
Published: (2025)
Towards Long Video Understanding via Fine-detailed Video Story Generation
by: You, Zeng, et al.
Published: (2024)
by: You, Zeng, et al.
Published: (2024)
Disentangled Representations for Short-Term and Long-Term Person Re-Identification
by: Eom, Chanho, et al.
Published: (2024)
by: Eom, Chanho, et al.
Published: (2024)
Caption This, Reason That: VLMs Caught in the Middle
by: Weng, Zihan, et al.
Published: (2025)
by: Weng, Zihan, et al.
Published: (2025)
SEED-Story: Multimodal Long Story Generation with Large Language Model
by: Yang, Shuai, et al.
Published: (2024)
by: Yang, Shuai, et al.
Published: (2024)
Addressing the ID-Matching Challenge in Long Video Captioning
by: Yang, Zhantao, et al.
Published: (2025)
by: Yang, Zhantao, et al.
Published: (2025)
Story-Iter: A Training-free Iterative Paradigm for Long Story Visualization
by: Mao, Jiawei, et al.
Published: (2024)
by: Mao, Jiawei, et al.
Published: (2024)
VITRIX-CLIPIN: Enhancing Fine-Grained Visual Understanding in CLIP via Instruction Editing Data and Long Captions
by: Wang, Ziteng, et al.
Published: (2025)
by: Wang, Ziteng, et al.
Published: (2025)
DreamLIP: Language-Image Pre-training with Long Captions
by: Zheng, Kecheng, et al.
Published: (2024)
by: Zheng, Kecheng, et al.
Published: (2024)
Video ReCap: Recursive Captioning of Hour-Long Videos
by: Islam, Md Mohaiminul, et al.
Published: (2024)
by: Islam, Md Mohaiminul, et al.
Published: (2024)
Stateful Token Reduction for Long-Video Hybrid VLMs
by: Jiang, Jindong, et al.
Published: (2026)
by: Jiang, Jindong, et al.
Published: (2026)
Aligning Forest and Trees in Images & Long Captions for Visually Grounded Understanding
by: Woo, Byeongju, et al.
Published: (2026)
by: Woo, Byeongju, et al.
Published: (2026)
Fine-Grained Captioning of Long Videos through Scene Graph Consolidation
by: Chu, Sanghyeok, et al.
Published: (2025)
by: Chu, Sanghyeok, et al.
Published: (2025)
OVFact: Measuring and Improving Open-Vocabulary Factuality for Long Caption Models
by: Wysoczańska, Monika, et al.
Published: (2025)
by: Wysoczańska, Monika, et al.
Published: (2025)
Frame-Level Captions for Long Video Generation with Complex Multi Scenes
by: Zheng, Guangcong, et al.
Published: (2025)
by: Zheng, Guangcong, et al.
Published: (2025)
ReDiStory: Region-Disentangled Diffusion for Consistent Visual Story Generation
by: Sarkar, Ayushman, et al.
Published: (2026)
by: Sarkar, Ayushman, et al.
Published: (2026)
StoryMem: Multi-shot Long Video Storytelling with Memory
by: Zhang, Kaiwen, et al.
Published: (2025)
by: Zhang, Kaiwen, et al.
Published: (2025)
Seeing What Tastes Good: Revisiting Multimodal Distributional Semantics in the Billion Parameter Era
by: Oneata, Dan, et al.
Published: (2025)
by: Oneata, Dan, et al.
Published: (2025)
Causal Disentanglement for Robust Long-tail Medical Image Generation
by: Nie, Weizhi, et al.
Published: (2025)
by: Nie, Weizhi, et al.
Published: (2025)
Aligned Contrastive Loss for Long-Tailed Recognition
by: Ma, Jiali, et al.
Published: (2025)
by: Ma, Jiali, et al.
Published: (2025)
Decoupled Contrastive Learning for Long-Tailed Recognition
by: Xuan, Shiyu, et al.
Published: (2024)
by: Xuan, Shiyu, et al.
Published: (2024)
MiraData: A Large-Scale Video Dataset with Long Durations and Structured Captions
by: Ju, Xuan, et al.
Published: (2024)
by: Ju, Xuan, et al.
Published: (2024)
MAMBA4D: Efficient Long-Sequence Point Cloud Video Understanding with Disentangled Spatial-Temporal State Space Models
by: Liu, Jiuming, et al.
Published: (2024)
by: Liu, Jiuming, et al.
Published: (2024)
Similar Items
-
SPECS: Specificity-Enhanced CLIP-Score for Long Image Caption Evaluation
by: Chen, Xiaofu, et al.
Published: (2025) -
The Devil is in the EOS: Sequence Training for Detailed Image Captioning
by: Mohamed, Abdelrahman, et al.
Published: (2025) -
A Simple Data Augmentation Strategy for Text-in-Image Scientific VQA
by: Shoer, Belal, et al.
Published: (2025) -
EFSA: Episodic Few-Shot Adaptation for Text-to-Image Retrieval
by: Huzaifa, Muhammad, et al.
Published: (2024) -
VLMs Need Words: Vision Language Models Ignore Visual Detail In Favor of Semantic Anchors
by: Shahgir, Haz Sameen, et al.
Published: (2026)