Learning to Rank Caption Chains for Video-Text Alignment
Fuente:
arXiv
Saved in:
| Main Authors: | Blume, Ansel, Uzkent, Burak, Chaudhuri, Shalini, Kessler, Garin |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Narrative Aligned Long Form Video Question Answering
by: Jain, Rahul, et al.
Published: (2026)
by: Jain, Rahul, et al.
Published: (2026)
From Frames to Clips: Training-free Adaptive Key Clip Selection for Long-Form Video Understanding
by: Sun, Guangyu, et al.
Published: (2025)
by: Sun, Guangyu, et al.
Published: (2025)
Pretrained Image-Text Models are Secretly Video Captioners
by: Zhang, Chunhui, et al.
Published: (2025)
by: Zhang, Chunhui, et al.
Published: (2025)
Perceptio: Perception Enhanced Vision Language Models via Spatial Token Generation
by: Li, Yuchen, et al.
Published: (2026)
by: Li, Yuchen, et al.
Published: (2026)
Differentially Private Representation Learning via Image Captioning
by: Sander, Tom, et al.
Published: (2024)
by: Sander, Tom, et al.
Published: (2024)
CounterVid: Counterfactual Video Generation for Mitigating Action and Temporal Hallucinations in Video-Language Models
by: Poppi, Tobia, et al.
Published: (2026)
by: Poppi, Tobia, et al.
Published: (2026)
TALC: Time-Aligned Captions for Multi-Scene Text-to-Video Generation
by: Bansal, Hritik, et al.
Published: (2024)
by: Bansal, Hritik, et al.
Published: (2024)
Linear Alignment of Vision-language Models for Image Captioning
by: Paischer, Fabian, et al.
Published: (2023)
by: Paischer, Fabian, et al.
Published: (2023)
Infusing Environmental Captions for Long-Form Video Language Grounding
by: Lee, Hyogun, et al.
Published: (2024)
by: Lee, Hyogun, et al.
Published: (2024)
Feedback Alignment Meets Low-Rank Manifolds: A Structured Recipe for Local Learning
by: Roy, Arani, et al.
Published: (2025)
by: Roy, Arani, et al.
Published: (2025)
Temporal Object Captioning for Street Scene Videos from LiDAR Tracks
by: Gopinathan, Vignesh, et al.
Published: (2025)
by: Gopinathan, Vignesh, et al.
Published: (2025)
Whats in a Video: Factorized Autoregressive Decoding for Online Dense Video Captioning
by: Piergiovanni, AJ, et al.
Published: (2024)
by: Piergiovanni, AJ, et al.
Published: (2024)
Text-centric Alignment for Multi-Modality Learning
by: Tsai, Yun-Da, et al.
Published: (2024)
by: Tsai, Yun-Da, et al.
Published: (2024)
Learning Conditional Invariances through Non-Commutativity
by: Chaudhuri, Abhra, et al.
Published: (2024)
by: Chaudhuri, Abhra, et al.
Published: (2024)
Information Theoretic Text-to-Image Alignment
by: Wang, Chao, et al.
Published: (2024)
by: Wang, Chao, et al.
Published: (2024)
Wolf: Dense Video Captioning with a World Summarization Framework
by: Li, Boyi, et al.
Published: (2024)
by: Li, Boyi, et al.
Published: (2024)
Cycle Consistency as Reward: Learning Image-Text Alignment without Human Preferences
by: Bahng, Hyojin, et al.
Published: (2025)
by: Bahng, Hyojin, et al.
Published: (2025)
Learn to Rank: Visual Attribution by Learning Importance Ranking
by: Schinagl, David, et al.
Published: (2026)
by: Schinagl, David, et al.
Published: (2026)
Image Captions are Natural Prompts for Text-to-Image Models
by: Lei, Shiye, et al.
Published: (2023)
by: Lei, Shiye, et al.
Published: (2023)
Revised Regularization for Efficient Continual Learning through Correlation-Based Parameter Update in Bayesian Neural Networks
by: Palit, Sanchar, et al.
Published: (2024)
by: Palit, Sanchar, et al.
Published: (2024)
MTA: Multimodal Task Alignment for BEV Perception and Captioning
by: Ma, Yunsheng, et al.
Published: (2024)
by: Ma, Yunsheng, et al.
Published: (2024)
QG-CoC: Question-Guided Chain-of-Captions for Large Multimodal Models
by: Kao, Kuei-Chun, et al.
Published: (2025)
by: Kao, Kuei-Chun, et al.
Published: (2025)
Text Embedding is Not All You Need: Attention Control for Text-to-Image Semantic Alignment with Text Self-Attention Maps
by: Kim, Jeeyung, et al.
Published: (2024)
by: Kim, Jeeyung, et al.
Published: (2024)
Patch Ranking: Efficient CLIP by Learning to Rank Local Patches
by: Wu, Cheng-En, et al.
Published: (2024)
by: Wu, Cheng-En, et al.
Published: (2024)
Improving Long-Text Alignment for Text-to-Image Diffusion Models
by: Liu, Luping, et al.
Published: (2024)
by: Liu, Luping, et al.
Published: (2024)
Unleashing Text-to-Image Diffusion Prior for Zero-Shot Image Captioning
by: Luo, Jianjie, et al.
Published: (2024)
by: Luo, Jianjie, et al.
Published: (2024)
Structured Captions Improve Prompt Adherence in Text-to-Image Models (Re-LAION-Caption 19M)
by: Merchant, Nicholas, et al.
Published: (2025)
by: Merchant, Nicholas, et al.
Published: (2025)
Image Captioning as an Assistive Technology: Lessons Learned from VizWiz 2020 Challenge
by: Dognin, Pierre, et al.
Published: (2020)
by: Dognin, Pierre, et al.
Published: (2020)
Hyperdimensional Cross-Modal Alignment of Frozen Language and Image Models for Efficient Image Captioning
by: Dalvi, Abhishek, et al.
Published: (2026)
by: Dalvi, Abhishek, et al.
Published: (2026)
Free$^2$Guide: Training-Free Text-to-Video Alignment using Image LVLM
by: Kim, Jaemin, et al.
Published: (2024)
by: Kim, Jaemin, et al.
Published: (2024)
RFMI: Estimating Mutual Information on Rectified Flow for Text-to-Image Alignment
by: Wang, Chao, et al.
Published: (2025)
by: Wang, Chao, et al.
Published: (2025)
Fine-Grained Alignment and Noise Refinement for Compositional Text-to-Image Generation
by: Izadi, Amir Mohammad, et al.
Published: (2025)
by: Izadi, Amir Mohammad, et al.
Published: (2025)
Mapping Land Naturalness from Sentinel-2 using Deep Contextual and Geographical Priors
by: Ekim, Burak, et al.
Published: (2024)
by: Ekim, Burak, et al.
Published: (2024)
VIRL: Volume-Informed Representation Learning towards Few-shot Manufacturability Estimation
by: Chen, Yu-hsuan, et al.
Published: (2024)
by: Chen, Yu-hsuan, et al.
Published: (2024)
Large VLM-based Stylized Sports Captioning
by: Dhar, Sauptik, et al.
Published: (2025)
by: Dhar, Sauptik, et al.
Published: (2025)
Pixels to Prose: Understanding the art of Image Captioning
by: Singh, Hrishikesh, et al.
Published: (2024)
by: Singh, Hrishikesh, et al.
Published: (2024)
Improving Image Captioning Descriptiveness by Ranking and LLM-based Fusion
by: Celona, Luigi, et al.
Published: (2023)
by: Celona, Luigi, et al.
Published: (2023)
Contrast-Aware Calibration for Fine-Tuned CLIP: Leveraging Image-Text Alignment
by: Lv, Song-Lin, et al.
Published: (2025)
by: Lv, Song-Lin, et al.
Published: (2025)
Corruption-Aware Training of Latent Video Diffusion Models for Robust Text-to-Video Generation
by: Maduabuchi, Chika, et al.
Published: (2025)
by: Maduabuchi, Chika, et al.
Published: (2025)
TA-Prompting: Enhancing Video Large Language Models for Dense Video Captioning via Temporal Anchors
by: Cheng, Wei-Yuan, et al.
Published: (2026)
by: Cheng, Wei-Yuan, et al.
Published: (2026)
Similar Items
-
Narrative Aligned Long Form Video Question Answering
by: Jain, Rahul, et al.
Published: (2026) -
From Frames to Clips: Training-free Adaptive Key Clip Selection for Long-Form Video Understanding
by: Sun, Guangyu, et al.
Published: (2025) -
Pretrained Image-Text Models are Secretly Video Captioners
by: Zhang, Chunhui, et al.
Published: (2025) -
Perceptio: Perception Enhanced Vision Language Models via Spatial Token Generation
by: Li, Yuchen, et al.
Published: (2026) -
Differentially Private Representation Learning via Image Captioning
by: Sander, Tom, et al.
Published: (2024)