Pretrained Image-Text Models are Secretly Video Captioners
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Zhang, Chunhui, Jian, Yiren, Ouyang, Zhongyu, Vosoughi, Soroush |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Expedited Training of Visual Conditioned Language Generation via Redundancy Reduction
von: Jian, Yiren, et al.
Veröffentlicht: (2023)
von: Jian, Yiren, et al.
Veröffentlicht: (2023)
Text-to-Image GAN with Pretrained Representations
von: You, Xiaozhou, et al.
Veröffentlicht: (2024)
von: You, Xiaozhou, et al.
Veröffentlicht: (2024)
Semantic Compositions Enhance Vision-Language Contrastive Learning
von: Aladago, Maxwell, et al.
Veröffentlicht: (2024)
von: Aladago, Maxwell, et al.
Veröffentlicht: (2024)
Image Captions are Natural Prompts for Text-to-Image Models
von: Lei, Shiye, et al.
Veröffentlicht: (2023)
von: Lei, Shiye, et al.
Veröffentlicht: (2023)
Learning to Rank Caption Chains for Video-Text Alignment
von: Blume, Ansel, et al.
Veröffentlicht: (2026)
von: Blume, Ansel, et al.
Veröffentlicht: (2026)
Is Large-Scale Pretraining the Secret to Good Domain Generalization?
von: Teterwak, Piotr, et al.
Veröffentlicht: (2024)
von: Teterwak, Piotr, et al.
Veröffentlicht: (2024)
CLIP with Quality Captions: A Strong Pretraining for Vision Tasks
von: Vasu, Pavan Kumar Anasosalu, et al.
Veröffentlicht: (2024)
von: Vasu, Pavan Kumar Anasosalu, et al.
Veröffentlicht: (2024)
Your Image is Secretly the Last Frame of a Pseudo Video
von: Chen, Wenlong, et al.
Veröffentlicht: (2024)
von: Chen, Wenlong, et al.
Veröffentlicht: (2024)
Modeling Image-Caption Rating from Comparative Judgments
von: Minni, Kezia, et al.
Veröffentlicht: (2026)
von: Minni, Kezia, et al.
Veröffentlicht: (2026)
Modeling Caption Diversity in Contrastive Vision-Language Pretraining
von: Lavoie, Samuel, et al.
Veröffentlicht: (2024)
von: Lavoie, Samuel, et al.
Veröffentlicht: (2024)
Structured Captions Improve Prompt Adherence in Text-to-Image Models (Re-LAION-Caption 19M)
von: Merchant, Nicholas, et al.
Veröffentlicht: (2025)
von: Merchant, Nicholas, et al.
Veröffentlicht: (2025)
Unleashing Text-to-Image Diffusion Prior for Zero-Shot Image Captioning
von: Luo, Jianjie, et al.
Veröffentlicht: (2024)
von: Luo, Jianjie, et al.
Veröffentlicht: (2024)
TALC: Time-Aligned Captions for Multi-Scene Text-to-Video Generation
von: Bansal, Hritik, et al.
Veröffentlicht: (2024)
von: Bansal, Hritik, et al.
Veröffentlicht: (2024)
Foundation Model-oriented Robustness: Robust Image Model Evaluation with Pretrained Models
von: Zhang, Peiyan, et al.
Veröffentlicht: (2023)
von: Zhang, Peiyan, et al.
Veröffentlicht: (2023)
OSCaR: Object State Captioning and State Change Representation
von: Nguyen, Nguyen, et al.
Veröffentlicht: (2024)
von: Nguyen, Nguyen, et al.
Veröffentlicht: (2024)
Fine-Grained Alignment and Noise Refinement for Compositional Text-to-Image Generation
von: Izadi, Amir Mohammad, et al.
Veröffentlicht: (2025)
von: Izadi, Amir Mohammad, et al.
Veröffentlicht: (2025)
Temporal Working Memory: Query-Guided Segment Refinement for Enhanced Multimodal Understanding
von: Diao, Xingjian, et al.
Veröffentlicht: (2025)
von: Diao, Xingjian, et al.
Veröffentlicht: (2025)
Scaled Supervision is an Implicit Lipschitz Regularizer
von: Ouyang, Zhongyu, et al.
Veröffentlicht: (2025)
von: Ouyang, Zhongyu, et al.
Veröffentlicht: (2025)
Pixels to Prose: Understanding the art of Image Captioning
von: Singh, Hrishikesh, et al.
Veröffentlicht: (2024)
von: Singh, Hrishikesh, et al.
Veröffentlicht: (2024)
Recurrent Attention-based Token Selection for Efficient Streaming Video-LLMs
von: Dorovatas, Vaggelis, et al.
Veröffentlicht: (2025)
von: Dorovatas, Vaggelis, et al.
Veröffentlicht: (2025)
Infusing Environmental Captions for Long-Form Video Language Grounding
von: Lee, Hyogun, et al.
Veröffentlicht: (2024)
von: Lee, Hyogun, et al.
Veröffentlicht: (2024)
How to Train your Text-to-Image Model: Evaluating Design Choices for Synthetic Training Captions
von: Brack, Manuel, et al.
Veröffentlicht: (2025)
von: Brack, Manuel, et al.
Veröffentlicht: (2025)
Linear Alignment of Vision-language Models for Image Captioning
von: Paischer, Fabian, et al.
Veröffentlicht: (2023)
von: Paischer, Fabian, et al.
Veröffentlicht: (2023)
IMAGINE-E: Image Generation Intelligence Evaluation of State-of-the-art Text-to-Image Models
von: Lei, Jiayi, et al.
Veröffentlicht: (2025)
von: Lei, Jiayi, et al.
Veröffentlicht: (2025)
Differentially Private Representation Learning via Image Captioning
von: Sander, Tom, et al.
Veröffentlicht: (2024)
von: Sander, Tom, et al.
Veröffentlicht: (2024)
Image Clustering via the Principle of Rate Reduction in the Age of Pretrained Models
von: Chu, Tianzhe, et al.
Veröffentlicht: (2023)
von: Chu, Tianzhe, et al.
Veröffentlicht: (2023)
Temporal Object Captioning for Street Scene Videos from LiDAR Tracks
von: Gopinathan, Vignesh, et al.
Veröffentlicht: (2025)
von: Gopinathan, Vignesh, et al.
Veröffentlicht: (2025)
ICC: Quantifying Image Caption Concreteness for Multimodal Dataset Curation
von: Yanuka, Moran, et al.
Veröffentlicht: (2024)
von: Yanuka, Moran, et al.
Veröffentlicht: (2024)
Contextualized Diffusion Models for Text-Guided Image and Video Generation
von: Yang, Ling, et al.
Veröffentlicht: (2024)
von: Yang, Ling, et al.
Veröffentlicht: (2024)
Whats in a Video: Factorized Autoregressive Decoding for Online Dense Video Captioning
von: Piergiovanni, AJ, et al.
Veröffentlicht: (2024)
von: Piergiovanni, AJ, et al.
Veröffentlicht: (2024)
Data Attribution for Text-to-Image Models by Unlearning Synthesized Images
von: Wang, Sheng-Yu, et al.
Veröffentlicht: (2024)
von: Wang, Sheng-Yu, et al.
Veröffentlicht: (2024)
STAND: Semantic Anchoring Constraint with Dual-Granularity Disambiguation for Remote Sensing Image Change Captioning
von: Gong, Yanpei, et al.
Veröffentlicht: (2026)
von: Gong, Yanpei, et al.
Veröffentlicht: (2026)
FonTS: Text Rendering with Typography and Style Controls
von: Shi, Wenda, et al.
Veröffentlicht: (2024)
von: Shi, Wenda, et al.
Veröffentlicht: (2024)
Generalizable Geometric Image Caption Synthesis
von: Xin, Yue, et al.
Veröffentlicht: (2025)
von: Xin, Yue, et al.
Veröffentlicht: (2025)
Fast Data Attribution for Text-to-Image Models
von: Wang, Sheng-Yu, et al.
Veröffentlicht: (2025)
von: Wang, Sheng-Yu, et al.
Veröffentlicht: (2025)
Wolf: Dense Video Captioning with a World Summarization Framework
von: Li, Boyi, et al.
Veröffentlicht: (2024)
von: Li, Boyi, et al.
Veröffentlicht: (2024)
Stochastic Siamese MAE Pretraining for Longitudinal Medical Images
von: Emre, Taha, et al.
Veröffentlicht: (2025)
von: Emre, Taha, et al.
Veröffentlicht: (2025)
Detecting Backdoor Samples in Contrastive Language Image Pretraining
von: Huang, Hanxun, et al.
Veröffentlicht: (2025)
von: Huang, Hanxun, et al.
Veröffentlicht: (2025)
Time-to-Event Pretraining for 3D Medical Imaging
von: Huo, Zepeng, et al.
Veröffentlicht: (2024)
von: Huo, Zepeng, et al.
Veröffentlicht: (2024)
PMPGuard: Catching Pseudo-Matched Pairs in Remote Sensing Image-Text Retrieval
von: Ouyang, Pengxiang, et al.
Veröffentlicht: (2025)
von: Ouyang, Pengxiang, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Expedited Training of Visual Conditioned Language Generation via Redundancy Reduction
von: Jian, Yiren, et al.
Veröffentlicht: (2023) -
Text-to-Image GAN with Pretrained Representations
von: You, Xiaozhou, et al.
Veröffentlicht: (2024) -
Semantic Compositions Enhance Vision-Language Contrastive Learning
von: Aladago, Maxwell, et al.
Veröffentlicht: (2024) -
Image Captions are Natural Prompts for Text-to-Image Models
von: Lei, Shiye, et al.
Veröffentlicht: (2023) -
Learning to Rank Caption Chains for Video-Text Alignment
von: Blume, Ansel, et al.
Veröffentlicht: (2026)