Masked Diffusion Captioning for Visual Feature Learning
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Feng, Chao, Wei, Zihao, Owens, Andrew |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Efficient Vision-Language Pre-training by Cluster Masking
von: Wei, Zihao, et al.
Veröffentlicht: (2024)
von: Wei, Zihao, et al.
Veröffentlicht: (2024)
Visual Anagrams: Generating Multi-View Optical Illusions with Diffusion Models
von: Geng, Daniel, et al.
Veröffentlicht: (2023)
von: Geng, Daniel, et al.
Veröffentlicht: (2023)
New Encoder Learning for Captioning Heavy Rain Images via Semantic Visual Feature Matching
von: Son, Chang-Hwan, et al.
Veröffentlicht: (2021)
von: Son, Chang-Hwan, et al.
Veröffentlicht: (2021)
Motion Guidance: Diffusion-Based Image Editing with Differentiable Motion Estimators
von: Geng, Daniel, et al.
Veröffentlicht: (2024)
von: Geng, Daniel, et al.
Veröffentlicht: (2024)
RxnCaption: Reformulating Reaction Diagram Parsing as Visual Prompt Guided Captioning
von: Song, Jiahe, et al.
Veröffentlicht: (2025)
von: Song, Jiahe, et al.
Veröffentlicht: (2025)
Q-Adapter: Visual Query Adapter for Extracting Textually-related Features in Video Captioning
von: Chen, Junan, et al.
Veröffentlicht: (2025)
von: Chen, Junan, et al.
Veröffentlicht: (2025)
D-ORCA: Dialogue-Centric Optimization for Robust Audio-Visual Captioning
von: Tang, Changli, et al.
Veröffentlicht: (2026)
von: Tang, Changli, et al.
Veröffentlicht: (2026)
MaskDiME: Adaptive Masked Diffusion for Precise and Efficient Visual Counterfactual Explanations
von: Guo, Changlu, et al.
Veröffentlicht: (2026)
von: Guo, Changlu, et al.
Veröffentlicht: (2026)
MetaCaptioner: Towards Generalist Visual Captioning with Open-source Suites
von: Lei, Zhenxin, et al.
Veröffentlicht: (2025)
von: Lei, Zhenxin, et al.
Veröffentlicht: (2025)
Factorized Diffusion: Perceptual Illusions by Noise Decomposition
von: Geng, Daniel, et al.
Veröffentlicht: (2024)
von: Geng, Daniel, et al.
Veröffentlicht: (2024)
MaskMatch: Boosting Semi-Supervised Learning Through Mask Autoencoder-Driven Feature Learning
von: Zhang, Wenjin, et al.
Veröffentlicht: (2024)
von: Zhang, Wenjin, et al.
Veröffentlicht: (2024)
TimeChat-Captioner: Scripting Multi-Scene Videos with Time-Aware and Structural Audio-Visual Captions
von: Yao, Linli, et al.
Veröffentlicht: (2026)
von: Yao, Linli, et al.
Veröffentlicht: (2026)
MV-CC: Mask Enhanced Video Model for Remote Sensing Change Caption
von: Liu, Ruixun, et al.
Veröffentlicht: (2024)
von: Liu, Ruixun, et al.
Veröffentlicht: (2024)
DEVICE: Depth and Visual Concepts Aware Transformer for OCR-based Image Captioning
von: Xu, Dongsheng, et al.
Veröffentlicht: (2023)
von: Xu, Dongsheng, et al.
Veröffentlicht: (2023)
Dual Caption Preference Optimization for Diffusion Models
von: Saeidi, Amir, et al.
Veröffentlicht: (2025)
von: Saeidi, Amir, et al.
Veröffentlicht: (2025)
video-SALMONN 2: Caption-Enhanced Audio-Visual Large Language Models
von: Tang, Changli, et al.
Veröffentlicht: (2025)
von: Tang, Changli, et al.
Veröffentlicht: (2025)
From Image Captioning to Visual Storytelling
von: Passadakis, Admitos, et al.
Veröffentlicht: (2025)
von: Passadakis, Admitos, et al.
Veröffentlicht: (2025)
Point Prompting: Counterfactual Tracking with Video Diffusion Models
von: Shrivastava, Ayush, et al.
Veröffentlicht: (2025)
von: Shrivastava, Ayush, et al.
Veröffentlicht: (2025)
GPS as a Control Signal for Image Generation
von: Feng, Chao, et al.
Veröffentlicht: (2025)
von: Feng, Chao, et al.
Veröffentlicht: (2025)
Visually-Aware Context Modeling for News Image Captioning
von: Qu, Tingyu, et al.
Veröffentlicht: (2023)
von: Qu, Tingyu, et al.
Veröffentlicht: (2023)
LocCa: Visual Pretraining with Location-aware Captioners
von: Wan, Bo, et al.
Veröffentlicht: (2024)
von: Wan, Bo, et al.
Veröffentlicht: (2024)
Enhancing Descriptive Captions with Visual Attributes for Multimodal Perception
von: Sun, Yanpeng, et al.
Veröffentlicht: (2024)
von: Sun, Yanpeng, et al.
Veröffentlicht: (2024)
SC-Captioner: Improving Image Captioning with Self-Correction by Reinforcement Learning
von: Zhang, Lin, et al.
Veröffentlicht: (2025)
von: Zhang, Lin, et al.
Veröffentlicht: (2025)
Diffusion Based Augmentation for Captioning and Retrieval in Cultural Heritage
von: Cioni, Dario, et al.
Veröffentlicht: (2023)
von: Cioni, Dario, et al.
Veröffentlicht: (2023)
Enhancing Prompt Following with Visual Control Through Training-Free Mask-Guided Diffusion
von: Chen, Hongyu, et al.
Veröffentlicht: (2024)
von: Chen, Hongyu, et al.
Veröffentlicht: (2024)
Evolved Hierarchical Masking for Self-Supervised Learning
von: Feng, Zhanzhou, et al.
Veröffentlicht: (2025)
von: Feng, Zhanzhou, et al.
Veröffentlicht: (2025)
Masked Generative Story Transformer with Character Guidance and Caption Augmentation
von: Papadimitriou, Christos, et al.
Veröffentlicht: (2024)
von: Papadimitriou, Christos, et al.
Veröffentlicht: (2024)
Coefficients-Preserving Sampling for Reinforcement Learning with Flow Matching
von: Wang, Feng, et al.
Veröffentlicht: (2025)
von: Wang, Feng, et al.
Veröffentlicht: (2025)
Unleashing Text-to-Image Diffusion Prior for Zero-Shot Image Captioning
von: Luo, Jianjie, et al.
Veröffentlicht: (2024)
von: Luo, Jianjie, et al.
Veröffentlicht: (2024)
Analyzing Image Beyond Visual Aspect: Image Emotion Classification via Multiple-Affective Captioning
von: Zhou, Zibo, et al.
Veröffentlicht: (2025)
von: Zhou, Zibo, et al.
Veröffentlicht: (2025)
Frequency-Guided Masking for Enhanced Vision Self-Supervised Learning
von: Monsefi, Amin Karimi, et al.
Veröffentlicht: (2024)
von: Monsefi, Amin Karimi, et al.
Veröffentlicht: (2024)
Learning Speaker-Invariant Visual Features for Lipreading
von: Li, Yu, et al.
Veröffentlicht: (2025)
von: Li, Yu, et al.
Veröffentlicht: (2025)
Mask$^2$DiT: Dual Mask-based Diffusion Transformer for Multi-Scene Long Video Generation
von: Qi, Tianhao, et al.
Veröffentlicht: (2025)
von: Qi, Tianhao, et al.
Veröffentlicht: (2025)
Caption Anything in Video: Fine-grained Object-centric Captioning via Spatiotemporal Multimodal Prompting
von: Tang, Yunlong, et al.
Veröffentlicht: (2025)
von: Tang, Yunlong, et al.
Veröffentlicht: (2025)
View Selection for 3D Captioning via Diffusion Ranking
von: Luo, Tiange, et al.
Veröffentlicht: (2024)
von: Luo, Tiange, et al.
Veröffentlicht: (2024)
MaskSem: Semantic-Guided Masking for Learning 3D Hybrid High-Order Motion Representation
von: Wei, Wei, et al.
Veröffentlicht: (2025)
von: Wei, Wei, et al.
Veröffentlicht: (2025)
Diff-Prompt: Diffusion-Driven Prompt Generator with Mask Supervision
von: Yan, Weicai, et al.
Veröffentlicht: (2025)
von: Yan, Weicai, et al.
Veröffentlicht: (2025)
LongCaptioning: Unlocking the Power of Long Video Caption Generation in Large Multimodal Models
von: Wei, Hongchen, et al.
Veröffentlicht: (2025)
von: Wei, Hongchen, et al.
Veröffentlicht: (2025)
Sample-specific Masks for Visual Reprogramming-based Prompting
von: Cai, Chengyi, et al.
Veröffentlicht: (2024)
von: Cai, Chengyi, et al.
Veröffentlicht: (2024)
LazyMAR: Accelerating Masked Autoregressive Models via Feature Caching
von: Yan, Feihong, et al.
Veröffentlicht: (2025)
von: Yan, Feihong, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Efficient Vision-Language Pre-training by Cluster Masking
von: Wei, Zihao, et al.
Veröffentlicht: (2024) -
Visual Anagrams: Generating Multi-View Optical Illusions with Diffusion Models
von: Geng, Daniel, et al.
Veröffentlicht: (2023) -
New Encoder Learning for Captioning Heavy Rain Images via Semantic Visual Feature Matching
von: Son, Chang-Hwan, et al.
Veröffentlicht: (2021) -
Motion Guidance: Diffusion-Based Image Editing with Differentiable Motion Estimators
von: Geng, Daniel, et al.
Veröffentlicht: (2024) -
RxnCaption: Reformulating Reaction Diagram Parsing as Visual Prompt Guided Captioning
von: Song, Jiahe, et al.
Veröffentlicht: (2025)