Controllable Contextualized Image Captioning: Directing the Visual Narrative through User-Defined Highlights
Fuente:
arXiv
Saved in:
| Main Authors: | Mao, Shunqi, Zhang, Chaoyi, Su, Hang, Song, Hwanjun, Shalyminov, Igor, Cai, Weidong |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Through the Magnifying Glass: Adaptive Perception Magnification for Hallucination-Free VLM Decoding
by: Mao, Shunqi, et al.
Published: (2025)
by: Mao, Shunqi, et al.
Published: (2025)
Ctrl-Z Sampling: Scaling Diffusion Sampling with Controlled Random Zigzag Explorations
by: Mao, Shunqi, et al.
Published: (2025)
by: Mao, Shunqi, et al.
Published: (2025)
Exploring Annotation-free Image Captioning with Retrieval-augmented Pseudo Sentence Generation
by: Li, Zhiyuan, et al.
Published: (2023)
by: Li, Zhiyuan, et al.
Published: (2023)
The Collapse of Patches
by: Guo, Wei, et al.
Published: (2025)
by: Guo, Wei, et al.
Published: (2025)
Beyond Random Masking: A Dual-Stream Approach for Rotation-Invariant Point Cloud Masked Autoencoders
by: Yin, Xuanhua, et al.
Published: (2025)
by: Yin, Xuanhua, et al.
Published: (2025)
Reasoning over Video: Evaluating How MLLMs Extract, Integrate, and Reconstruct Spatiotemporal Evidence
by: Bang, Seunghwan, et al.
Published: (2026)
by: Bang, Seunghwan, et al.
Published: (2026)
Gene-DML: Dual-Pathway Multi-Level Discrimination for Gene Expression Prediction from Histopathology Images
by: Song, Yaxuan, et al.
Published: (2025)
by: Song, Yaxuan, et al.
Published: (2025)
The Quadratic Geometry of Flow Matching: Semantic Granularity Alignment for Text-to-Image Synthesis
by: Xiong, Zhinan, et al.
Published: (2026)
by: Xiong, Zhinan, et al.
Published: (2026)
DeepIcon: A Hierarchical Network for Layer-wise Icon Vectorization
by: Bing, Qi, et al.
Published: (2024)
by: Bing, Qi, et al.
Published: (2024)
PaRot: Patch-Wise Rotation-Invariant Network via Feature Disentanglement and Pose Restoration
by: Zhang, Dingxin, et al.
Published: (2023)
by: Zhang, Dingxin, et al.
Published: (2023)
From Image Captioning to Visual Storytelling
by: Passadakis, Admitos, et al.
Published: (2025)
by: Passadakis, Admitos, et al.
Published: (2025)
Learning to Synthesize Graphics Programs for Geometric Artworks
by: Bing, Qi, et al.
Published: (2024)
by: Bing, Qi, et al.
Published: (2024)
Dual-path Collaborative Generation Network for Emotional Video Captioning
by: Ye, Cheng, et al.
Published: (2024)
by: Ye, Cheng, et al.
Published: (2024)
Enhancing Advanced Visual Reasoning Ability of Large Language Models
by: Li, Zhiyuan, et al.
Published: (2024)
by: Li, Zhiyuan, et al.
Published: (2024)
CultureCLIP: Empowering CLIP with Cultural Awareness through Synthetic Images and Contextualized Captions
by: Huang, Yuchen, et al.
Published: (2025)
by: Huang, Yuchen, et al.
Published: (2025)
CaptionSmiths: Flexibly Controlling Language Pattern in Image Captioning
by: Saito, Kuniaki, et al.
Published: (2025)
by: Saito, Kuniaki, et al.
Published: (2025)
DEVICE: Depth and Visual Concepts Aware Transformer for OCR-based Image Captioning
by: Xu, Dongsheng, et al.
Published: (2023)
by: Xu, Dongsheng, et al.
Published: (2023)
FACE-net: Factual Calibration and Emotion Augmentation for Retrieval-enhanced Emotional Video Captioning
by: Chen, Weidong, et al.
Published: (2026)
by: Chen, Weidong, et al.
Published: (2026)
Visually-Aware Context Modeling for News Image Captioning
by: Qu, Tingyu, et al.
Published: (2023)
by: Qu, Tingyu, et al.
Published: (2023)
GuiDINO: Rethinking Vision Foundation Model in Medical Image Segmentation
by: Liang, Zhuonan, et al.
Published: (2026)
by: Liang, Zhuonan, et al.
Published: (2026)
Enhancing Visual Question Answering through Question-Driven Image Captions as Prompts
by: Özdemir, Övgü, et al.
Published: (2024)
by: Özdemir, Övgü, et al.
Published: (2024)
Enhancing Robustness to Noise Corruption for Point Cloud Recognition via Spatial Sorting and Set-Mixing Aggregation Module
by: Zhang, Dingxin, et al.
Published: (2024)
by: Zhang, Dingxin, et al.
Published: (2024)
FineSurE: Fine-grained Summarization Evaluation using LLMs
by: Song, Hwanjun, et al.
Published: (2024)
by: Song, Hwanjun, et al.
Published: (2024)
User-Aware Prefix-Tuning is a Good Learner for Personalized Image Captioning
by: Wang, Xuan, et al.
Published: (2023)
by: Wang, Xuan, et al.
Published: (2023)
RxnCaption: Reformulating Reaction Diagram Parsing as Visual Prompt Guided Captioning
by: Song, Jiahe, et al.
Published: (2025)
by: Song, Jiahe, et al.
Published: (2025)
Continual Learning for Image Captioning through Improved Image-Text Alignment
by: Taetz, Bertram, et al.
Published: (2025)
by: Taetz, Bertram, et al.
Published: (2025)
Bi-directional Contextual Attention for 3D Dense Captioning
by: Kim, Minjung, et al.
Published: (2024)
by: Kim, Minjung, et al.
Published: (2024)
Improving Text Generation on Images with Synthetic Captions
by: Koh, Jun Young, et al.
Published: (2024)
by: Koh, Jun Young, et al.
Published: (2024)
Contextual AD Narration with Interleaved Multimodal Sequence
by: Wang, Hanlin, et al.
Published: (2024)
by: Wang, Hanlin, et al.
Published: (2024)
MetaCaptioner: Towards Generalist Visual Captioning with Open-source Suites
by: Lei, Zhenxin, et al.
Published: (2025)
by: Lei, Zhenxin, et al.
Published: (2025)
CaptionQA: Is Your Caption as Useful as the Image Itself?
by: Yang, Shijia, et al.
Published: (2025)
by: Yang, Shijia, et al.
Published: (2025)
Progressive Text-to-Image Diffusion with Soft Latent Direction
by: Ye, YuTeng, et al.
Published: (2023)
by: Ye, YuTeng, et al.
Published: (2023)
Top-Down Framework for Weakly-supervised Grounded Image Captioning
by: Cai, Chen, et al.
Published: (2023)
by: Cai, Chen, et al.
Published: (2023)
Learning to Generalize over Subpartitions for Heterogeneity-aware Domain Adaptive Nuclei Segmentation
by: Fan, Jianan, et al.
Published: (2024)
by: Fan, Jianan, et al.
Published: (2024)
NarrativeBridge: Enhancing Video Captioning with Causal-Temporal Narrative
by: Nadeem, Asmar, et al.
Published: (2024)
by: Nadeem, Asmar, et al.
Published: (2024)
Edit As You Wish: Video Caption Editing with Multi-grained User Control
by: Yao, Linli, et al.
Published: (2023)
by: Yao, Linli, et al.
Published: (2023)
Exploiting Structural Consistency of Chest Anatomy for Unsupervised Anomaly Detection in Radiography Images
by: Xiang, Tiange, et al.
Published: (2024)
by: Xiang, Tiange, et al.
Published: (2024)
ECCV Caption: Correcting False Negatives by Collecting Machine-and-Human-verified Image-Caption Associations for MS-COCO
by: Chun, Sanghyuk, et al.
Published: (2022)
by: Chun, Sanghyuk, et al.
Published: (2022)
KFFocus: Highlighting Keyframes for Enhanced Video Understanding
by: Nie, Ming, et al.
Published: (2025)
by: Nie, Ming, et al.
Published: (2025)
HICEScore: A Hierarchical Metric for Image Captioning Evaluation
by: Zeng, Zequn, et al.
Published: (2024)
by: Zeng, Zequn, et al.
Published: (2024)
Similar Items
-
Through the Magnifying Glass: Adaptive Perception Magnification for Hallucination-Free VLM Decoding
by: Mao, Shunqi, et al.
Published: (2025) -
Ctrl-Z Sampling: Scaling Diffusion Sampling with Controlled Random Zigzag Explorations
by: Mao, Shunqi, et al.
Published: (2025) -
Exploring Annotation-free Image Captioning with Retrieval-augmented Pseudo Sentence Generation
by: Li, Zhiyuan, et al.
Published: (2023) -
The Collapse of Patches
by: Guo, Wei, et al.
Published: (2025) -
Beyond Random Masking: A Dual-Stream Approach for Rotation-Invariant Point Cloud Masked Autoencoders
by: Yin, Xuanhua, et al.
Published: (2025)