Predicting Winning Captions for Weekly New Yorker Comics
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Cao, Stanley, Young, Sonny |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Figuring out Figures: Using Textual References to Caption Scientific Figures
von: Cao, Stanley, et al.
Veröffentlicht: (2024)
von: Cao, Stanley, et al.
Veröffentlicht: (2024)
Zooming into Comics: Region-Aware RL Improves Fine-Grained Comic Understanding in Vision-Language Models
von: Chen, Yule, et al.
Veröffentlicht: (2025)
von: Chen, Yule, et al.
Veröffentlicht: (2025)
Towards Faithful Reasoning in Comics for Small MLLMs
von: Feng, Chengcheng, et al.
Veröffentlicht: (2026)
von: Feng, Chengcheng, et al.
Veröffentlicht: (2026)
CaptionFool: Universal Image Captioning Model Attacks
von: Parekh, Swapnil
Veröffentlicht: (2026)
von: Parekh, Swapnil
Veröffentlicht: (2026)
AGIC: Attention-Guided Image Captioning to Improve Caption Relevance
von: Teja, L. D. M. S. Sai, et al.
Veröffentlicht: (2025)
von: Teja, L. D. M. S. Sai, et al.
Veröffentlicht: (2025)
Any2Caption:Interpreting Any Condition to Caption for Controllable Video Generation
von: Wu, Shengqiong, et al.
Veröffentlicht: (2025)
von: Wu, Shengqiong, et al.
Veröffentlicht: (2025)
Culture In a Frame: C$^3$B as a Comic-Based Benchmark for Multimodal Culturally Awareness
von: Song, Yuchen, et al.
Veröffentlicht: (2025)
von: Song, Yuchen, et al.
Veröffentlicht: (2025)
Humor in Pixels: Benchmarking Large Multimodal Models Understanding of Online Comics
von: Ryan, Yuriel, et al.
Veröffentlicht: (2025)
von: Ryan, Yuriel, et al.
Veröffentlicht: (2025)
Image Captioning in news report scenario
von: Liu, Tianrui, et al.
Veröffentlicht: (2024)
von: Liu, Tianrui, et al.
Veröffentlicht: (2024)
Automated Image Captioning with CNNs and Transformers
von: Cahyono, Joshua Adrian, et al.
Veröffentlicht: (2024)
von: Cahyono, Joshua Adrian, et al.
Veröffentlicht: (2024)
Caption This, Reason That: VLMs Caught in the Middle
von: Weng, Zihan, et al.
Veröffentlicht: (2025)
von: Weng, Zihan, et al.
Veröffentlicht: (2025)
URECA: Unique Region Caption Anything
von: Lim, Sangbeom, et al.
Veröffentlicht: (2025)
von: Lim, Sangbeom, et al.
Veröffentlicht: (2025)
Accurate and Fast Compressed Video Captioning
von: Shen, Yaojie, et al.
Veröffentlicht: (2023)
von: Shen, Yaojie, et al.
Veröffentlicht: (2023)
Mitigating Open-Vocabulary Caption Hallucinations
von: Ben-Kish, Assaf, et al.
Veröffentlicht: (2023)
von: Ben-Kish, Assaf, et al.
Veröffentlicht: (2023)
Imagine How To Change: Explicit Procedure Modeling for Change Captioning
von: Sun, Jiayang, et al.
Veröffentlicht: (2026)
von: Sun, Jiayang, et al.
Veröffentlicht: (2026)
SAIL: Similarity-Aware Guidance and Inter-Caption Augmentation-based Learning for Weakly-Supervised Dense Video Captioning
von: Kim, Ye-Chan, et al.
Veröffentlicht: (2026)
von: Kim, Ye-Chan, et al.
Veröffentlicht: (2026)
Knowledge Completes the Vision: A Multimodal Entity-aware Retrieval-Augmented Generation Framework for News Image Captioning
von: You, Xiaoxing, et al.
Veröffentlicht: (2025)
von: You, Xiaoxing, et al.
Veröffentlicht: (2025)
Self-Explainable Affordance Learning with Embodied Caption
von: Zhang, Zhipeng, et al.
Veröffentlicht: (2024)
von: Zhang, Zhipeng, et al.
Veröffentlicht: (2024)
Image Embedding Sampling Method for Diverse Captioning
von: Waheed, Sania, et al.
Veröffentlicht: (2025)
von: Waheed, Sania, et al.
Veröffentlicht: (2025)
Top-Down Semantic Refinement for Image Captioning
von: Zhang, Jusheng, et al.
Veröffentlicht: (2025)
von: Zhang, Jusheng, et al.
Veröffentlicht: (2025)
Parrot Captions Teach CLIP to Spot Text
von: Lin, Yiqi, et al.
Veröffentlicht: (2023)
von: Lin, Yiqi, et al.
Veröffentlicht: (2023)
Target-Dependent Multimodal Sentiment Analysis Via Employing Visual-to Emotional-Caption Translation Network using Visual-Caption Pairs
von: Pandey, Ananya, et al.
Veröffentlicht: (2024)
von: Pandey, Ananya, et al.
Veröffentlicht: (2024)
WinT3R: Window-Based Streaming Reconstruction with Camera Token Pool
von: Li, Zizun, et al.
Veröffentlicht: (2025)
von: Li, Zizun, et al.
Veröffentlicht: (2025)
IG Captioner: Information Gain Captioners are Strong Zero-shot Classifiers
von: Yang, Chenglin, et al.
Veröffentlicht: (2023)
von: Yang, Chenglin, et al.
Veröffentlicht: (2023)
Is Your Text-to-Image Model Robust to Caption Noise?
von: Yu, Weichen, et al.
Veröffentlicht: (2024)
von: Yu, Weichen, et al.
Veröffentlicht: (2024)
ReflectCAP: Detailed Image Captioning with Reflective Memory
von: Min, Kyungmin, et al.
Veröffentlicht: (2026)
von: Min, Kyungmin, et al.
Veröffentlicht: (2026)
An Ensemble Model with Attention Based Mechanism for Image Captioning
von: Badarneh, Israa Al, et al.
Veröffentlicht: (2025)
von: Badarneh, Israa Al, et al.
Veröffentlicht: (2025)
Describe Anything: Detailed Localized Image and Video Captioning
von: Lian, Long, et al.
Veröffentlicht: (2025)
von: Lian, Long, et al.
Veröffentlicht: (2025)
Towards Fine-Grained Human Motion Video Captioning
von: Song, Guorui, et al.
Veröffentlicht: (2025)
von: Song, Guorui, et al.
Veröffentlicht: (2025)
Generating Accurate and Detailed Captions for High-Resolution Images
von: Lee, Hankyeol, et al.
Veröffentlicht: (2025)
von: Lee, Hankyeol, et al.
Veröffentlicht: (2025)
Leveraging Textual Compositional Reasoning for Robust Change Captioning
von: Park, Kyu Ri, et al.
Veröffentlicht: (2025)
von: Park, Kyu Ri, et al.
Veröffentlicht: (2025)
On Explaining Visual Captioning with Hybrid Markov Logic Networks
von: Shah, Monika, et al.
Veröffentlicht: (2025)
von: Shah, Monika, et al.
Veröffentlicht: (2025)
XMeCap: Meme Caption Generation with Sub-Image Adaptability
von: Chen, Yuyan, et al.
Veröffentlicht: (2024)
von: Chen, Yuyan, et al.
Veröffentlicht: (2024)
Uterine Ultrasound Image Captioning Using Deep Learning Techniques
von: Boulesnane, Abdennour, et al.
Veröffentlicht: (2024)
von: Boulesnane, Abdennour, et al.
Veröffentlicht: (2024)
KALE: An Artwork Image Captioning System Augmented with Heterogeneous Graph
von: Jiang, Yanbei, et al.
Veröffentlicht: (2024)
von: Jiang, Yanbei, et al.
Veröffentlicht: (2024)
KTVIC: A Vietnamese Image Captioning Dataset on the Life Domain
von: Pham, Anh-Cuong, et al.
Veröffentlicht: (2024)
von: Pham, Anh-Cuong, et al.
Veröffentlicht: (2024)
Masked Generative Story Transformer with Character Guidance and Caption Augmentation
von: Papadimitriou, Christos, et al.
Veröffentlicht: (2024)
von: Papadimitriou, Christos, et al.
Veröffentlicht: (2024)
Dense Video Captioning using Graph-based Sentence Summarization
von: Zhang, Zhiwang, et al.
Veröffentlicht: (2025)
von: Zhang, Zhiwang, et al.
Veröffentlicht: (2025)
MAMS: Model-Agnostic Module Selection Framework for Video Captioning
von: Lee, Sangho, et al.
Veröffentlicht: (2025)
von: Lee, Sangho, et al.
Veröffentlicht: (2025)
Controllable Hybrid Captioner for Improved Long-form Video Understanding
von: Sasse, Kuleen, et al.
Veröffentlicht: (2025)
von: Sasse, Kuleen, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Figuring out Figures: Using Textual References to Caption Scientific Figures
von: Cao, Stanley, et al.
Veröffentlicht: (2024) -
Zooming into Comics: Region-Aware RL Improves Fine-Grained Comic Understanding in Vision-Language Models
von: Chen, Yule, et al.
Veröffentlicht: (2025) -
Towards Faithful Reasoning in Comics for Small MLLMs
von: Feng, Chengcheng, et al.
Veröffentlicht: (2026) -
CaptionFool: Universal Image Captioning Model Attacks
von: Parekh, Swapnil
Veröffentlicht: (2026) -
AGIC: Attention-Guided Image Captioning to Improve Caption Relevance
von: Teja, L. D. M. S. Sai, et al.
Veröffentlicht: (2025)