XMeCap: Meme Caption Generation with Sub-Image Adaptability
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Chen, Yuyan, Yan, Songzhou, Zhu, Zhihong, Li, Zhixu, Xiao, Yanghua |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
HOTVCOM: Generating Buzzworthy Comments for Videos
von: Chen, Yuyan, et al.
Veröffentlicht: (2024)
von: Chen, Yuyan, et al.
Veröffentlicht: (2024)
DiffCap-Bench: A Comprehensive, Challenging, Robust Benchmark for Image Difference Captioning
von: Wei, Yuancheng, et al.
Veröffentlicht: (2026)
von: Wei, Yuancheng, et al.
Veröffentlicht: (2026)
InstanceCap: Improving Text-to-Video Generation via Instance-aware Structured Caption
von: Fan, Tiehan, et al.
Veröffentlicht: (2024)
von: Fan, Tiehan, et al.
Veröffentlicht: (2024)
BalCapRL: A Balanced Framework for RL-Based MLLM Image Captioning
von: Ye, Shaokai, et al.
Veröffentlicht: (2026)
von: Ye, Shaokai, et al.
Veröffentlicht: (2026)
Detecting and Understanding Hateful Contents in Memes Through Captioning and Visual Question-Answering
von: Anaissi, Ali, et al.
Veröffentlicht: (2025)
von: Anaissi, Ali, et al.
Veröffentlicht: (2025)
DualCap: Enhancing Lightweight Image Captioning via Dual Retrieval with Similar Scenes Visual Prompts
von: Li, Binbin, et al.
Veröffentlicht: (2025)
von: Li, Binbin, et al.
Veröffentlicht: (2025)
Enhancing Multimodal In-Context Learning via Inductive-Deductive Reasoning
von: Wang, Haoyu, et al.
Veröffentlicht: (2026)
von: Wang, Haoyu, et al.
Veröffentlicht: (2026)
CapRL: Stimulating Dense Image Caption Capabilities via Reinforcement Learning
von: Xing, Long, et al.
Veröffentlicht: (2025)
von: Xing, Long, et al.
Veröffentlicht: (2025)
CapGeo: A Caption-Assisted Approach to Geometric Reasoning
von: Li, Yuying, et al.
Veröffentlicht: (2025)
von: Li, Yuying, et al.
Veröffentlicht: (2025)
LaMP-Cap: Personalized Figure Caption Generation With Multimodal Figure Profiles
von: Ng, Ho Yin 'Sam', et al.
Veröffentlicht: (2025)
von: Ng, Ho Yin 'Sam', et al.
Veröffentlicht: (2025)
RubiCap: Rubric-Guided Reinforcement Learning for Dense Image Captioning
von: Huang, Tzu-Heng, et al.
Veröffentlicht: (2026)
von: Huang, Tzu-Heng, et al.
Veröffentlicht: (2026)
ChartCap: Mitigating Hallucination of Dense Chart Captioning
von: Lim, Junyoung, et al.
Veröffentlicht: (2025)
von: Lim, Junyoung, et al.
Veröffentlicht: (2025)
GL-PGENet: A Parameterized Generation Framework for Robust Document Image Enhancement
von: Tang, Zhihong
Veröffentlicht: (2025)
von: Tang, Zhihong
Veröffentlicht: (2025)
CompCap: Improving Multimodal Large Language Models with Composite Captions
von: Chen, Xiaohui, et al.
Veröffentlicht: (2024)
von: Chen, Xiaohui, et al.
Veröffentlicht: (2024)
Evaluating Semantic Variation in Text-to-Image Synthesis: A Causal Perspective
von: Zhu, Xiangru, et al.
Veröffentlicht: (2024)
von: Zhu, Xiangru, et al.
Veröffentlicht: (2024)
WsiCaption: Multiple Instance Generation of Pathology Reports for Gigapixel Whole-Slide Images
von: Chen, Pingyi, et al.
Veröffentlicht: (2023)
von: Chen, Pingyi, et al.
Veröffentlicht: (2023)
CaptionFool: Universal Image Captioning Model Attacks
von: Parekh, Swapnil
Veröffentlicht: (2026)
von: Parekh, Swapnil
Veröffentlicht: (2026)
Generating Accurate and Detailed Captions for High-Resolution Images
von: Lee, Hankyeol, et al.
Veröffentlicht: (2025)
von: Lee, Hankyeol, et al.
Veröffentlicht: (2025)
AGIC: Attention-Guided Image Captioning to Improve Caption Relevance
von: Teja, L. D. M. S. Sai, et al.
Veröffentlicht: (2025)
von: Teja, L. D. M. S. Sai, et al.
Veröffentlicht: (2025)
Any2Caption:Interpreting Any Condition to Caption for Controllable Video Generation
von: Wu, Shengqiong, et al.
Veröffentlicht: (2025)
von: Wu, Shengqiong, et al.
Veröffentlicht: (2025)
The Role of Data Curation in Image Captioning
von: Li, Wenyan, et al.
Veröffentlicht: (2023)
von: Li, Wenyan, et al.
Veröffentlicht: (2023)
Exploring Annotation-free Image Captioning with Retrieval-augmented Pseudo Sentence Generation
von: Li, Zhiyuan, et al.
Veröffentlicht: (2023)
von: Li, Zhiyuan, et al.
Veröffentlicht: (2023)
Do Large Multimodal Models Solve Caption Generation for Scientific Figures? Lessons Learned from SciCap Challenge 2023
von: Hsu, Ting-Yao E., et al.
Veröffentlicht: (2025)
von: Hsu, Ting-Yao E., et al.
Veröffentlicht: (2025)
Enhancing Image Caption Generation Using Reinforcement Learning with Human Feedback
von: L, Adarsh N, et al.
Veröffentlicht: (2024)
von: L, Adarsh N, et al.
Veröffentlicht: (2024)
Image Captioning in news report scenario
von: Liu, Tianrui, et al.
Veröffentlicht: (2024)
von: Liu, Tianrui, et al.
Veröffentlicht: (2024)
Automated Image Captioning with CNNs and Transformers
von: Cahyono, Joshua Adrian, et al.
Veröffentlicht: (2024)
von: Cahyono, Joshua Adrian, et al.
Veröffentlicht: (2024)
MoColl: Agent-Based Specific and General Model Collaboration for Image Captioning
von: Yang, Pu, et al.
Veröffentlicht: (2025)
von: Yang, Pu, et al.
Veröffentlicht: (2025)
good4cir: Generating Detailed Synthetic Captions for Composed Image Retrieval
von: Kolouju, Pranavi, et al.
Veröffentlicht: (2025)
von: Kolouju, Pranavi, et al.
Veröffentlicht: (2025)
Pre-Trained CNN Architecture for Transformer-Based Image Caption Generation Model
von: Dufera, Amanuel Tafese
Veröffentlicht: (2025)
von: Dufera, Amanuel Tafese
Veröffentlicht: (2025)
LAViTeR: Learning Aligned Visual and Textual Representations Assisted by Image and Caption Generation
von: Hashemi, Mohammad Abuzar, et al.
Veröffentlicht: (2021)
von: Hashemi, Mohammad Abuzar, et al.
Veröffentlicht: (2021)
Describe Anything: Detailed Localized Image and Video Captioning
von: Lian, Long, et al.
Veröffentlicht: (2025)
von: Lian, Long, et al.
Veröffentlicht: (2025)
Text-only Synthesis for Image Captioning
von: Zhou, Qing, et al.
Veröffentlicht: (2024)
von: Zhou, Qing, et al.
Veröffentlicht: (2024)
Pix2Cap-COCO: Advancing Visual Comprehension via Pixel-Level Captioning
von: You, Zuyao, et al.
Veröffentlicht: (2025)
von: You, Zuyao, et al.
Veröffentlicht: (2025)
Image Embedding Sampling Method for Diverse Captioning
von: Waheed, Sania, et al.
Veröffentlicht: (2025)
von: Waheed, Sania, et al.
Veröffentlicht: (2025)
Top-Down Semantic Refinement for Image Captioning
von: Zhang, Jusheng, et al.
Veröffentlicht: (2025)
von: Zhang, Jusheng, et al.
Veröffentlicht: (2025)
UnSCAR: Universal, Scalable, Controllable, and Adaptable Image Restoration
von: Mandal, Debabrata, et al.
Veröffentlicht: (2026)
von: Mandal, Debabrata, et al.
Veröffentlicht: (2026)
Captioning Daily Activity Images in Early Childhood Education: Benchmark and Algorithm
von: Li, Sixing, et al.
Veröffentlicht: (2026)
von: Li, Sixing, et al.
Veröffentlicht: (2026)
Explainable, Multi-modal Wound Infection Classification from Images Augmented with Generated Captions
von: Busaranuvong, Palawat, et al.
Veröffentlicht: (2025)
von: Busaranuvong, Palawat, et al.
Veröffentlicht: (2025)
MemeMind: A Large-Scale Multimodal Dataset with Chain-of-Thought Reasoning for Harmful Meme Detection
von: Gu, Hexiang, et al.
Veröffentlicht: (2025)
von: Gu, Hexiang, et al.
Veröffentlicht: (2025)
World to Code: Multi-modal Data Generation via Self-Instructed Compositional Captioning and Filtering
von: Wang, Jiacong, et al.
Veröffentlicht: (2024)
von: Wang, Jiacong, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
HOTVCOM: Generating Buzzworthy Comments for Videos
von: Chen, Yuyan, et al.
Veröffentlicht: (2024) -
DiffCap-Bench: A Comprehensive, Challenging, Robust Benchmark for Image Difference Captioning
von: Wei, Yuancheng, et al.
Veröffentlicht: (2026) -
InstanceCap: Improving Text-to-Video Generation via Instance-aware Structured Caption
von: Fan, Tiehan, et al.
Veröffentlicht: (2024) -
BalCapRL: A Balanced Framework for RL-Based MLLM Image Captioning
von: Ye, Shaokai, et al.
Veröffentlicht: (2026) -
Detecting and Understanding Hateful Contents in Memes Through Captioning and Visual Question-Answering
von: Anaissi, Ali, et al.
Veröffentlicht: (2025)