Caption Generation for Dongba Paintings via Prompt Learning and Semantic Fusion
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Qian, Shuangwu, Yuan, Xiaochan, Liu, Pengfei |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
DongbaMIE: A Multimodal Information Extraction Dataset for Evaluating Semantic Understanding of Dongba Pictograms
von: Bi, Xiaojun, et al.
Veröffentlicht: (2025)
von: Bi, Xiaojun, et al.
Veröffentlicht: (2025)
Multi-View Deformable Convolution Meets Visual Mamba for Coronary Artery Segmentation
von: Yuan, Xiaochan, et al.
Veröffentlicht: (2026)
von: Yuan, Xiaochan, et al.
Veröffentlicht: (2026)
Painting with Words: Elevating Detailed Image Captioning with Benchmark and Alignment Learning
von: Ye, Qinghao, et al.
Veröffentlicht: (2025)
von: Ye, Qinghao, et al.
Veröffentlicht: (2025)
Semantic-Spatial Feature Fusion with Dynamic Graph Refinement for Remote Sensing Image Captioning
von: Liu, Maofu, et al.
Veröffentlicht: (2025)
von: Liu, Maofu, et al.
Veröffentlicht: (2025)
Caption Anything in Video: Fine-grained Object-centric Captioning via Spatiotemporal Multimodal Prompting
von: Tang, Yunlong, et al.
Veröffentlicht: (2025)
von: Tang, Yunlong, et al.
Veröffentlicht: (2025)
CtrlFuse: Mask-Prompt Guided Controllable Infrared and Visible Image Fusion
von: Sun, Yiming, et al.
Veröffentlicht: (2026)
von: Sun, Yiming, et al.
Veröffentlicht: (2026)
RxnCaption: Reformulating Reaction Diagram Parsing as Visual Prompt Guided Captioning
von: Song, Jiahe, et al.
Veröffentlicht: (2025)
von: Song, Jiahe, et al.
Veröffentlicht: (2025)
JSSFF: A Joint Structural-Semantic Fusion Framework for Remote Sensing Image Captioning
von: Das, Swadhin, et al.
Veröffentlicht: (2026)
von: Das, Swadhin, et al.
Veröffentlicht: (2026)
Learning Transferable Facial Emotion Representations from Large-Scale Semantically Rich Captions
von: Sun, Licai, et al.
Veröffentlicht: (2025)
von: Sun, Licai, et al.
Veröffentlicht: (2025)
OPCap:Object-aware Prompting Captioning
von: Huang, Feiyang
Veröffentlicht: (2024)
von: Huang, Feiyang
Veröffentlicht: (2024)
FreeKD: Knowledge Distillation via Semantic Frequency Prompt
von: Zhang, Yuan, et al.
Veröffentlicht: (2023)
von: Zhang, Yuan, et al.
Veröffentlicht: (2023)
Semantic-CC: Boosting Remote Sensing Image Change Captioning via Foundational Knowledge and Semantic Guidance
von: Zhu, Yongshuo, et al.
Veröffentlicht: (2024)
von: Zhu, Yongshuo, et al.
Veröffentlicht: (2024)
Every Painting Awakened: A Training-free Framework for Painting-to-Animation Generation
von: Liu, Lingyu, et al.
Veröffentlicht: (2025)
von: Liu, Lingyu, et al.
Veröffentlicht: (2025)
I Dream My Painting: Connecting MLLMs and Diffusion Models via Prompt Generation for Text-Guided Multi-Mask Inpainting
von: Fanelli, Nicola, et al.
Veröffentlicht: (2024)
von: Fanelli, Nicola, et al.
Veröffentlicht: (2024)
Set Prediction Guided by Semantic Concepts for Diverse Video Captioning
von: Lu, Yifan, et al.
Veröffentlicht: (2023)
von: Lu, Yifan, et al.
Veröffentlicht: (2023)
New Encoder Learning for Captioning Heavy Rain Images via Semantic Visual Feature Matching
von: Son, Chang-Hwan, et al.
Veröffentlicht: (2021)
von: Son, Chang-Hwan, et al.
Veröffentlicht: (2021)
PaintFlow: A Unified Framework for Interactive Oil Paintings Editing and Generation
von: Hu, Zhangli, et al.
Veröffentlicht: (2025)
von: Hu, Zhangli, et al.
Veröffentlicht: (2025)
Inverse Painting: Reconstructing The Painting Process
von: Chen, Bowei, et al.
Veröffentlicht: (2024)
von: Chen, Bowei, et al.
Veröffentlicht: (2024)
Semantic-aware SAM for Point-Prompted Instance Segmentation
von: Wei, Zhaoyang, et al.
Veröffentlicht: (2023)
von: Wei, Zhaoyang, et al.
Veröffentlicht: (2023)
Visual and Semantic Prompt Collaboration for Generalized Zero-Shot Learning
von: Jiang, Huajie, et al.
Veröffentlicht: (2025)
von: Jiang, Huajie, et al.
Veröffentlicht: (2025)
Rewrite Caption Semantics: Bridging Semantic Gaps for Language-Supervised Semantic Segmentation
von: Xing, Yun, et al.
Veröffentlicht: (2023)
von: Xing, Yun, et al.
Veröffentlicht: (2023)
Any2Caption:Interpreting Any Condition to Caption for Controllable Video Generation
von: Wu, Shengqiong, et al.
Veröffentlicht: (2025)
von: Wu, Shengqiong, et al.
Veröffentlicht: (2025)
Text Data-Centric Image Captioning with Interactive Prompts
von: Wang, Yiyu, et al.
Veröffentlicht: (2024)
von: Wang, Yiyu, et al.
Veröffentlicht: (2024)
DSPFusion: Image Fusion via Degradation and Semantic Dual-Prior Guidance
von: Tang, Linfeng, et al.
Veröffentlicht: (2025)
von: Tang, Linfeng, et al.
Veröffentlicht: (2025)
Robust RGB-T Tracking via Learnable Visual Fourier Prompt Fine-tuning and Modality Fusion Prompt Generation
von: Yang, Hongtao, et al.
Veröffentlicht: (2025)
von: Yang, Hongtao, et al.
Veröffentlicht: (2025)
CritiFusion: Semantic Critique and Spectral Alignment for Faithful Text-to-Image Generation
von: Chen, ZhenQi, et al.
Veröffentlicht: (2025)
von: Chen, ZhenQi, et al.
Veröffentlicht: (2025)
Semantic-Aware Prefix Learning for Token-Efficient Image Generation
von: Li, Qingfeng, et al.
Veröffentlicht: (2026)
von: Li, Qingfeng, et al.
Veröffentlicht: (2026)
Prompt-Based Caption Generation for Single-Tooth Dental Images Using Vision-Language Models
von: Sukhanova, Anastasiia, et al.
Veröffentlicht: (2026)
von: Sukhanova, Anastasiia, et al.
Veröffentlicht: (2026)
Vision Graph Prompting via Semantic Low-Rank Decomposition
von: Ai, Zixiang, et al.
Veröffentlicht: (2025)
von: Ai, Zixiang, et al.
Veröffentlicht: (2025)
PTGCF: Printing Texture Guided Color Fusion for Impressionism Oil Painting Style Rendering
von: Geng, Jing, et al.
Veröffentlicht: (2022)
von: Geng, Jing, et al.
Veröffentlicht: (2022)
It's Just Another Day: Unique Video Captioning by Discriminative Prompting
von: Perrett, Toby, et al.
Veröffentlicht: (2024)
von: Perrett, Toby, et al.
Veröffentlicht: (2024)
RORPCap: Retrieval-based Objects and Relations Prompt for Image Captioning
von: Gu, Jinjing, et al.
Veröffentlicht: (2025)
von: Gu, Jinjing, et al.
Veröffentlicht: (2025)
HowToCaption: Prompting LLMs to Transform Video Annotations at Scale
von: Shvetsova, Nina, et al.
Veröffentlicht: (2023)
von: Shvetsova, Nina, et al.
Veröffentlicht: (2023)
Dense Video Captioning Using Unsupervised Semantic Information
von: Estevam, Valter, et al.
Veröffentlicht: (2021)
von: Estevam, Valter, et al.
Veröffentlicht: (2021)
Easier Painting Than Thinking: Can Text-to-Image Models Set the Stage, but Not Direct the Play?
von: Li, Ouxiang, et al.
Veröffentlicht: (2025)
von: Li, Ouxiang, et al.
Veröffentlicht: (2025)
PMCE: Probabilistic Multi-Granularity Semantics with Caption-Guided Enhancement for Few-Shot Learning
von: Wu, Jiaying, et al.
Veröffentlicht: (2026)
von: Wu, Jiaying, et al.
Veröffentlicht: (2026)
TimeChat-Captioner: Scripting Multi-Scene Videos with Time-Aware and Structural Audio-Visual Captions
von: Yao, Linli, et al.
Veröffentlicht: (2026)
von: Yao, Linli, et al.
Veröffentlicht: (2026)
Task-Customized Mixture of Adapters for General Image Fusion
von: Zhu, Pengfei, et al.
Veröffentlicht: (2024)
von: Zhu, Pengfei, et al.
Veröffentlicht: (2024)
Joint Semantic Token Selection and Prompt Optimization for Interpretable Prompt Learning
von: Wang, Yating, et al.
Veröffentlicht: (2026)
von: Wang, Yating, et al.
Veröffentlicht: (2026)
PromptFusion: Decoupling Stability and Plasticity for Continual Learning
von: Chen, Haoran, et al.
Veröffentlicht: (2023)
von: Chen, Haoran, et al.
Veröffentlicht: (2023)
Ähnliche Einträge
-
DongbaMIE: A Multimodal Information Extraction Dataset for Evaluating Semantic Understanding of Dongba Pictograms
von: Bi, Xiaojun, et al.
Veröffentlicht: (2025) -
Multi-View Deformable Convolution Meets Visual Mamba for Coronary Artery Segmentation
von: Yuan, Xiaochan, et al.
Veröffentlicht: (2026) -
Painting with Words: Elevating Detailed Image Captioning with Benchmark and Alignment Learning
von: Ye, Qinghao, et al.
Veröffentlicht: (2025) -
Semantic-Spatial Feature Fusion with Dynamic Graph Refinement for Remote Sensing Image Captioning
von: Liu, Maofu, et al.
Veröffentlicht: (2025) -
Caption Anything in Video: Fine-grained Object-centric Captioning via Spatiotemporal Multimodal Prompting
von: Tang, Yunlong, et al.
Veröffentlicht: (2025)