Multi-task Prompt Words Learning for Social Media Content Generation
Fuente:
arXiv
Saved in:
| Main Authors: | Xue, Haochen, Zhang, Chong, Liu, Chengzhi, Wu, Fangyu, Jin, Xiaobo |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Prompt-A-Video: Prompt Your Video Diffusion Model via Preference-Aligned LLM
by: Ji, Yatai, et al.
Published: (2024)
by: Ji, Yatai, et al.
Published: (2024)
HAIC: Improving Human Action Understanding and Generation with Better Captions for Multi-modal Large Language Models
by: Wang, Xiao, et al.
Published: (2025)
by: Wang, Xiao, et al.
Published: (2025)
WordArt Designer API: User-Driven Artistic Typography Synthesis with Large Language Models on ModelScope
by: He, Jun-Yan, et al.
Published: (2024)
by: He, Jun-Yan, et al.
Published: (2024)
Co-Reinforcement Learning for Unified Multimodal Understanding and Generation
by: Jiang, Jingjing, et al.
Published: (2025)
by: Jiang, Jingjing, et al.
Published: (2025)
Teach Me Sign: Stepwise Prompting LLM for Sign Language Production
by: An, Zhaoyi, et al.
Published: (2025)
by: An, Zhaoyi, et al.
Published: (2025)
Dynamic Self-adaptive Multiscale Distillation from Pre-trained Multimodal Large Model for Efficient Cross-modal Representation Learning
by: Liang, Zhengyang, et al.
Published: (2024)
by: Liang, Zhengyang, et al.
Published: (2024)
Multi-Modal Semantic Parsing for the Interpretation of Tombstone Inscriptions
by: Zhang, Xiao, et al.
Published: (2025)
by: Zhang, Xiao, et al.
Published: (2025)
Graph-Driven Multimodal Feature Learning Framework for Apparent Personality Assessment
by: Wang, Kangsheng, et al.
Published: (2025)
by: Wang, Kangsheng, et al.
Published: (2025)
CPSL: Representing Volumetric Video via Content-Promoted Scene Layers
by: Hu, Kaiyuan, et al.
Published: (2025)
by: Hu, Kaiyuan, et al.
Published: (2025)
M$^3$Face: A Unified Multi-Modal Multilingual Framework for Human Face Generation and Editing
by: Mofayezi, Mohammadreza, et al.
Published: (2024)
by: Mofayezi, Mohammadreza, et al.
Published: (2024)
Bi-VLDoc: Bidirectional Vision-Language Modeling for Visually-Rich Document Understanding
by: Luo, Chuwei, et al.
Published: (2022)
by: Luo, Chuwei, et al.
Published: (2022)
Fit and Prune: Fast and Training-free Visual Token Pruning for Multi-modal Large Language Models
by: Ye, Weihao, et al.
Published: (2024)
by: Ye, Weihao, et al.
Published: (2024)
Enhancing the Learning Experience: Using Vision-Language Models to Generate Questions for Educational Videos
by: Stamatakis, Markos, et al.
Published: (2025)
by: Stamatakis, Markos, et al.
Published: (2025)
Rethinking Radiology Report Generation via Causal Inspired Counterfactual Augmentation
by: Song, Xiao, et al.
Published: (2023)
by: Song, Xiao, et al.
Published: (2023)
Seeing Beyond Words: Self-Supervised Visual Learning for Multimodal Large Language Models
by: Caffagni, Davide, et al.
Published: (2025)
by: Caffagni, Davide, et al.
Published: (2025)
Movie101v2: Improved Movie Narration Benchmark
by: Yue, Zihao, et al.
Published: (2024)
by: Yue, Zihao, et al.
Published: (2024)
KeyVideoLLM: Towards Large-scale Video Keyframe Selection
by: Liang, Hao, et al.
Published: (2024)
by: Liang, Hao, et al.
Published: (2024)
MLANet: Multi-Level Attention Network with Sub-instruction for Continuous Vision-and-Language Navigation
by: He, Zongtao, et al.
Published: (2023)
by: He, Zongtao, et al.
Published: (2023)
Ask Questions with Double Hints: Visual Question Generation with Answer-awareness and Region-reference
by: Shen, Kai, et al.
Published: (2024)
by: Shen, Kai, et al.
Published: (2024)
MaVEn: An Effective Multi-granularity Hybrid Visual Encoding Framework for Multimodal Large Language Model
by: Jiang, Chaoya, et al.
Published: (2024)
by: Jiang, Chaoya, et al.
Published: (2024)
MuLTI: Efficient Video-and-Language Understanding with Text-Guided MultiWay-Sampler and Multiple Choice Modeling
by: Xu, Jiaqi, et al.
Published: (2023)
by: Xu, Jiaqi, et al.
Published: (2023)
Multi-proposal Collaboration and Multi-task Training for Weakly-supervised Video Moment Retrieval
by: Zhang, Bolin, et al.
Published: (2026)
by: Zhang, Bolin, et al.
Published: (2026)
MERIT: Multilingual Semantic Retrieval with Interleaved Multi-Condition Query
by: Chow, Wei, et al.
Published: (2025)
by: Chow, Wei, et al.
Published: (2025)
Failures to Surface Harmful Contents in Video Large Language Models
by: Cao, Yuxin, et al.
Published: (2025)
by: Cao, Yuxin, et al.
Published: (2025)
EasyAnimate: High-Performance Video Generation Framework with Hybrid Windows Attention and Reward Backpropagation
by: Xu, Jiaqi, et al.
Published: (2024)
by: Xu, Jiaqi, et al.
Published: (2024)
Anim-Director: A Large Multimodal Model Powered Agent for Controllable Animation Video Generation
by: Li, Yunxin, et al.
Published: (2024)
by: Li, Yunxin, et al.
Published: (2024)
AeSlides: Incentivizing Aesthetic Layout in LLM-Based Slide Generation via Verifiable Rewards
by: Pan, Yiming, et al.
Published: (2026)
by: Pan, Yiming, et al.
Published: (2026)
Recipe Generation from Unsegmented Cooking Videos
by: Nishimura, Taichi, et al.
Published: (2022)
by: Nishimura, Taichi, et al.
Published: (2022)
Analyzing Images of Legal Documents: Toward Multi-Modal LLMs for Access to Justice
by: Westermann, Hannes, et al.
Published: (2024)
by: Westermann, Hannes, et al.
Published: (2024)
GameTileNet: A Semantic Dataset for Low-Resolution Game Art in Procedural Content Generation
by: Chen, Yi-Chun, et al.
Published: (2025)
by: Chen, Yi-Chun, et al.
Published: (2025)
VideoCLIP-XL: Advancing Long Description Understanding for Video CLIP Models
by: Wang, Jiapeng, et al.
Published: (2024)
by: Wang, Jiapeng, et al.
Published: (2024)
How Far Are We from Generating Missing Modalities with Foundation Models?
by: Ke, Guanzhou, et al.
Published: (2025)
by: Ke, Guanzhou, et al.
Published: (2025)
Quizzard@INOVA Challenge 2025 -- Track A: Plug-and-Play Technique in Interleaved Multi-Image Model
by: Cuong, Dinh Viet, et al.
Published: (2025)
by: Cuong, Dinh Viet, et al.
Published: (2025)
PDA: Text-Augmented Defense Framework for Robust Vision-Language Models against Adversarial Image Attacks
by: Xu, Jingning, et al.
Published: (2026)
by: Xu, Jingning, et al.
Published: (2026)
LocoMotion: Learning Motion-Focused Video-Language Representations
by: Doughty, Hazel, et al.
Published: (2024)
by: Doughty, Hazel, et al.
Published: (2024)
DreamArtist++: Controllable One-Shot Text-to-Image Generation via Positive-Negative Adapter
by: Dong, Ziyi, et al.
Published: (2022)
by: Dong, Ziyi, et al.
Published: (2022)
MambaPro: Multi-Modal Object Re-Identification with Mamba Aggregation and Synergistic Prompt
by: Wang, Yuhao, et al.
Published: (2024)
by: Wang, Yuhao, et al.
Published: (2024)
Magnifier Prompt: Tackling Multimodal Hallucination via Extremely Simple Instructions
by: Fu, Yuhan, et al.
Published: (2024)
by: Fu, Yuhan, et al.
Published: (2024)
Empowering Multimodal LLMs with External Tools: A Comprehensive Survey
by: An, Wenbin, et al.
Published: (2025)
by: An, Wenbin, et al.
Published: (2025)
Seeing Through Deception: Uncovering Misleading Creator Intent in Multimodal News with Vision-Language Models
by: Wu, Jiaying, et al.
Published: (2025)
by: Wu, Jiaying, et al.
Published: (2025)
Similar Items
-
Prompt-A-Video: Prompt Your Video Diffusion Model via Preference-Aligned LLM
by: Ji, Yatai, et al.
Published: (2024) -
HAIC: Improving Human Action Understanding and Generation with Better Captions for Multi-modal Large Language Models
by: Wang, Xiao, et al.
Published: (2025) -
WordArt Designer API: User-Driven Artistic Typography Synthesis with Large Language Models on ModelScope
by: He, Jun-Yan, et al.
Published: (2024) -
Co-Reinforcement Learning for Unified Multimodal Understanding and Generation
by: Jiang, Jingjing, et al.
Published: (2025) -
Teach Me Sign: Stepwise Prompting LLM for Sign Language Production
by: An, Zhaoyi, et al.
Published: (2025)