Redefining <Creative> in Dictionary: Towards an Enhanced Semantic Understanding of Creative Generation
Fuente:
arXiv
Guardado en:
| Autores principales: | Feng, Fu, Xie, Yucheng, Yang, Xu, Wang, Jing, Geng, Xin |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Distribution-Conditional Generation: From Class Distribution to Creative Generation
por: Feng, Fu, et al.
Publicado: (2025)
por: Feng, Fu, et al.
Publicado: (2025)
A Creative Agent is Worth a 64-Token Template
por: Shi, Ruixiao, et al.
Publicado: (2026)
por: Shi, Ruixiao, et al.
Publicado: (2026)
Mining Contextualized Visual Associations from Images for Creativity Understanding
por: Sahu, Ananya, et al.
Publicado: (2025)
por: Sahu, Ananya, et al.
Publicado: (2025)
CAP: Evaluation of Persuasive and Creative Image Generation
por: Aghazadeh, Aysan, et al.
Publicado: (2024)
por: Aghazadeh, Aysan, et al.
Publicado: (2024)
VisuCraft: Enhancing Large Vision-Language Models for Complex Visual-Guided Creative Content Generation via Structured Information Extraction
por: Jiang, Rongxin, et al.
Publicado: (2025)
por: Jiang, Rongxin, et al.
Publicado: (2025)
Probing and Inducing Combinational Creativity in Vision-Language Models
por: Peng, Yongqian, et al.
Publicado: (2025)
por: Peng, Yongqian, et al.
Publicado: (2025)
CogOmniControl: Reasoning-Driven Controllable Video Generation via Creative Intent Cognition
por: Yang, Hongji, et al.
Publicado: (2026)
por: Yang, Hongji, et al.
Publicado: (2026)
KIND: Knowledge Integration and Diversion for Training Decomposable Models
por: Xie, Yucheng, et al.
Publicado: (2024)
por: Xie, Yucheng, et al.
Publicado: (2024)
DivControl: Knowledge Diversion for Controllable Image Generation
por: Xie, Yucheng, et al.
Publicado: (2025)
por: Xie, Yucheng, et al.
Publicado: (2025)
FAD: Frequency Adaptation and Diversion for Cross-domain Few-shot Learning
por: Shi, Ruixiao, et al.
Publicado: (2025)
por: Shi, Ruixiao, et al.
Publicado: (2025)
Self-Creative Text-to-Object Generation using Semantic-Aware Spatial Weighting
por: Yu, Yue, et al.
Publicado: (2026)
por: Yu, Yue, et al.
Publicado: (2026)
Creative4U: MLLMs-based Advertising Creative Image Selector with Comparative Reasoning
por: Lin, Yukang, et al.
Publicado: (2025)
por: Lin, Yukang, et al.
Publicado: (2025)
DualToken: Towards Unifying Visual Understanding and Generation with Dual Visual Vocabularies
por: Song, Wei, et al.
Publicado: (2025)
por: Song, Wei, et al.
Publicado: (2025)
Enhancing Creative Generation on Stable Diffusion-based Models
por: Han, Jiyeon, et al.
Publicado: (2025)
por: Han, Jiyeon, et al.
Publicado: (2025)
Let's Think Outside the Box: Exploring Leap-of-Thought in Large Language Models with Creative Humor Generation
por: Zhong, Shanshan, et al.
Publicado: (2023)
por: Zhong, Shanshan, et al.
Publicado: (2023)
FINE: Factorizing Knowledge for Initialization of Variable-sized Diffusion Models
por: Xie, Yucheng, et al.
Publicado: (2024)
por: Xie, Yucheng, et al.
Publicado: (2024)
Draw ALL Your Imagine: A Holistic Benchmark and Agent Framework for Complex Instruction-based Image Generation
por: Zhou, Yucheng, et al.
Publicado: (2025)
por: Zhou, Yucheng, et al.
Publicado: (2025)
Self-Supervised Weight Templates for Scalable Vision Model Initialization
por: Xie, Yucheng, et al.
Publicado: (2026)
por: Xie, Yucheng, et al.
Publicado: (2026)
Creative Image Generation with Diffusion Models
por: Song, Kunpeng, et al.
Publicado: (2026)
por: Song, Kunpeng, et al.
Publicado: (2026)
Multi-Object Advertisement Creative Generation
por: Gao, Jialu, et al.
Publicado: (2026)
por: Gao, Jialu, et al.
Publicado: (2026)
DoraemonGPT: Toward Understanding Dynamic Scenes with Large Language Models (Exemplified as A Video Agent)
por: Yang, Zongxin, et al.
Publicado: (2024)
por: Yang, Zongxin, et al.
Publicado: (2024)
LADR: Locality-Aware Dynamic Rescue for Efficient Text-to-Image Generation with Diffusion Large Language Models
por: Wang, Chenglin, et al.
Publicado: (2026)
por: Wang, Chenglin, et al.
Publicado: (2026)
Towards Event-oriented Long Video Understanding
por: Du, Yifan, et al.
Publicado: (2024)
por: Du, Yifan, et al.
Publicado: (2024)
M4U: Evaluating Multilingual Understanding and Reasoning for Large Multimodal Models
por: Wang, Hongyu, et al.
Publicado: (2024)
por: Wang, Hongyu, et al.
Publicado: (2024)
From Text to Pixel: Advancing Long-Context Understanding in MLLMs
por: Lu, Yujie, et al.
Publicado: (2024)
por: Lu, Yujie, et al.
Publicado: (2024)
VideoLLaMB: Long Streaming Video Understanding with Recurrent Memory Bridges
por: Wang, Yuxuan, et al.
Publicado: (2024)
por: Wang, Yuxuan, et al.
Publicado: (2024)
Can Large Vision-Language Models Understand Multimodal Sarcasm?
por: Wang, Xinyu, et al.
Publicado: (2025)
por: Wang, Xinyu, et al.
Publicado: (2025)
Creative Portraiture: Exploring Creative Adversarial Networks and Conditional Creative Adversarial Networks
por: Hereu, Sebastian, et al.
Publicado: (2024)
por: Hereu, Sebastian, et al.
Publicado: (2024)
SemVink: Advancing VLMs' Semantic Understanding of Optical Illusions via Visual Global Thinking
por: Li, Sifan, et al.
Publicado: (2025)
por: Li, Sifan, et al.
Publicado: (2025)
Video Understanding with Large Language Models: A Survey
por: Tang, Yolo Y., et al.
Publicado: (2023)
por: Tang, Yolo Y., et al.
Publicado: (2023)
Vibe Spaces for Creatively Connecting and Expressing Visual Concepts
por: Yang, Huzheng, et al.
Publicado: (2025)
por: Yang, Huzheng, et al.
Publicado: (2025)
Towards Effective Long-Video Event Prediction via Multi-Level Event Semantics Mining
por: Peng, Bo, et al.
Publicado: (2026)
por: Peng, Bo, et al.
Publicado: (2026)
KD-CVG: A Knowledge-Driven Approach for Creative Video Generation
por: Liu, Linkai, et al.
Publicado: (2026)
por: Liu, Linkai, et al.
Publicado: (2026)
Multi-modal Semantic Understanding with Contrastive Cross-modal Feature Alignment
por: Zhang, Ming, et al.
Publicado: (2024)
por: Zhang, Ming, et al.
Publicado: (2024)
Realizing Video Summarization from the Path of Language-based Semantic Understanding
por: Mu, Kuan-Chen, et al.
Publicado: (2024)
por: Mu, Kuan-Chen, et al.
Publicado: (2024)
DocScope: Benchmarking Verifiable Reasoning for Trustworthy Long-Document Understanding
por: Feng, Xiang, et al.
Publicado: (2026)
por: Feng, Xiang, et al.
Publicado: (2026)
UniWorld-V1: High-Resolution Semantic Encoders for Unified Visual Understanding and Generation
por: Lin, Bin, et al.
Publicado: (2025)
por: Lin, Bin, et al.
Publicado: (2025)
Vision-G1: Towards General Vision Language Reasoning with Multi-Domain Data Curation
por: Zha, Yuheng, et al.
Publicado: (2025)
por: Zha, Yuheng, et al.
Publicado: (2025)
Commonsense-T2I Challenge: Can Text-to-Image Generation Models Understand Commonsense?
por: Fu, Xingyu, et al.
Publicado: (2024)
por: Fu, Xingyu, et al.
Publicado: (2024)
Adaptive Markup Language Generation for Contextually-Grounded Visual Document Understanding
por: Xiao, Han, et al.
Publicado: (2025)
por: Xiao, Han, et al.
Publicado: (2025)
Ejemplares similares
-
Distribution-Conditional Generation: From Class Distribution to Creative Generation
por: Feng, Fu, et al.
Publicado: (2025) -
A Creative Agent is Worth a 64-Token Template
por: Shi, Ruixiao, et al.
Publicado: (2026) -
Mining Contextualized Visual Associations from Images for Creativity Understanding
por: Sahu, Ananya, et al.
Publicado: (2025) -
CAP: Evaluation of Persuasive and Creative Image Generation
por: Aghazadeh, Aysan, et al.
Publicado: (2024) -
VisuCraft: Enhancing Large Vision-Language Models for Complex Visual-Guided Creative Content Generation via Structured Information Extraction
por: Jiang, Rongxin, et al.
Publicado: (2025)