VisuCraft: Enhancing Large Vision-Language Models for Complex Visual-Guided Creative Content Generation via Structured Information Extraction
Fuente:
arXiv
Guardado en:
| Autores principales: | Jiang, Rongxin, Long, Robert, Gu, Chenghao, Yan, Mingrui |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Contextualized Multimodal Lifelong Person Re-Identification in Hybrid Clothing States
por: Long, Robert, et al.
Publicado: (2025)
por: Long, Robert, et al.
Publicado: (2025)
VisuLogic: A Benchmark for Evaluating Visual Reasoning in Multi-modal Large Language Models
por: Xu, Weiye, et al.
Publicado: (2025)
por: Xu, Weiye, et al.
Publicado: (2025)
VisuRiddles: Fine-grained Perception is a Primary Bottleneck for Multimodal Large Language Models in Abstract Visual Reasoning
por: Yan, Hao, et al.
Publicado: (2025)
por: Yan, Hao, et al.
Publicado: (2025)
PartCraft: Crafting Creative Objects by Parts
por: Ng, Kam Woh, et al.
Publicado: (2024)
por: Ng, Kam Woh, et al.
Publicado: (2024)
Enhancing Visual Continual Learning with Language-Guided Supervision
por: Ni, Bolin, et al.
Publicado: (2024)
por: Ni, Bolin, et al.
Publicado: (2024)
Interpreting and Enhancing Emotional Circuits in Large Vision-Language Models via Cross-Modal Information Flow
por: Zhang, Chengsheng, et al.
Publicado: (2026)
por: Zhang, Chengsheng, et al.
Publicado: (2026)
Crafting Large Language Models for Enhanced Interpretability
por: Sun, Chung-En, et al.
Publicado: (2024)
por: Sun, Chung-En, et al.
Publicado: (2024)
Empowering Visual Creativity: A Vision-Language Assistant to Image Editing Recommendations
por: Shen, Tiancheng, et al.
Publicado: (2024)
por: Shen, Tiancheng, et al.
Publicado: (2024)
Adaptive Reinforcement Learning Planning: Harnessing Large Language Models for Complex Information Extraction
por: Ding, Zepeng, et al.
Publicado: (2024)
por: Ding, Zepeng, et al.
Publicado: (2024)
Structure-aware Domain Knowledge Injection for Large Language Models
por: Liu, Kai, et al.
Publicado: (2024)
por: Liu, Kai, et al.
Publicado: (2024)
LDP: Generalizing to Multilingual Visual Information Extraction by Language Decoupled Pretraining
por: Shen, Huawen, et al.
Publicado: (2024)
por: Shen, Huawen, et al.
Publicado: (2024)
Prefill-Time Intervention for Mitigating Hallucination in Large Vision-Language Models
por: Zhang, Chengsheng, et al.
Publicado: (2026)
por: Zhang, Chengsheng, et al.
Publicado: (2026)
Real-Time World Crafting: Generating Structured Game Behaviors from Natural Language with Large Language Models
por: Drake, Austin, et al.
Publicado: (2025)
por: Drake, Austin, et al.
Publicado: (2025)
Relation-Rich Visual Document Generator for Visual Information Extraction
por: Jiang, Zi-Han, et al.
Publicado: (2025)
por: Jiang, Zi-Han, et al.
Publicado: (2025)
UNIKIE-BENCH: Benchmarking Large Multimodal Models for Key Information Extraction in Visual Documents
por: Ji, Yifan, et al.
Publicado: (2026)
por: Ji, Yifan, et al.
Publicado: (2026)
Structure-Aware Decoding Mechanisms for Complex Entity Extraction with Large-Scale Language Models
por: Qiu, Zhimin, et al.
Publicado: (2025)
por: Qiu, Zhimin, et al.
Publicado: (2025)
Anatomical Structure-Guided Medical Vision-Language Pre-training
por: Li, Qingqiu, et al.
Publicado: (2024)
por: Li, Qingqiu, et al.
Publicado: (2024)
Large Language Models for Generative Information Extraction: A Survey
por: Xu, Derong, et al.
Publicado: (2023)
por: Xu, Derong, et al.
Publicado: (2023)
ViSA-Enhanced Aerial VLN: A Visual-Spatial Reasoning Enhanced Framework for Aerial Vision-Language Navigation
por: Tong, Haoyu, et al.
Publicado: (2026)
por: Tong, Haoyu, et al.
Publicado: (2026)
On the Creativity of Large Language Models
por: Franceschelli, Giorgio, et al.
Publicado: (2023)
por: Franceschelli, Giorgio, et al.
Publicado: (2023)
PlotCraft: Pushing the Limits of LLMs for Complex and Interactive Data Visualization
por: Zhang, Jiajun, et al.
Publicado: (2025)
por: Zhang, Jiajun, et al.
Publicado: (2025)
To Sink or Not to Sink: Visual Information Pathways in Large Vision-Language Models
por: Luo, Jiayun, et al.
Publicado: (2025)
por: Luo, Jiayun, et al.
Publicado: (2025)
RedVTP: Training-Free Acceleration of Diffusion Vision-Language Models Inference via Masked Token-Guided Visual Token Pruning
por: Xu, Jingqi, et al.
Publicado: (2025)
por: Xu, Jingqi, et al.
Publicado: (2025)
LacAIDes: Generative AI-Supported Creative Interactive Circuits Crafting to Enliven Traditional Lacquerware
por: Li, Yaning, et al.
Publicado: (2025)
por: Li, Yaning, et al.
Publicado: (2025)
Crafting Generative Art through Genetic Improvement: Managing Creative Outputs in Diverse Fitness Landscapes
por: Fredericks, Erik M., et al.
Publicado: (2024)
por: Fredericks, Erik M., et al.
Publicado: (2024)
Redefining <Creative> in Dictionary: Towards an Enhanced Semantic Understanding of Creative Generation
por: Feng, Fu, et al.
Publicado: (2024)
por: Feng, Fu, et al.
Publicado: (2024)
Enhancing LLM's Cognition via Structurization
por: Liu, Kai, et al.
Publicado: (2024)
por: Liu, Kai, et al.
Publicado: (2024)
STAR: Boosting Low-Resource Information Extraction by Structure-to-Text Data Generation with Large Language Models
por: Ma, Mingyu Derek, et al.
Publicado: (2023)
por: Ma, Mingyu Derek, et al.
Publicado: (2023)
Enhancing Large Vision Language Models with Self-Training on Image Comprehension
por: Deng, Yihe, et al.
Publicado: (2024)
por: Deng, Yihe, et al.
Publicado: (2024)
More Than the Final Answer: Improving Visual Extraction and Logical Consistency in Vision-Language Models
por: Just, Hoang Anh, et al.
Publicado: (2025)
por: Just, Hoang Anh, et al.
Publicado: (2025)
Information Extraction from Electricity Invoices with General-Purpose Large Language Models
por: Gómez, Javier, et al.
Publicado: (2026)
por: Gómez, Javier, et al.
Publicado: (2026)
Finer: Investigating and Enhancing Fine-Grained Visual Concept Recognition in Large Vision Language Models
por: Kim, Jeonghwan, et al.
Publicado: (2024)
por: Kim, Jeonghwan, et al.
Publicado: (2024)
GuideX: Guided Synthetic Data Generation for Zero-Shot Information Extraction
por: De La Fuente, Neil, et al.
Publicado: (2025)
por: De La Fuente, Neil, et al.
Publicado: (2025)
Evaluating and Analyzing Relationship Hallucinations in Large Vision-Language Models
por: Wu, Mingrui, et al.
Publicado: (2024)
por: Wu, Mingrui, et al.
Publicado: (2024)
Enhancing Visual-Language Modality Alignment in Large Vision Language Models via Self-Improvement
por: Wang, Xiyao, et al.
Publicado: (2024)
por: Wang, Xiyao, et al.
Publicado: (2024)
Visual In-Context Learning for Large Vision-Language Models
por: Zhou, Yucheng, et al.
Publicado: (2024)
por: Zhou, Yucheng, et al.
Publicado: (2024)
Recurrent Visual Feature Extraction and Stereo Attentions for CT Report Generation
por: Tian, Yuanhe, et al.
Publicado: (2025)
por: Tian, Yuanhe, et al.
Publicado: (2025)
Probing and Inducing Combinational Creativity in Vision-Language Models
por: Peng, Yongqian, et al.
Publicado: (2025)
por: Peng, Yongqian, et al.
Publicado: (2025)
QueryCraft: Transformer-Guided Query Initialization for Enhanced Human-Object Interaction Detection
por: Wang, Yuxiao, et al.
Publicado: (2025)
por: Wang, Yuxiao, et al.
Publicado: (2025)
Quantifying Generalization Complexity for Large Language Models
por: Qi, Zhenting, et al.
Publicado: (2024)
por: Qi, Zhenting, et al.
Publicado: (2024)
Ejemplares similares
-
Contextualized Multimodal Lifelong Person Re-Identification in Hybrid Clothing States
por: Long, Robert, et al.
Publicado: (2025) -
VisuLogic: A Benchmark for Evaluating Visual Reasoning in Multi-modal Large Language Models
por: Xu, Weiye, et al.
Publicado: (2025) -
VisuRiddles: Fine-grained Perception is a Primary Bottleneck for Multimodal Large Language Models in Abstract Visual Reasoning
por: Yan, Hao, et al.
Publicado: (2025) -
PartCraft: Crafting Creative Objects by Parts
por: Ng, Kam Woh, et al.
Publicado: (2024) -
Enhancing Visual Continual Learning with Language-Guided Supervision
por: Ni, Bolin, et al.
Publicado: (2024)