HiPrompt: Tuning-free Higher-Resolution Generation with Hierarchical MLLM Prompts
Fuente:
arXiv
Saved in:
| Main Authors: | Liu, Xinyu, He, Yingqing, Guo, Lanqing, Li, Xiang, Jin, Bu, Li, Peng, Li, Yan, Chan, Chi-Min, Chen, Qifeng, Xue, Wei, Luo, Wenhan, Liu, Qifeng, Guo, Yike |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
VFX Creator: Animated Visual Effect Generation with Controllable Diffusion Transformer
by: Liu, Xinyu, et al.
Published: (2025)
by: Liu, Xinyu, et al.
Published: (2025)
Importance Weighting Can Help Large Language Models Self-Improve
by: Jiang, Chunyang, et al.
Published: (2024)
by: Jiang, Chunyang, et al.
Published: (2024)
Weakly-Supervised Emotion Transition Learning for Diverse 3D Co-speech Gesture Generation
by: Qi, Xingqun, et al.
Published: (2023)
by: Qi, Xingqun, et al.
Published: (2023)
Make a Cheap Scaling: A Self-Cascade Diffusion Model for Higher-Resolution Adaptation
by: Guo, Lanqing, et al.
Published: (2024)
by: Guo, Lanqing, et al.
Published: (2024)
Foundation Cures Personalization: Improving Personalized Models' Prompt Consistency via Hidden Foundation Knowledge
by: Cai, Yiyang, et al.
Published: (2024)
by: Cai, Yiyang, et al.
Published: (2024)
Reflection in the Dark: Exposing and Escaping the Black Box in Reflective Prompt Optimization
by: Liu, Shiyan, et al.
Published: (2026)
by: Liu, Shiyan, et al.
Published: (2026)
Era3D: High-Resolution Multiview Diffusion using Efficient Row-wise Attention
by: Li, Peng, et al.
Published: (2024)
by: Li, Peng, et al.
Published: (2024)
Co$^{3}$Gesture: Towards Coherent Concurrent Co-speech 3D Gesture Generation with Interactive Diffusion
by: Qi, Xingqun, et al.
Published: (2025)
by: Qi, Xingqun, et al.
Published: (2025)
MMTrail: A Multimodal Trailer Video Dataset with Language and Music Descriptions
by: Chi, Xiaowei, et al.
Published: (2024)
by: Chi, Xiaowei, et al.
Published: (2024)
VidMuse: A Simple Video-to-Music Generation Framework with Long-Short-Term Modeling
by: Tian, Zeyue, et al.
Published: (2024)
by: Tian, Zeyue, et al.
Published: (2024)
M$^{2}$Chat: Empowering VLM for Multimodal LLM Interleaved Text-Image Generation
by: Chi, Xiaowei, et al.
Published: (2023)
by: Chi, Xiaowei, et al.
Published: (2023)
CoCoGesture: Toward Coherent Co-speech 3D Gesture Generation in the Wild
by: Qi, Xingqun, et al.
Published: (2024)
by: Qi, Xingqun, et al.
Published: (2024)
A Data-embedded Solution Paradigm for Nonconvex Probable Event Constrained Optimization
by: Li, Qifeng
Published: (2026)
by: Li, Qifeng
Published: (2026)
Probable Event Constrained Optimization and A Data-embedded Solution Paradigm
by: Li, Qifeng
Published: (2022)
by: Li, Qifeng
Published: (2022)
CoMoSVC: Consistency Model-based Singing Voice Conversion
by: Lu, Yiwen, et al.
Published: (2024)
by: Lu, Yiwen, et al.
Published: (2024)
FastSAG: Towards Fast Non-Autoregressive Singing Accompaniment Generation
by: Chen, Jianyi, et al.
Published: (2024)
by: Chen, Jianyi, et al.
Published: (2024)
Follow-Your-Click: Open-domain Regional Image Animation via Short Prompts
by: Ma, Yue, et al.
Published: (2024)
by: Ma, Yue, et al.
Published: (2024)
UniSH: Unifying Scene and Human Reconstruction in a Feed-Forward Pass
by: Li, Mengfei, et al.
Published: (2026)
by: Li, Mengfei, et al.
Published: (2026)
LLMs Meet Multimodal Generation and Editing: A Survey
by: He, Yingqing, et al.
Published: (2024)
by: He, Yingqing, et al.
Published: (2024)
Characteristic conic connections and torsion-free principal connections
by: Hwang, Jun-Muk, et al.
Published: (2024)
by: Hwang, Jun-Muk, et al.
Published: (2024)
Robust Depth Enhancement via Polarization Prompt Fusion Tuning
by: Ikemura, Kei, et al.
Published: (2024)
by: Ikemura, Kei, et al.
Published: (2024)
Optimization of Barrier Properties of PLA/PBAT Polymers in Biodegradable Materials
by: Wang Xinyu, et al.
Published: (2024)
by: Wang Xinyu, et al.
Published: (2024)
HiStyle: Hierarchical Style Embedding Predictor for Text-Prompt-Guided Controllable Speech Synthesis
by: Zhang, Ziyu, et al.
Published: (2025)
by: Zhang, Ziyu, et al.
Published: (2025)
CogniEdit: Dense Gradient Flow Optimization for Fine-Grained Image Editing
by: Li, Yan, et al.
Published: (2025)
by: Li, Yan, et al.
Published: (2025)
Attention Hijacking: Response Manipulation Across Queries in Vision-Language Models
by: Wang, Zhiqiang, et al.
Published: (2026)
by: Wang, Zhiqiang, et al.
Published: (2026)
SINGER: Vivid Audio-driven Singing Video Generation with Multi-scale Spectral Diffusion Model
by: Li, Yan, et al.
Published: (2024)
by: Li, Yan, et al.
Published: (2024)
PROMETHEE-based Modeling of Endogenous Behavioral Uncertainty of EV Owners
by: Sarkar, Dipayan, et al.
Published: (2026)
by: Sarkar, Dipayan, et al.
Published: (2026)
Real-time Optimization for Wind-to-H2 Driven Critical Infrastructures Based on Active Constraints Identification and Integer Variables Prediction
by: Goodarzi, Mostafa, et al.
Published: (2025)
by: Goodarzi, Mostafa, et al.
Published: (2025)
Identification and role of differentially expressed genes/proteins between pulmonary tuberculosis patients and controls across lung tissues and blood samples
by: Qifeng Li, et al.
Published: (2024)
by: Qifeng Li, et al.
Published: (2024)
EVA: An Embodied World Model for Future Video Anticipation
by: Chi, Xiaowei, et al.
Published: (2024)
by: Chi, Xiaowei, et al.
Published: (2024)
Fake it till You Make it: Reward Modeling as Discriminative Prediction
by: Liu, Runtao, et al.
Published: (2025)
by: Liu, Runtao, et al.
Published: (2025)
Efficient Prompt Tuning by Multi-Space Projection and Prompt Fusion
by: Lan, Pengxiang, et al.
Published: (2024)
by: Lan, Pengxiang, et al.
Published: (2024)
Hi-Mamba: Hierarchical Mamba for Efficient Image Super-Resolution
by: Qiao, Junbo, et al.
Published: (2024)
by: Qiao, Junbo, et al.
Published: (2024)
Seeing and Hearing: Open-domain Visual-Audio Generation with Diffusion Latent Aligners
by: Xing, Yazhou, et al.
Published: (2024)
by: Xing, Yazhou, et al.
Published: (2024)
SpA2V: Harnessing Spatial Auditory Cues for Audio-driven Spatially-aware Video Generation
by: Pham, Kien T., et al.
Published: (2025)
by: Pham, Kien T., et al.
Published: (2025)
Hierarchical Interaction Summarization and Contrastive Prompting for Explainable Recommendations
by: Liu, Yibin, et al.
Published: (2025)
by: Liu, Yibin, et al.
Published: (2025)
Visual Position Prompt for MLLM based Visual Grounding
by: Tang, Wei, et al.
Published: (2025)
by: Tang, Wei, et al.
Published: (2025)
Generating Synthetic Datasets for Few-shot Prompt Tuning
by: Guo, Xu, et al.
Published: (2024)
by: Guo, Xu, et al.
Published: (2024)
AudioX: A Unified Framework for Anything-to-Audio Generation
by: Tian, Zeyue, et al.
Published: (2025)
by: Tian, Zeyue, et al.
Published: (2025)
A conditional normalizing flow for domain decomposed uncertainty quantification
by: Li, Sen, et al.
Published: (2024)
by: Li, Sen, et al.
Published: (2024)
Similar Items
-
VFX Creator: Animated Visual Effect Generation with Controllable Diffusion Transformer
by: Liu, Xinyu, et al.
Published: (2025) -
Importance Weighting Can Help Large Language Models Self-Improve
by: Jiang, Chunyang, et al.
Published: (2024) -
Weakly-Supervised Emotion Transition Learning for Diverse 3D Co-speech Gesture Generation
by: Qi, Xingqun, et al.
Published: (2023) -
Make a Cheap Scaling: A Self-Cascade Diffusion Model for Higher-Resolution Adaptation
by: Guo, Lanqing, et al.
Published: (2024) -
Foundation Cures Personalization: Improving Personalized Models' Prompt Consistency via Hidden Foundation Knowledge
by: Cai, Yiyang, et al.
Published: (2024)