Saved in:
| Main Authors: | Yao, Yuxuan, Chen, Yuxuan, Li, Hui, Cheng, Kaihui, Guo, Qipeng, Sun, Yuwei, Dong, Zilong, Wang, Jingdong, Zhu, Siyu |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2602.06886 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
The Thinking Pixel: Recursive Sparse Reasoning in Multimodal Diffusion Latents
by: Sun, Yuwei, et al.
Published: (2026)
by: Sun, Yuwei, et al.
Published: (2026)
MixFlow Training: Alleviating Exposure Bias with Slowed Interpolation Mixture
by: Li, Hui, et al.
Published: (2025)
by: Li, Hui, et al.
Published: (2025)
Hallo4: High-Fidelity Dynamic Portrait Animation via Direct Preference Optimization
by: Cui, Jiahao, et al.
Published: (2025)
by: Cui, Jiahao, et al.
Published: (2025)
OpenHumanVid: A Large-Scale High-Quality Dataset for Enhancing Human-Centric Video Generation
by: Li, Hui, et al.
Published: (2024)
by: Li, Hui, et al.
Published: (2024)
Head-wise Adaptive Rotary Positional Encoding for Fine-Grained Image Generation
by: Li, Jiaye, et al.
Published: (2025)
by: Li, Jiaye, et al.
Published: (2025)
Hallo2: Long-Duration and High-Resolution Audio-Driven Portrait Image Animation
by: Cui, Jiahao, et al.
Published: (2024)
by: Cui, Jiahao, et al.
Published: (2024)
Hallo3: Highly Dynamic and Realistic Portrait Image Animation with Video Diffusion Transformer
by: Cui, Jiahao, et al.
Published: (2024)
by: Cui, Jiahao, et al.
Published: (2024)
UniAPO: Unified Multimodal Automated Prompt Optimization
by: Zhu, Qipeng, et al.
Published: (2025)
by: Zhu, Qipeng, et al.
Published: (2025)
DicFace: Dirichlet-Constrained Variational Codebook Learning for Temporally Coherent Video Face Restoration
by: Chen, Yan, et al.
Published: (2025)
by: Chen, Yan, et al.
Published: (2025)
Panther: Illuminate the Sight of Multimodal LLMs with Instruction-Guided Visual Prompts
by: Li, Honglin, et al.
Published: (2024)
by: Li, Honglin, et al.
Published: (2024)
ModalPrompt: Towards Efficient Multimodal Continual Instruction Tuning with Dual-Modality Guided Prompt
by: Zeng, Fanhu, et al.
Published: (2024)
by: Zeng, Fanhu, et al.
Published: (2024)
Pixel-Perfect Depth with Semantics-Prompted Diffusion Transformers
by: Xu, Gangwei, et al.
Published: (2025)
by: Xu, Gangwei, et al.
Published: (2025)
Unleashing the Power of Prompt-driven Nucleus Instance Segmentation
by: Shui, Zhongyi, et al.
Published: (2023)
by: Shui, Zhongyi, et al.
Published: (2023)
Skip and Skip: Segmenting Medical Images with Prompts
by: Chen, Jiawei, et al.
Published: (2024)
by: Chen, Jiawei, et al.
Published: (2024)
MoviePuzzle: Visual Narrative Reasoning through Multimodal Order Learning
by: Wang, Jianghui, et al.
Published: (2023)
by: Wang, Jianghui, et al.
Published: (2023)
The CLIP Model is Secretly an Image-to-Prompt Converter
by: Ding, Yuxuan, et al.
Published: (2023)
by: Ding, Yuxuan, et al.
Published: (2023)
TP-Eval: Tap Multimodal LLMs' Potential in Evaluation by Customizing Prompts
by: Xie, Yuxuan, et al.
Published: (2024)
by: Xie, Yuxuan, et al.
Published: (2024)
Prompting Forgetting: Unlearning in GANs via Textual Guidance
by: Nagasubramaniam, Piyush, et al.
Published: (2025)
by: Nagasubramaniam, Piyush, et al.
Published: (2025)
DPStyler: Dynamic PromptStyler for Source-Free Domain Generalization
by: Tang, Yunlong, et al.
Published: (2024)
by: Tang, Yunlong, et al.
Published: (2024)
Exploring Multimodal Diffusion Transformers for Enhanced Prompt-based Image Editing
by: Shin, Joonghyuk, et al.
Published: (2025)
by: Shin, Joonghyuk, et al.
Published: (2025)
SPT: Sequence Prompt Transformer for Interactive Image Segmentation
by: Cheng, Senlin, et al.
Published: (2024)
by: Cheng, Senlin, et al.
Published: (2024)
OmniMotion: Multimodal Motion Generation with Continuous Masked Autoregression
by: Li, Zhe, et al.
Published: (2025)
by: Li, Zhe, et al.
Published: (2025)
Training-free Regional Prompting for Diffusion Transformers
by: Chen, Anthony, et al.
Published: (2024)
by: Chen, Anthony, et al.
Published: (2024)
EVF-SAM: Early Vision-Language Fusion for Text-Prompted Segment Anything Model
by: Zhang, Yuxuan, et al.
Published: (2024)
by: Zhang, Yuxuan, et al.
Published: (2024)
Open-Set Video-based Facial Expression Recognition with Human Expression-sensitive Prompting
by: Liu, Yuanyuan, et al.
Published: (2024)
by: Liu, Yuanyuan, et al.
Published: (2024)
UAOR: Uncertainty-aware Observation Reinjection for Vision-Language-Action Models
by: Yang, Jiabing, et al.
Published: (2026)
by: Yang, Jiabing, et al.
Published: (2026)
VRP-SAM: SAM with Visual Reference Prompt
by: Sun, Yanpeng, et al.
Published: (2024)
by: Sun, Yanpeng, et al.
Published: (2024)
Local-Prompt: Extensible Local Prompts for Few-Shot Out-of-Distribution Detection
by: Zeng, Fanhu, et al.
Published: (2024)
by: Zeng, Fanhu, et al.
Published: (2024)
SCING:Towards More Efficient and Robust Person Re-Identification through Selective Cross-modal Prompt Tuning
by: Xie, Yunfei, et al.
Published: (2025)
by: Xie, Yunfei, et al.
Published: (2025)
Test-Time Multimodal Backdoor Detection by Contrastive Prompting
by: Niu, Yuwei, et al.
Published: (2024)
by: Niu, Yuwei, et al.
Published: (2024)
Prompt-Guided Adaptive Model Transformation for Whole Slide Image Classification
by: Lin, Yi, et al.
Published: (2024)
by: Lin, Yi, et al.
Published: (2024)
Pyramidal Patchification Flow for Visual Generation
by: Li, Hui, et al.
Published: (2025)
by: Li, Hui, et al.
Published: (2025)
Video Virtual Try-on with Conditional Diffusion Transformer Inpainter
by: Zou, Cheng, et al.
Published: (2025)
by: Zou, Cheng, et al.
Published: (2025)
Diff-Prompt: Diffusion-Driven Prompt Generator with Mask Supervision
by: Yan, Weicai, et al.
Published: (2025)
by: Yan, Weicai, et al.
Published: (2025)
ExoViP: Step-by-step Verification and Exploration with Exoskeleton Modules for Compositional Visual Reasoning
by: Wang, Yuxuan, et al.
Published: (2024)
by: Wang, Yuxuan, et al.
Published: (2024)
Video-As-Prompt: Unified Semantic Control for Video Generation
by: Bian, Yuxuan, et al.
Published: (2025)
by: Bian, Yuxuan, et al.
Published: (2025)
Unsupervised Pre-training with Language-Vision Prompts for Low-Data Instance Segmentation
by: Zhang, Dingwen, et al.
Published: (2024)
by: Zhang, Dingwen, et al.
Published: (2024)
Prompt Diffusion Robustifies Any-Modality Prompt Learning
by: Du, Yingjun, et al.
Published: (2024)
by: Du, Yingjun, et al.
Published: (2024)
Exploring Interpretability for Visual Prompt Tuning with Cross-layer Concepts
by: Wang, Yubin, et al.
Published: (2025)
by: Wang, Yubin, et al.
Published: (2025)
See&Trek: Training-Free Spatial Prompting for Multimodal Large Language Model
by: Li, Pengteng, et al.
Published: (2025)
by: Li, Pengteng, et al.
Published: (2025)
Similar Items
-
The Thinking Pixel: Recursive Sparse Reasoning in Multimodal Diffusion Latents
by: Sun, Yuwei, et al.
Published: (2026) -
MixFlow Training: Alleviating Exposure Bias with Slowed Interpolation Mixture
by: Li, Hui, et al.
Published: (2025) -
Hallo4: High-Fidelity Dynamic Portrait Animation via Direct Preference Optimization
by: Cui, Jiahao, et al.
Published: (2025) -
OpenHumanVid: A Large-Scale High-Quality Dataset for Enhancing Human-Centric Video Generation
by: Li, Hui, et al.
Published: (2024) -
Head-wise Adaptive Rotary Positional Encoding for Fine-Grained Image Generation
by: Li, Jiaye, et al.
Published: (2025)