Self-Prompting Diffusion Transformer for Open-Vocabulary Scene Text Editing via In-Context Learning
Fuente:
arXiv
Saved in:
| Main Authors: | Li, Hongxi, Wang, Tong, Wu, Chengjing, Liu, Tianbao, Yao, Jiangtao, Qu, Xiaochao, Wu, Xinxiao, Liu, Luoqi, Liu, Ting |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
GlyphMastero: A Glyph Encoder for High-Fidelity Scene Text Editing
by: Wang, Tong, et al.
Published: (2025)
by: Wang, Tong, et al.
Published: (2025)
Learning Stochastic Bridges for Video Object Removal via Video-to-Video Translation
by: Lou, Zijie, et al.
Published: (2026)
by: Lou, Zijie, et al.
Published: (2026)
MiVE: Multiscale Vision-language features for reference-guided video Editing
by: Wang, Tong, et al.
Published: (2026)
by: Wang, Tong, et al.
Published: (2026)
TextMastero: Mastering High-Quality Scene Text Editing in Diverse Languages and Styles
by: Wang, Tong, et al.
Published: (2024)
by: Wang, Tong, et al.
Published: (2024)
SAM-REF: Introducing Image-Prompt Synergy during Interaction for Detail Enhancement in the Segment Anything Model
by: Yu, Chongkai, et al.
Published: (2024)
by: Yu, Chongkai, et al.
Published: (2024)
2nd Place Solution for MOSE Track in CVPR 2024 PVUW workshop: Complex Video Object Segmentation
by: Xu, Zhensong, et al.
Published: (2024)
by: Xu, Zhensong, et al.
Published: (2024)
EVPGS: Enhanced View Prior Guidance for Splatting-based Extrapolated View Synthesis
by: Li, Jiahe, et al.
Published: (2025)
by: Li, Jiahe, et al.
Published: (2025)
MTADiffusion: Mask Text Alignment Diffusion Model for Object Inpainting
by: Huang, Jun, et al.
Published: (2025)
by: Huang, Jun, et al.
Published: (2025)
Draw Like an Artist: Complex Scene Generation with Diffusion Model via Composition, Painting, and Retouching
by: Liu, Minghao, et al.
Published: (2024)
by: Liu, Minghao, et al.
Published: (2024)
DCEdit: Dual-Level Controlled Image Editing via Precisely Localized Semantics
by: Hu, Yihan, et al.
Published: (2025)
by: Hu, Yihan, et al.
Published: (2025)
Rethinking Video Segmentation with Masked Video Consistency: Did the Model Learn as Intended?
by: Liang, Chen, et al.
Published: (2024)
by: Liang, Chen, et al.
Published: (2024)
On Exact Editing of Flow-Based Diffusion Models
by: Li, Zixiang, et al.
Published: (2025)
by: Li, Zixiang, et al.
Published: (2025)
3rd Place Solution for PVUW Challenge 2024: Video Panoptic Segmentation
by: Wu, Ruipu, et al.
Published: (2024)
by: Wu, Ruipu, et al.
Published: (2024)
Semantic Segmentation on VSPW Dataset through Masked Video Consistency
by: Liang, Chen, et al.
Published: (2024)
by: Liang, Chen, et al.
Published: (2024)
AlignGen: Boosting Personalized Image Generation with Cross-Modality Prior Alignment
by: Lin, Yiheng, et al.
Published: (2025)
by: Lin, Yiheng, et al.
Published: (2025)
Memory Efficient Matting with Adaptive Token Routing
by: Lin, Yiheng, et al.
Published: (2024)
by: Lin, Yiheng, et al.
Published: (2024)
Video Summarization using Denoising Diffusion Probabilistic Model
by: Shang, Zirui, et al.
Published: (2024)
by: Shang, Zirui, et al.
Published: (2024)
LLM-enhanced Action-aware Multi-modal Prompt Tuning for Image-Text Matching
by: Tian, Mengxiao, et al.
Published: (2025)
by: Tian, Mengxiao, et al.
Published: (2025)
PortraitCraft: A Benchmark for Portrait Composition Understanding and Generation
by: Sha, Yuyang, et al.
Published: (2026)
by: Sha, Yuyang, et al.
Published: (2026)
DepthSSC: Monocular 3D Semantic Scene Completion via Depth-Spatial Alignment and Voxel Adaptation
by: Yao, Jiawei, et al.
Published: (2023)
by: Yao, Jiawei, et al.
Published: (2023)
End-to-end Open-vocabulary Video Visual Relationship Detection using Multi-modal Prompting
by: Wang, Yongqi, et al.
Published: (2024)
by: Wang, Yongqi, et al.
Published: (2024)
FlowSeg: Dynamic Semantic Guidance for LLM-Conditioned Segmentation
by: Zhang, Zekang, et al.
Published: (2026)
by: Zhang, Zekang, et al.
Published: (2026)
GHOST: Grounded Human Motion Generation with Open Vocabulary Scene-and-Text Contexts
by: Milacski, Zoltán Á., et al.
Published: (2024)
by: Milacski, Zoltán Á., et al.
Published: (2024)
How Vision-Language Tasks Benefit from Large Pre-trained Models: A Survey
by: Qi, Yayun, et al.
Published: (2024)
by: Qi, Yayun, et al.
Published: (2024)
HarmonicNeRF: Geometry-Informed Synthetic View Augmentation for 3D Scene Reconstruction in Driving Scenarios
by: Pan, Xiaochao, et al.
Published: (2023)
by: Pan, Xiaochao, et al.
Published: (2023)
Open-Vocabulary 3D Semantic Segmentation with Text-to-Image Diffusion Models
by: Zhu, Xiaoyu, et al.
Published: (2024)
by: Zhu, Xiaoyu, et al.
Published: (2024)
DiffTF++: 3D-aware Diffusion Transformer for Large-Vocabulary 3D Generation
by: Cao, Ziang, et al.
Published: (2024)
by: Cao, Ziang, et al.
Published: (2024)
Open-Vocabulary HOI Detection with Interaction-aware Prompt and Concept Calibration
by: Lei, Ting, et al.
Published: (2025)
by: Lei, Ting, et al.
Published: (2025)
GaussEdit: Adaptive 3D Scene Editing with Text and Image Prompts
by: Shu, Zhenyu, et al.
Published: (2025)
by: Shu, Zhenyu, et al.
Published: (2025)
Editing Massive Concepts in Text-to-Image Diffusion Models
by: Xiong, Tianwei, et al.
Published: (2024)
by: Xiong, Tianwei, et al.
Published: (2024)
In-Context Deep Learning via Transformer Models
by: Wu, Weimin, et al.
Published: (2024)
by: Wu, Weimin, et al.
Published: (2024)
Large-Vocabulary Segmentation for Medical Images with Text Prompts
by: Zhao, Ziheng, et al.
Published: (2023)
by: Zhao, Ziheng, et al.
Published: (2023)
FLUX-Text: A Simple and Advanced Diffusion Transformer Baseline for Scene Text Editing
by: Lan, Rui, et al.
Published: (2025)
by: Lan, Rui, et al.
Published: (2025)
METOR: A Unified Framework for Mutual Enhancement of Objects and Relationships in Open-vocabulary Video Visual Relationship Detection
by: Wang, Yongqi, et al.
Published: (2025)
by: Wang, Yongqi, et al.
Published: (2025)
Open-Vocabulary Functional 3D Human-Scene Interaction Generation
by: Liu, Jie, et al.
Published: (2026)
by: Liu, Jie, et al.
Published: (2026)
Classifying the Unknown: In-Context Learning for Open-Vocabulary Text and Symbol Recognition
by: Simon, Tom, et al.
Published: (2025)
by: Simon, Tom, et al.
Published: (2025)
Region-Constraint In-Context Generation for Instructional Video Editing
by: Zhang, Zhongwei, et al.
Published: (2025)
by: Zhang, Zhongwei, et al.
Published: (2025)
OpenESS: Event-based Semantic Scene Understanding with Open Vocabularies
by: Kong, Lingdong, et al.
Published: (2024)
by: Kong, Lingdong, et al.
Published: (2024)
KUDA: Keypoints to Unify Dynamics Learning and Visual Prompting for Open-Vocabulary Robotic Manipulation
by: Liu, Zixian, et al.
Published: (2025)
by: Liu, Zixian, et al.
Published: (2025)
Open-Vocabulary Scene Text Recognition via Pseudo-Image Labeling and Margin Loss
by: Ren, Xuhua, et al.
Published: (2024)
by: Ren, Xuhua, et al.
Published: (2024)
Similar Items
-
GlyphMastero: A Glyph Encoder for High-Fidelity Scene Text Editing
by: Wang, Tong, et al.
Published: (2025) -
Learning Stochastic Bridges for Video Object Removal via Video-to-Video Translation
by: Lou, Zijie, et al.
Published: (2026) -
MiVE: Multiscale Vision-language features for reference-guided video Editing
by: Wang, Tong, et al.
Published: (2026) -
TextMastero: Mastering High-Quality Scene Text Editing in Diverse Languages and Styles
by: Wang, Tong, et al.
Published: (2024) -
SAM-REF: Introducing Image-Prompt Synergy during Interaction for Detail Enhancement in the Segment Anything Model
by: Yu, Chongkai, et al.
Published: (2024)