On Manipulating Scene Text in the Wild with Diffusion Models
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Santoso, Joshua, Simon, Christian, Williem |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2023
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Ef-QuantFace: Streamlined Face Recognition with Small Data and Low-Bit Precision
von: Gazali, William, et al.
Veröffentlicht: (2024)
von: Gazali, William, et al.
Veröffentlicht: (2024)
SceneVTG++: Controllable Multilingual Visual Text Generation in the Wild
von: Liu, Jiawei, et al.
Veröffentlicht: (2025)
von: Liu, Jiawei, et al.
Veröffentlicht: (2025)
Scene Grounding In the Wild
von: Cohen, Tamir, et al.
Veröffentlicht: (2026)
von: Cohen, Tamir, et al.
Veröffentlicht: (2026)
JaWildText: A Benchmark for Vision-Language Models on Japanese Scene Text Understanding
von: Maeda, Koki, et al.
Veröffentlicht: (2026)
von: Maeda, Koki, et al.
Veröffentlicht: (2026)
Layout Agnostic Scene Text Image Synthesis with Diffusion Models
von: Zhangli, Qilong, et al.
Veröffentlicht: (2024)
von: Zhangli, Qilong, et al.
Veröffentlicht: (2024)
DiffSTR: Controlled Diffusion Models for Scene Text Removal
von: Pathak, Sanhita, et al.
Veröffentlicht: (2024)
von: Pathak, Sanhita, et al.
Veröffentlicht: (2024)
GeoSceneGraph: Geometric Scene Graph Diffusion Model for Text-guided 3D Indoor Scene Synthesis
von: Ruiz, Antonio, et al.
Veröffentlicht: (2025)
von: Ruiz, Antonio, et al.
Veröffentlicht: (2025)
MDiff4STR: Mask Diffusion Model for Scene Text Recognition
von: Du, Yongkun, et al.
Veröffentlicht: (2025)
von: Du, Yongkun, et al.
Veröffentlicht: (2025)
TextDiff: Mask-Guided Residual Diffusion Models for Scene Text Image Super-Resolution
von: Liu, Baolin, et al.
Veröffentlicht: (2023)
von: Liu, Baolin, et al.
Veröffentlicht: (2023)
TextSSR: Diffusion-based Data Synthesis for Scene Text Recognition
von: Ye, Xingsong, et al.
Veröffentlicht: (2024)
von: Ye, Xingsong, et al.
Veröffentlicht: (2024)
ViewDelta: Scaling Scene Change Detection through Text-Conditioning
von: Varghese, Subin, et al.
Veröffentlicht: (2024)
von: Varghese, Subin, et al.
Veröffentlicht: (2024)
TextCtrl: Diffusion-based Scene Text Editing with Prior Guidance Control
von: Zeng, Weichao, et al.
Veröffentlicht: (2024)
von: Zeng, Weichao, et al.
Veröffentlicht: (2024)
Visual Text Generation in the Wild
von: Zhu, Yuanzhi, et al.
Veröffentlicht: (2024)
von: Zhu, Yuanzhi, et al.
Veröffentlicht: (2024)
FLUX-Text: A Simple and Advanced Diffusion Transformer Baseline for Scene Text Editing
von: Lan, Rui, et al.
Veröffentlicht: (2025)
von: Lan, Rui, et al.
Veröffentlicht: (2025)
Zero-Shot Monocular Scene Flow Estimation in the Wild
von: Liang, Yiqing, et al.
Veröffentlicht: (2025)
von: Liang, Yiqing, et al.
Veröffentlicht: (2025)
$L^3$:Scene-agnostic Visual Localization in the Wild
von: Zhang, Yu, et al.
Veröffentlicht: (2026)
von: Zhang, Yu, et al.
Veröffentlicht: (2026)
Evaluating Hallucination in Text-to-Image Diffusion Models with Scene-Graph based Question-Answering Agent
von: Qin, Ziyuan, et al.
Veröffentlicht: (2024)
von: Qin, Ziyuan, et al.
Veröffentlicht: (2024)
Anchored Video Generation: Decoupling Scene Construction and Temporal Synthesis in Text-to-Video Diffusion Models
von: Hassan, Mariam, et al.
Veröffentlicht: (2025)
von: Hassan, Mariam, et al.
Veröffentlicht: (2025)
EmbodMocap: In-the-Wild 4D Human-Scene Reconstruction for Embodied Agents
von: Wang, Wenjia, et al.
Veröffentlicht: (2026)
von: Wang, Wenjia, et al.
Veröffentlicht: (2026)
IPAD: Iterative, Parallel, and Diffusion-based Network for Scene Text Recognition
von: Yang, Xiaomeng, et al.
Veröffentlicht: (2023)
von: Yang, Xiaomeng, et al.
Veröffentlicht: (2023)
Improving Diffusion Models for Authentic Virtual Try-on in the Wild
von: Choi, Yisol, et al.
Veröffentlicht: (2024)
von: Choi, Yisol, et al.
Veröffentlicht: (2024)
SceneTextStylizer: A Training-Free Scene Text Style Transfer Framework with Diffusion Model
von: Yuan, Honghui, et al.
Veröffentlicht: (2025)
von: Yuan, Honghui, et al.
Veröffentlicht: (2025)
Joint Optimization for 4D Human-Scene Reconstruction in the Wild
von: Liu, Zhizheng, et al.
Veröffentlicht: (2025)
von: Liu, Zhizheng, et al.
Veröffentlicht: (2025)
WildVidFit: Video Virtual Try-On in the Wild via Image-Based Controlled Diffusion Models
von: He, Zijian, et al.
Veröffentlicht: (2024)
von: He, Zijian, et al.
Veröffentlicht: (2024)
A Self-Conditioned Representation Guided Diffusion Model for Realistic Text-to-LiDAR Scene Generation
von: Qu, Wentao, et al.
Veröffentlicht: (2025)
von: Qu, Wentao, et al.
Veröffentlicht: (2025)
SceneBooth: Diffusion-based Framework for Subject-preserved Text-to-Image Generation
von: Chai, Shang, et al.
Veröffentlicht: (2025)
von: Chai, Shang, et al.
Veröffentlicht: (2025)
ChildPlay-Hand: A Dataset of Hand Manipulations in the Wild
von: Farkhondeh, Arya, et al.
Veröffentlicht: (2024)
von: Farkhondeh, Arya, et al.
Veröffentlicht: (2024)
TITAN-Guide: Taming Inference-Time AligNment for Guided Text-to-Video Diffusion Models
von: Simon, Christian, et al.
Veröffentlicht: (2025)
von: Simon, Christian, et al.
Veröffentlicht: (2025)
DiffuScene: Denoising Diffusion Models for Generative Indoor Scene Synthesis
von: Tang, Jiapeng, et al.
Veröffentlicht: (2023)
von: Tang, Jiapeng, et al.
Veröffentlicht: (2023)
TDEdit: A Unified Diffusion Framework for Text-Drag Guided Image Manipulation
von: Wang, Qihang, et al.
Veröffentlicht: (2025)
von: Wang, Qihang, et al.
Veröffentlicht: (2025)
Prometheus: 3D-Aware Latent Diffusion Models for Feed-Forward Text-to-3D Scene Generation
von: Yang, Yuanbo, et al.
Veröffentlicht: (2024)
von: Yang, Yuanbo, et al.
Veröffentlicht: (2024)
WildCAT3D: Appearance-Aware Multi-View Diffusion in the Wild
von: Alper, Morris, et al.
Veröffentlicht: (2025)
von: Alper, Morris, et al.
Veröffentlicht: (2025)
Leveraging Text Localization for Scene Text Removal via Text-aware Masked Image Modeling
von: Wang, Zixiao, et al.
Veröffentlicht: (2024)
von: Wang, Zixiao, et al.
Veröffentlicht: (2024)
WildLMa: Long Horizon Loco-Manipulation in the Wild
von: Qiu, Ri-Zhao, et al.
Veröffentlicht: (2024)
von: Qiu, Ri-Zhao, et al.
Veröffentlicht: (2024)
Aggregated Text Transformer for Scene Text Detection
von: Zhou, Zhao, et al.
Veröffentlicht: (2022)
von: Zhou, Zhao, et al.
Veröffentlicht: (2022)
MOoSE: Multi-Orientation Sharing Experts for Open-set Scene Text Recognition
von: Liu, Chang, et al.
Veröffentlicht: (2024)
von: Liu, Chang, et al.
Veröffentlicht: (2024)
Partial Scene Text Retrieval
von: Wang, Hao, et al.
Veröffentlicht: (2024)
von: Wang, Hao, et al.
Veröffentlicht: (2024)
Inverse Scene Text Removal
von: Yoshimatsu, Takumi, et al.
Veröffentlicht: (2025)
von: Yoshimatsu, Takumi, et al.
Veröffentlicht: (2025)
R3GW: Relightable 3D Gaussians for Outdoor Scenes in the Wild
von: Corona, Margherita Lea, et al.
Veröffentlicht: (2026)
von: Corona, Margherita Lea, et al.
Veröffentlicht: (2026)
Audio-Visual Camera Pose Estimation with Passive Scene Sounds and In-the-Wild Video
von: Adebi, Daniel, et al.
Veröffentlicht: (2025)
von: Adebi, Daniel, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Ef-QuantFace: Streamlined Face Recognition with Small Data and Low-Bit Precision
von: Gazali, William, et al.
Veröffentlicht: (2024) -
SceneVTG++: Controllable Multilingual Visual Text Generation in the Wild
von: Liu, Jiawei, et al.
Veröffentlicht: (2025) -
Scene Grounding In the Wild
von: Cohen, Tamir, et al.
Veröffentlicht: (2026) -
JaWildText: A Benchmark for Vision-Language Models on Japanese Scene Text Understanding
von: Maeda, Koki, et al.
Veröffentlicht: (2026) -
Layout Agnostic Scene Text Image Synthesis with Diffusion Models
von: Zhangli, Qilong, et al.
Veröffentlicht: (2024)