Saved in:
| Main Authors: | Santoso, Joshua, Simon, Christian, Williem |
|---|---|
| Format: | Preprint |
| Published: |
2023
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2311.00734 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Ef-QuantFace: Streamlined Face Recognition with Small Data and Low-Bit Precision
by: Gazali, William, et al.
Published: (2024)
by: Gazali, William, et al.
Published: (2024)
SceneVTG++: Controllable Multilingual Visual Text Generation in the Wild
by: Liu, Jiawei, et al.
Published: (2025)
by: Liu, Jiawei, et al.
Published: (2025)
Scene Grounding In the Wild
by: Cohen, Tamir, et al.
Published: (2026)
by: Cohen, Tamir, et al.
Published: (2026)
JaWildText: A Benchmark for Vision-Language Models on Japanese Scene Text Understanding
by: Maeda, Koki, et al.
Published: (2026)
by: Maeda, Koki, et al.
Published: (2026)
Layout Agnostic Scene Text Image Synthesis with Diffusion Models
by: Zhangli, Qilong, et al.
Published: (2024)
by: Zhangli, Qilong, et al.
Published: (2024)
DiffSTR: Controlled Diffusion Models for Scene Text Removal
by: Pathak, Sanhita, et al.
Published: (2024)
by: Pathak, Sanhita, et al.
Published: (2024)
GeoSceneGraph: Geometric Scene Graph Diffusion Model for Text-guided 3D Indoor Scene Synthesis
by: Ruiz, Antonio, et al.
Published: (2025)
by: Ruiz, Antonio, et al.
Published: (2025)
MDiff4STR: Mask Diffusion Model for Scene Text Recognition
by: Du, Yongkun, et al.
Published: (2025)
by: Du, Yongkun, et al.
Published: (2025)
TextDiff: Mask-Guided Residual Diffusion Models for Scene Text Image Super-Resolution
by: Liu, Baolin, et al.
Published: (2023)
by: Liu, Baolin, et al.
Published: (2023)
TextSSR: Diffusion-based Data Synthesis for Scene Text Recognition
by: Ye, Xingsong, et al.
Published: (2024)
by: Ye, Xingsong, et al.
Published: (2024)
ViewDelta: Scaling Scene Change Detection through Text-Conditioning
by: Varghese, Subin, et al.
Published: (2024)
by: Varghese, Subin, et al.
Published: (2024)
TextCtrl: Diffusion-based Scene Text Editing with Prior Guidance Control
by: Zeng, Weichao, et al.
Published: (2024)
by: Zeng, Weichao, et al.
Published: (2024)
Visual Text Generation in the Wild
by: Zhu, Yuanzhi, et al.
Published: (2024)
by: Zhu, Yuanzhi, et al.
Published: (2024)
FLUX-Text: A Simple and Advanced Diffusion Transformer Baseline for Scene Text Editing
by: Lan, Rui, et al.
Published: (2025)
by: Lan, Rui, et al.
Published: (2025)
Zero-Shot Monocular Scene Flow Estimation in the Wild
by: Liang, Yiqing, et al.
Published: (2025)
by: Liang, Yiqing, et al.
Published: (2025)
$L^3$:Scene-agnostic Visual Localization in the Wild
by: Zhang, Yu, et al.
Published: (2026)
by: Zhang, Yu, et al.
Published: (2026)
SceneTextStylizer: A Training-Free Scene Text Style Transfer Framework with Diffusion Model
by: Yuan, Honghui, et al.
Published: (2025)
by: Yuan, Honghui, et al.
Published: (2025)
EmbodMocap: In-the-Wild 4D Human-Scene Reconstruction for Embodied Agents
by: Wang, Wenjia, et al.
Published: (2026)
by: Wang, Wenjia, et al.
Published: (2026)
Improving Diffusion Models for Authentic Virtual Try-on in the Wild
by: Choi, Yisol, et al.
Published: (2024)
by: Choi, Yisol, et al.
Published: (2024)
Evaluating Hallucination in Text-to-Image Diffusion Models with Scene-Graph based Question-Answering Agent
by: Qin, Ziyuan, et al.
Published: (2024)
by: Qin, Ziyuan, et al.
Published: (2024)
Anchored Video Generation: Decoupling Scene Construction and Temporal Synthesis in Text-to-Video Diffusion Models
by: Hassan, Mariam, et al.
Published: (2025)
by: Hassan, Mariam, et al.
Published: (2025)
IPAD: Iterative, Parallel, and Diffusion-based Network for Scene Text Recognition
by: Yang, Xiaomeng, et al.
Published: (2023)
by: Yang, Xiaomeng, et al.
Published: (2023)
WildLMa: Long Horizon Loco-Manipulation in the Wild
by: Qiu, Ri-Zhao, et al.
Published: (2024)
by: Qiu, Ri-Zhao, et al.
Published: (2024)
Joint Optimization for 4D Human-Scene Reconstruction in the Wild
by: Liu, Zhizheng, et al.
Published: (2025)
by: Liu, Zhizheng, et al.
Published: (2025)
WildVidFit: Video Virtual Try-On in the Wild via Image-Based Controlled Diffusion Models
by: He, Zijian, et al.
Published: (2024)
by: He, Zijian, et al.
Published: (2024)
ChildPlay-Hand: A Dataset of Hand Manipulations in the Wild
by: Farkhondeh, Arya, et al.
Published: (2024)
by: Farkhondeh, Arya, et al.
Published: (2024)
A Self-Conditioned Representation Guided Diffusion Model for Realistic Text-to-LiDAR Scene Generation
by: Qu, Wentao, et al.
Published: (2025)
by: Qu, Wentao, et al.
Published: (2025)
TITAN-Guide: Taming Inference-Time AligNment for Guided Text-to-Video Diffusion Models
by: Simon, Christian, et al.
Published: (2025)
by: Simon, Christian, et al.
Published: (2025)
SceneBooth: Diffusion-based Framework for Subject-preserved Text-to-Image Generation
by: Chai, Shang, et al.
Published: (2025)
by: Chai, Shang, et al.
Published: (2025)
DiffuScene: Denoising Diffusion Models for Generative Indoor Scene Synthesis
by: Tang, Jiapeng, et al.
Published: (2023)
by: Tang, Jiapeng, et al.
Published: (2023)
WildCAT3D: Appearance-Aware Multi-View Diffusion in the Wild
by: Alper, Morris, et al.
Published: (2025)
by: Alper, Morris, et al.
Published: (2025)
TDEdit: A Unified Diffusion Framework for Text-Drag Guided Image Manipulation
by: Wang, Qihang, et al.
Published: (2025)
by: Wang, Qihang, et al.
Published: (2025)
Prometheus: 3D-Aware Latent Diffusion Models for Feed-Forward Text-to-3D Scene Generation
by: Yang, Yuanbo, et al.
Published: (2024)
by: Yang, Yuanbo, et al.
Published: (2024)
Aggregated Text Transformer for Scene Text Detection
by: Zhou, Zhao, et al.
Published: (2022)
by: Zhou, Zhao, et al.
Published: (2022)
Leveraging Text Localization for Scene Text Removal via Text-aware Masked Image Modeling
by: Wang, Zixiao, et al.
Published: (2024)
by: Wang, Zixiao, et al.
Published: (2024)
MOoSE: Multi-Orientation Sharing Experts for Open-set Scene Text Recognition
by: Liu, Chang, et al.
Published: (2024)
by: Liu, Chang, et al.
Published: (2024)
WildScenes: A Benchmark for 2D and 3D Semantic Segmentation in Large-scale Natural Environments
by: Vidanapathirana, Kavisha, et al.
Published: (2023)
by: Vidanapathirana, Kavisha, et al.
Published: (2023)
Partial Scene Text Retrieval
by: Wang, Hao, et al.
Published: (2024)
by: Wang, Hao, et al.
Published: (2024)
Inverse Scene Text Removal
by: Yoshimatsu, Takumi, et al.
Published: (2025)
by: Yoshimatsu, Takumi, et al.
Published: (2025)
R3GW: Relightable 3D Gaussians for Outdoor Scenes in the Wild
by: Corona, Margherita Lea, et al.
Published: (2026)
by: Corona, Margherita Lea, et al.
Published: (2026)
Similar Items
-
Ef-QuantFace: Streamlined Face Recognition with Small Data and Low-Bit Precision
by: Gazali, William, et al.
Published: (2024) -
SceneVTG++: Controllable Multilingual Visual Text Generation in the Wild
by: Liu, Jiawei, et al.
Published: (2025) -
Scene Grounding In the Wild
by: Cohen, Tamir, et al.
Published: (2026) -
JaWildText: A Benchmark for Vision-Language Models on Japanese Scene Text Understanding
by: Maeda, Koki, et al.
Published: (2026) -
Layout Agnostic Scene Text Image Synthesis with Diffusion Models
by: Zhangli, Qilong, et al.
Published: (2024)