Watch Your Steps: Local Image and Scene Editing by Text Instructions
Fuente:
arXiv
Saved in:
| Main Authors: | Mirzaei, Ashkan, Aumentado-Armstrong, Tristan, Brubaker, Marcus A., Kelly, Jonathan, Levinshtein, Alex, Derpanis, Konstantinos G., Gilitschenski, Igor |
|---|---|
| Format: | Preprint |
| Published: |
2023
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Geometry-Aware Diffusion Models for Multiview Scene Inpainting
by: Salimi, Ahmad, et al.
Published: (2025)
by: Salimi, Ahmad, et al.
Published: (2025)
Learn Your Scales: Towards Scale-Consistent Generative Novel View Synthesis
by: Forghani, Fereshteh, et al.
Published: (2025)
by: Forghani, Fereshteh, et al.
Published: (2025)
PolyOculus: Simultaneous Multi-view Image-based Novel View Synthesis
by: Yu, Jason J., et al.
Published: (2024)
by: Yu, Jason J., et al.
Published: (2024)
Face2Scene: Using Facial Degradation as an Oracle for Diffusion-Based Scene Restoration
by: Kazerouni, Amirhossein, et al.
Published: (2026)
by: Kazerouni, Amirhossein, et al.
Published: (2026)
Augmenting Perceptual Super-Resolution via Image Quality Predictors
by: Zhang, Fengjia, et al.
Published: (2025)
by: Zhang, Fengjia, et al.
Published: (2025)
Track, Inpaint, Resplat: Subject-driven 3D and 4D Generation with Progressive Texture Infilling
by: Zheng, Shuhong, et al.
Published: (2025)
by: Zheng, Shuhong, et al.
Published: (2025)
GaussianCut: Interactive segmentation via graph cut for 3D Gaussian Splatting
by: Jain, Umangi, et al.
Published: (2024)
by: Jain, Umangi, et al.
Published: (2024)
EventSplat: 3D Gaussian Splatting from Moving Event Cameras for Real-time Rendering
by: Yura, Toshiya, et al.
Published: (2024)
by: Yura, Toshiya, et al.
Published: (2024)
Hallucination Score: Towards Mitigating Hallucinations in Generative Image Super-Resolution
by: Ren, Weiming, et al.
Published: (2025)
by: Ren, Weiming, et al.
Published: (2025)
BurstGP: Enhancing Raw Burst Image Super Resolution with Generative Priors
by: Huo, Dong, et al.
Published: (2026)
by: Huo, Dong, et al.
Published: (2026)
Towards Unsupervised Blind Face Restoration using Diffusion Prior
by: Kuai, Tianshu, et al.
Published: (2024)
by: Kuai, Tianshu, et al.
Published: (2024)
Towards High-Fidelity Gaussian Splatting with Queried-Convolution Neural Networks
by: Kumar, Abhinav, et al.
Published: (2025)
by: Kumar, Abhinav, et al.
Published: (2025)
RefFusion: Reference Adapted Diffusion Models for 3D Scene Inpainting
by: Mirzaei, Ashkan, et al.
Published: (2024)
by: Mirzaei, Ashkan, et al.
Published: (2024)
CTRL-D: Controllable Dynamic 3D Scene Editing with Personalized 2D Diffusion
by: He, Kai, et al.
Published: (2024)
by: He, Kai, et al.
Published: (2024)
Probabilistic Directed Distance Fields for Ray-Based Shape Representations
by: Aumentado-Armstrong, Tristan, et al.
Published: (2024)
by: Aumentado-Armstrong, Tristan, et al.
Published: (2024)
Feed-Forward Bullet-Time Reconstruction of Dynamic Scenes from Monocular Videos
by: Liang, Hanxue, et al.
Published: (2024)
by: Liang, Hanxue, et al.
Published: (2024)
S$^2$Edit: Text-Guided Image Editing with Precise Semantic and Spatial Control
by: Liu, Xudong, et al.
Published: (2025)
by: Liu, Xudong, et al.
Published: (2025)
DenseDPO: Fine-Grained Temporal Preference Optimization for Video Diffusion Models
by: Wu, Ziyi, et al.
Published: (2025)
by: Wu, Ziyi, et al.
Published: (2025)
Watch Your Steps: Observable and Modular Chains of Thought
by: Cohen, Cassandra A., et al.
Published: (2024)
by: Cohen, Cassandra A., et al.
Published: (2024)
Watch Your Step: Optimal Retrieval for Continual Learning at Scale
by: Hickok, Truman, et al.
Published: (2024)
by: Hickok, Truman, et al.
Published: (2024)
Global-Local Aware Scene Text Editing
by: Yang, Fuxiang, et al.
Published: (2025)
by: Yang, Fuxiang, et al.
Published: (2025)
EasyV2V: A High-quality Instruction-based Video Editing Framework
by: Mai, Jinjie, et al.
Published: (2025)
by: Mai, Jinjie, et al.
Published: (2025)
Watch Your Step: Learning Semantically-Guided Locomotion in Cluttered Environment
by: Liang, Denan, et al.
Published: (2026)
by: Liang, Denan, et al.
Published: (2026)
Visual Concept Connectome (VCC): Open World Concept Discovery and their Interlayer Connections in Deep Models
by: Kowal, Matthew, et al.
Published: (2024)
by: Kowal, Matthew, et al.
Published: (2024)
Is Byzantine Studies a Colonialist Discipline? Toward a Critical Historiography. International Center for Medieval Art, Viewpoints. Edited by BenjaminAnderson and MirelaIvanova. The Pennsylvania State University Press: University Park. 2023. $24.95. xvi + 200 pp. ISBN 9780271095264.
by: Leslie Brubaker
Published: (2024)
by: Leslie Brubaker
Published: (2024)
Dude, everyone wants pattern analysis tools (DEWPAT): Tools for measuring visual pattern complexity from digital images
by: Jillian A. Sanderson, et al.
Published: (2025)
by: Jillian A. Sanderson, et al.
Published: (2025)
Watch Your Step: Information Injection in Diffusion Models via Shadow Timestep Embedding
by: Huang, An, et al.
Published: (2026)
by: Huang, An, et al.
Published: (2026)
TPIE: Topology-Preserved Image Editing With Text Instructions
by: Jayakumar, Nivetha, et al.
Published: (2024)
by: Jayakumar, Nivetha, et al.
Published: (2024)
LatentEditor: Text Driven Local Editing of 3D Scenes
by: Khalid, Umar, et al.
Published: (2023)
by: Khalid, Umar, et al.
Published: (2023)
VibES: Induced Vibration for Persistent Event-Based Sensing
by: Polizzi, Vincenzo, et al.
Published: (2025)
by: Polizzi, Vincenzo, et al.
Published: (2025)
Revisiting Image Fusion for Multi-Illuminant White-Balance Correction
by: Serrano-Lozano, David, et al.
Published: (2025)
by: Serrano-Lozano, David, et al.
Published: (2025)
Text2Traffic: A Text-to-Image Generation and Editing Method for Traffic Scenes
by: Lv, Feng, et al.
Published: (2025)
by: Lv, Feng, et al.
Published: (2025)
Recognition-Synergistic Scene Text Editing
by: Fang, Zhengyao, et al.
Published: (2025)
by: Fang, Zhengyao, et al.
Published: (2025)
Instruction-based Image Manipulation by Watching How Things Move
by: Cao, Mingdeng, et al.
Published: (2024)
by: Cao, Mingdeng, et al.
Published: (2024)
Universal Sparse Autoencoders: Interpretable Cross-Model Concept Alignment
by: Thasarathan, Harrish, et al.
Published: (2025)
by: Thasarathan, Harrish, et al.
Published: (2025)
ShowHowTo: Generating Scene-Conditioned Step-by-Step Visual Instructions
by: Souček, Tomáš, et al.
Published: (2024)
by: Souček, Tomáš, et al.
Published: (2024)
TextSculptor: Training and Benchmarking Scene Text Editing
by: Lin, Yiheng, et al.
Published: (2026)
by: Lin, Yiheng, et al.
Published: (2026)
CLIPDrag: Combining Text-based and Drag-based Instructions for Image Editing
by: Jiang, Ziqi, et al.
Published: (2024)
by: Jiang, Ziqi, et al.
Published: (2024)
TESPEC: Temporally-Enhanced Self-Supervised Pretraining for Event Cameras
by: Mohammadi, Mohammad, et al.
Published: (2025)
by: Mohammadi, Mohammad, et al.
Published: (2025)
GeoMatch++: Morphology Conditioned Geometry Matching for Multi-Embodiment Grasping
by: Wei, Yunze, et al.
Published: (2024)
by: Wei, Yunze, et al.
Published: (2024)
Similar Items
-
Geometry-Aware Diffusion Models for Multiview Scene Inpainting
by: Salimi, Ahmad, et al.
Published: (2025) -
Learn Your Scales: Towards Scale-Consistent Generative Novel View Synthesis
by: Forghani, Fereshteh, et al.
Published: (2025) -
PolyOculus: Simultaneous Multi-view Image-based Novel View Synthesis
by: Yu, Jason J., et al.
Published: (2024) -
Face2Scene: Using Facial Degradation as an Oracle for Diffusion-Based Scene Restoration
by: Kazerouni, Amirhossein, et al.
Published: (2026) -
Augmenting Perceptual Super-Resolution via Image Quality Predictors
by: Zhang, Fengjia, et al.
Published: (2025)