From Statics to Dynamics: Physics-Aware Image Editing with Latent Transition Priors
Fuente:
arXiv
Saved in:
| Main Authors: | Zhao, Liangbing, Zhuo, Le, Paul, Sayak, Li, Hongsheng, Elhoseiny, Mohamed |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
From Reflection to Perfection: Scaling Inference-Time Optimization for Text-to-Image Diffusion Models via Reflection Tuning
by: Zhuo, Le, et al.
Published: (2025)
by: Zhuo, Le, et al.
Published: (2025)
ToddlerDiffusion: Interactive Structured Image Generation with Cascaded Schrödinger Bridge
by: Abdelrahman, Eslam, et al.
Published: (2023)
by: Abdelrahman, Eslam, et al.
Published: (2023)
Factuality Matters: When Image Generation and Editing Meet Structured Visuals
by: Zhuo, Le, et al.
Published: (2025)
by: Zhuo, Le, et al.
Published: (2025)
VRSBench: A Versatile Vision-Language Benchmark Dataset for Remote Sensing Image Understanding
by: Li, Xiang, et al.
Published: (2024)
by: Li, Xiang, et al.
Published: (2024)
WikiAutoGen: Towards Multi-Modal Wikipedia-Style Article Generation
by: Yang, Zhongyu, et al.
Published: (2025)
by: Yang, Zhongyu, et al.
Published: (2025)
Domain-Aware Continual Zero-Shot Learning
by: Yi, Kai, et al.
Published: (2021)
by: Yi, Kai, et al.
Published: (2021)
StoryGPT-V: Large Language Models as Consistent Story Visualizers
by: Shen, Xiaoqian, et al.
Published: (2023)
by: Shen, Xiaoqian, et al.
Published: (2023)
PIXART-δ: Fast and Controllable Image Generation with Latent Consistency Models
by: Chen, Junsong, et al.
Published: (2024)
by: Chen, Junsong, et al.
Published: (2024)
From Static to Dynamic: Adapting Landmark-Aware Image Models for Facial Expression Recognition in Videos
by: Chen, Yin, et al.
Published: (2023)
by: Chen, Yin, et al.
Published: (2023)
PICABench: How Far Are We from Physically Realistic Image Editing?
by: Pu, Yuandong, et al.
Published: (2025)
by: Pu, Yuandong, et al.
Published: (2025)
Overcoming Generic Knowledge Loss with Selective Parameter Update
by: Zhang, Wenxuan, et al.
Published: (2023)
by: Zhang, Wenxuan, et al.
Published: (2023)
Kestrel: 3D Multimodal LLM for Part-Aware Grounded Description
by: Ahmed, Mahmoud, et al.
Published: (2024)
by: Ahmed, Mahmoud, et al.
Published: (2024)
Degradation-Aware All-in-One Image Restoration via Latent Prior Encoding
by: Sharif, S M A, et al.
Published: (2025)
by: Sharif, S M A, et al.
Published: (2025)
InfinityStory: Unlimited Video Generation with World Consistency and Character-Aware Shot Transitions
by: Elmoghany, Mohamed, et al.
Published: (2026)
by: Elmoghany, Mohamed, et al.
Published: (2026)
FDS: Frequency-Aware Denoising Score for Text-Guided Latent Diffusion Image Editing
by: Ren, Yufan, et al.
Published: (2025)
by: Ren, Yufan, et al.
Published: (2025)
Efficient Self-supervised Vision Pretraining with Local Masked Reconstruction
by: Chen, Jun, et al.
Published: (2022)
by: Chen, Jun, et al.
Published: (2022)
From Static to Dynamic: Exploring Self-supervised Image-to-Video Representation Transfer Learning
by: Liu, Yang, et al.
Published: (2026)
by: Liu, Yang, et al.
Published: (2026)
PIXELS: Progressive Image Xemplar-based Editing with Latent Surgery
by: Biswas, Shristi Das, et al.
Published: (2025)
by: Biswas, Shristi Das, et al.
Published: (2025)
Image-to-Image Translation with Disentangled Latent Vectors for Face Editing
by: Dalva, Yusuf, et al.
Published: (2023)
by: Dalva, Yusuf, et al.
Published: (2023)
FramePainter: Endowing Interactive Image Editing with Video Diffusion Priors
by: Zhang, Yabo, et al.
Published: (2025)
by: Zhang, Yabo, et al.
Published: (2025)
Category-level Text-to-Image Retrieval Improved: Bridging the Domain Gap with Diffusion Models and Vision Encoders
by: Khan, Faizan Farooq, et al.
Published: (2025)
by: Khan, Faizan Farooq, et al.
Published: (2025)
3DCoMPaT200: Language-Grounded Compositional Understanding of Parts and Materials of 3D Shapes
by: Ahmed, Mahmoud, et al.
Published: (2025)
by: Ahmed, Mahmoud, et al.
Published: (2025)
Vgent: Graph-based Retrieval-Reasoning-Augmented Generation For Long Video Understanding
by: Shen, Xiaoqian, et al.
Published: (2025)
by: Shen, Xiaoqian, et al.
Published: (2025)
CoT3DRef: Chain-of-Thoughts Data-Efficient 3D Visual Grounding
by: Abdelrahman, Eslam, et al.
Published: (2023)
by: Abdelrahman, Eslam, et al.
Published: (2023)
From Priors to Perception: Grounding Video-LLMs in Physical Reality
by: Zhao, Zicheng, et al.
Published: (2026)
by: Zhao, Zicheng, et al.
Published: (2026)
Localized Latent Editing for Dose-Response Modeling in Botulinum Toxin Injection Planning
by: Arnaud, Estèphe, et al.
Published: (2026)
by: Arnaud, Estèphe, et al.
Published: (2026)
Breaking Latent Prior Bias in Detectors for Generalizable AIGC Image Detection
by: Zhou, Yue, et al.
Published: (2025)
by: Zhou, Yue, et al.
Published: (2025)
SphereDrag: Spherical Geometry-Aware Panoramic Image Editing
by: Feng, Zhiao, et al.
Published: (2025)
by: Feng, Zhiao, et al.
Published: (2025)
From Static to Dynamic: a Survey of Topology-Aware Perception in Autonomous Driving
by: Chen, Yixiao, et al.
Published: (2025)
by: Chen, Yixiao, et al.
Published: (2025)
DRNet: All-in-One Image Restoration via Prior-Guided Dynamic Reparameterization
by: Li, Ao, et al.
Published: (2026)
by: Li, Ao, et al.
Published: (2026)
PhysEdit: Physically-Consistent Region-Aware Image Editing via Adaptive Spatio-Temporal Reasoning
by: Li, Guandong, et al.
Published: (2026)
by: Li, Guandong, et al.
Published: (2026)
Multimodal-Conditioned Latent Diffusion Models for Fashion Image Editing
by: Baldrati, Alberto, et al.
Published: (2024)
by: Baldrati, Alberto, et al.
Published: (2024)
PixWizard: Versatile Image-to-Image Visual Assistant with Open-Language Instructions
by: Lin, Weifeng, et al.
Published: (2024)
by: Lin, Weifeng, et al.
Published: (2024)
3D-Fixup: Advancing Photo Editing with 3D Priors
by: Cheng, Yen-Chi, et al.
Published: (2025)
by: Cheng, Yen-Chi, et al.
Published: (2025)
LIVE: Leveraging Image Manipulation Priors for Instruction-based Video Editing
by: Wang, Weicheng, et al.
Published: (2026)
by: Wang, Weicheng, et al.
Published: (2026)
LIPE: Learning Personalized Identity Prior for Non-rigid Image Editing
by: Liu, Aoyang, et al.
Published: (2024)
by: Liu, Aoyang, et al.
Published: (2024)
Lumina-mGPT: Illuminate Flexible Photorealistic Text-to-Image Generation with Multimodal Generative Pretraining
by: Liu, Dongyang, et al.
Published: (2024)
by: Liu, Dongyang, et al.
Published: (2024)
NEP: Autoregressive Image Editing via Next Editing Token Prediction
by: Wu, Huimin, et al.
Published: (2025)
by: Wu, Huimin, et al.
Published: (2025)
DailyArt: Discovering Articulation from Single Static Images via Latent Dynamics
by: Zhang, Hang, et al.
Published: (2026)
by: Zhang, Hang, et al.
Published: (2026)
FlexiEdit: Frequency-Aware Latent Refinement for Enhanced Non-Rigid Editing
by: Koo, Gwanhyeong, et al.
Published: (2024)
by: Koo, Gwanhyeong, et al.
Published: (2024)
Similar Items
-
From Reflection to Perfection: Scaling Inference-Time Optimization for Text-to-Image Diffusion Models via Reflection Tuning
by: Zhuo, Le, et al.
Published: (2025) -
ToddlerDiffusion: Interactive Structured Image Generation with Cascaded Schrödinger Bridge
by: Abdelrahman, Eslam, et al.
Published: (2023) -
Factuality Matters: When Image Generation and Editing Meet Structured Visuals
by: Zhuo, Le, et al.
Published: (2025) -
VRSBench: A Versatile Vision-Language Benchmark Dataset for Remote Sensing Image Understanding
by: Li, Xiang, et al.
Published: (2024) -
WikiAutoGen: Towards Multi-Modal Wikipedia-Style Article Generation
by: Yang, Zhongyu, et al.
Published: (2025)