Unveil Inversion and Invariance in Flow Transformer for Versatile Image Editing
Fuente:
arXiv
Saved in:
| Main Authors: | Xu, Pengcheng, Jiang, Boyuan, Hu, Xiaobin, Luo, Donghao, He, Qingdong, Zhang, Jiangning, Wang, Chengjie, Wu, Yunsheng, Ling, Charles, Wang, Boyu |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Towards Generalized Multi-Image Editing for Unified Multimodal Models
by: Xu, Pengcheng, et al.
Published: (2026)
by: Xu, Pengcheng, et al.
Published: (2026)
FitDiT: Advancing the Authentic Garment Details for High-fidelity Virtual Try-on
by: Jiang, Boyuan, et al.
Published: (2024)
by: Jiang, Boyuan, et al.
Published: (2024)
DynamicControl: Adaptive Condition Selection for Improved Text-to-Image Generation
by: He, Qingdong, et al.
Published: (2024)
by: He, Qingdong, et al.
Published: (2024)
Reasoning to Edit: Hypothetical Instruction-Based Image Editing with Visual Reasoning
by: He, Qingdong, et al.
Published: (2025)
by: He, Qingdong, et al.
Published: (2025)
NoiseBoost: Alleviating Hallucination with Noise Perturbation for Multimodal Large Language Models
by: Wu, Kai, et al.
Published: (2024)
by: Wu, Kai, et al.
Published: (2024)
FFP-300K: Scaling First-Frame Propagation for Generalizable Video Editing
by: Huang, Xijie, et al.
Published: (2026)
by: Huang, Xijie, et al.
Published: (2026)
IVEBench: Modern Benchmark Suite for Instruction-Guided Video Editing Assessment
by: Chen, Yinan, et al.
Published: (2025)
by: Chen, Yinan, et al.
Published: (2025)
Transform Trained Transformer: Accelerating Naive 4K Video Generation Over 10$\times$
by: Zhang, Jiangning, et al.
Published: (2025)
by: Zhang, Jiangning, et al.
Published: (2025)
PointSeg: A Training-Free Paradigm for 3D Scene Segmentation via Foundation Models
by: He, Qingdong, et al.
Published: (2024)
by: He, Qingdong, et al.
Published: (2024)
ArtWeaver: Advanced Dynamic Style Integration via Diffusion Model
by: Xu, Chengming, et al.
Published: (2024)
by: Xu, Chengming, et al.
Published: (2024)
Towards One-step Causal Video Generation via Adversarial Self-Distillation
by: Yang, Yongqi, et al.
Published: (2025)
by: Yang, Yongqi, et al.
Published: (2025)
The devil is in the details: Enhancing Video Virtual Try-On via Keyframe-Driven Details Injection
by: He, Qingdong, et al.
Published: (2025)
by: He, Qingdong, et al.
Published: (2025)
VTON-HandFit: Virtual Try-on for Arbitrary Hand Pose Guided by Hand Priors Embedding
by: Liang, Yujie, et al.
Published: (2024)
by: Liang, Yujie, et al.
Published: (2024)
UniM-OV3D: Uni-Modality Open-Vocabulary 3D Scene Understanding with Fine-Grained Feature Representation
by: He, Qingdong, et al.
Published: (2024)
by: He, Qingdong, et al.
Published: (2024)
Sonic: Shifting Focus to Global Audio Perception in Portrait Animation
by: Ji, Xiaozhong, et al.
Published: (2024)
by: Ji, Xiaozhong, et al.
Published: (2024)
VTBench: Comprehensive Benchmark Suite Towards Real-World Virtual Try-on Models
by: Xiaobin, Hu, et al.
Published: (2025)
by: Xiaobin, Hu, et al.
Published: (2025)
CrossVTON: Mimicking the Logic Reasoning on Cross-category Virtual Try-on guided by Tri-zone Priors
by: Luo, Donghao, et al.
Published: (2025)
by: Luo, Donghao, et al.
Published: (2025)
MobileMamba: Lightweight Multi-Receptive Visual Mamba Network
by: He, Haoyang, et al.
Published: (2024)
by: He, Haoyang, et al.
Published: (2024)
DiffuMatting: Synthesizing Arbitrary Objects with Matting-level Annotation
by: Hu, Xiaobin, et al.
Published: (2024)
by: Hu, Xiaobin, et al.
Published: (2024)
PointRWKV: Efficient RWKV-Like Model for Hierarchical Point Cloud Learning
by: He, Qingdong, et al.
Published: (2024)
by: He, Qingdong, et al.
Published: (2024)
Human-MME: A Holistic Evaluation Benchmark for Human-Centric Multimodal Large Language Models
by: Liu, Yuansen, et al.
Published: (2025)
by: Liu, Yuansen, et al.
Published: (2025)
Class Overwhelms: Mutual Conditional Blended-Target Domain Adaptation
by: Xu, Pengcheng, et al.
Published: (2023)
by: Xu, Pengcheng, et al.
Published: (2023)
Exploring Real&Synthetic Dataset and Linear Attention in Image Restoration
by: Du, Yuzhen, et al.
Published: (2024)
by: Du, Yuzhen, et al.
Published: (2024)
FireFlow: Fast Inversion of Rectified Flow for Image Semantic Editing
by: Deng, Yingying, et al.
Published: (2024)
by: Deng, Yingying, et al.
Published: (2024)
Textualize Visual Prompt for Image Editing via Diffusion Bridge
by: Xu, Pengcheng, et al.
Published: (2025)
by: Xu, Pengcheng, et al.
Published: (2025)
Dual-Schedule Inversion: Training- and Tuning-Free Inversion for Real Image Editing
by: Huang, Jiancheng, et al.
Published: (2024)
by: Huang, Jiancheng, et al.
Published: (2024)
Boosting Reasoning in Large Multimodal Models via Activation Replay
by: Xing, Yun, et al.
Published: (2025)
by: Xing, Yun, et al.
Published: (2025)
CustAny: Customizing Anything from A Single Example
by: Kong, Lingjie, et al.
Published: (2024)
by: Kong, Lingjie, et al.
Published: (2024)
HumanVideo-MME: Benchmarking MLLMs for Human-Centric Video Understanding
by: Cai, Yuxuan, et al.
Published: (2025)
by: Cai, Yuxuan, et al.
Published: (2025)
Image Inversion: A Survey from GANs to Diffusion and Beyond
by: Chen, Yinan, et al.
Published: (2025)
by: Chen, Yinan, et al.
Published: (2025)
OracleFusion: Assisting the Decipherment of Oracle Bone Script with Structurally Constrained Semantic Typography
by: Li, Caoshuo, et al.
Published: (2025)
by: Li, Caoshuo, et al.
Published: (2025)
SplitFlow: Flow Decomposition for Inversion-Free Text-to-Image Editing
by: Yoon, Sung-Hoon, et al.
Published: (2025)
by: Yoon, Sung-Hoon, et al.
Published: (2025)
VIVID-10M: A Dataset and Baseline for Versatile and Interactive Video Local Editing
by: Hu, Jiahao, et al.
Published: (2024)
by: Hu, Jiahao, et al.
Published: (2024)
MARRS: Masked Autoregressive Unit-based Reaction Synthesis
by: Wang, Yabiao, et al.
Published: (2025)
by: Wang, Yabiao, et al.
Published: (2025)
SteerFlow: Steering Rectified Flows for Faithful Inversion-Based Image Editing
by: Dao, Thinh, et al.
Published: (2026)
by: Dao, Thinh, et al.
Published: (2026)
IRPO: Boosting Image Restoration via Post-training GRPO
by: Xu, Haoxuan, et al.
Published: (2025)
by: Xu, Haoxuan, et al.
Published: (2025)
M3DM-NR: RGB-3D Noisy-Resistant Industrial Anomaly Detection via Multimodal Denoising
by: Wang, Chengjie, et al.
Published: (2024)
by: Wang, Chengjie, et al.
Published: (2024)
CLEAR: Context-Aware Learning with End-to-End Mask-Free Inference for Adaptive Video Subtitle Removal
by: He, Qingdong, et al.
Published: (2026)
by: He, Qingdong, et al.
Published: (2026)
TokenAR: Multiple Subject Generation via Autoregressive Token-level enhancement
by: Sun, Haiyue, et al.
Published: (2025)
by: Sun, Haiyue, et al.
Published: (2025)
Open-Vocabulary SAM3D: Towards Training-free Open-Vocabulary 3D Scene Understanding
by: Tai, Hanchen, et al.
Published: (2024)
by: Tai, Hanchen, et al.
Published: (2024)
Similar Items
-
Towards Generalized Multi-Image Editing for Unified Multimodal Models
by: Xu, Pengcheng, et al.
Published: (2026) -
FitDiT: Advancing the Authentic Garment Details for High-fidelity Virtual Try-on
by: Jiang, Boyuan, et al.
Published: (2024) -
DynamicControl: Adaptive Condition Selection for Improved Text-to-Image Generation
by: He, Qingdong, et al.
Published: (2024) -
Reasoning to Edit: Hypothetical Instruction-Based Image Editing with Visual Reasoning
by: He, Qingdong, et al.
Published: (2025) -
NoiseBoost: Alleviating Hallucination with Noise Perturbation for Multimodal Large Language Models
by: Wu, Kai, et al.
Published: (2024)