Step1X-Edit: A Practical Framework for General Image Editing
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Liu, Shiyu, Han, Yucheng, Xing, Peng, Yin, Fukun, Wang, Rui, Cheng, Wei, Liao, Jiaqi, Wang, Yingming, Fu, Honghao, Han, Chunrui, Li, Guopeng, Peng, Yuang, Sun, Quan, Wu, Jingwei, Cai, Yan, Ge, Zheng, Ming, Ranchen, Xia, Lei, Zeng, Xianfang, Zhu, Yibo, Jiao, Binxing, Zhang, Xiangyu, Yu, Gang, Jiang, Daxin |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
ReasonEdit: Towards Reasoning-Enhanced Image Editing Models
von: Yin, Fukun, et al.
Veröffentlicht: (2025)
von: Yin, Fukun, et al.
Veröffentlicht: (2025)
NextStep-1: Toward Autoregressive Image Generation with Continuous Tokens at Scale
von: NextStep Team, et al.
Veröffentlicht: (2025)
von: NextStep Team, et al.
Veröffentlicht: (2025)
Step-Audio-EditX Technical Report
von: Yan, Chao, et al.
Veröffentlicht: (2025)
von: Yan, Chao, et al.
Veröffentlicht: (2025)
RealRestorer: Towards Generalizable Real-World Image Restoration with Large-Scale Image Editing Models
von: Yang, Yufeng, et al.
Veröffentlicht: (2026)
von: Yang, Yufeng, et al.
Veröffentlicht: (2026)
DGAE: Diffusion-Guided Autoencoder for Efficient Latent Representation Learning
von: Liu, Dongxu, et al.
Veröffentlicht: (2025)
von: Liu, Dongxu, et al.
Veröffentlicht: (2025)
Thinking by Doing: Building Efficient World Model Reasoning in LLMs via Multi-turn Interaction
von: Shu, Bao, et al.
Veröffentlicht: (2025)
von: Shu, Bao, et al.
Veröffentlicht: (2025)
Exploring Recurrent Long-term Temporal Fusion for Multi-view 3D Perception
von: Han, Chunrui, et al.
Veröffentlicht: (2023)
von: Han, Chunrui, et al.
Veröffentlicht: (2023)
Perception-R1: Pioneering Perception Policy with Reinforcement Learning
von: Yu, En, et al.
Veröffentlicht: (2025)
von: Yu, En, et al.
Veröffentlicht: (2025)
DeepEdit: Knowledge Editing as Decoding with Constraints
von: Wang, Yiwei, et al.
Veröffentlicht: (2024)
von: Wang, Yiwei, et al.
Veröffentlicht: (2024)
SeedEdit: Align Image Re-Generation to Image Editing
von: Shi, Yichun, et al.
Veröffentlicht: (2024)
von: Shi, Yichun, et al.
Veröffentlicht: (2024)
Edit Where You Mean: Region-Aware Adapter Injection for Mask-Free Local Image Editing
von: Cai, Honghao, et al.
Veröffentlicht: (2026)
von: Cai, Honghao, et al.
Veröffentlicht: (2026)
CARE-Edit: Condition-Aware Routing of Experts for Contextual Image Editing
von: Wang, Yucheng, et al.
Veröffentlicht: (2026)
von: Wang, Yucheng, et al.
Veröffentlicht: (2026)
FlexEdit: Marrying Free-Shape Masks to VLLM for Flexible Image Editing
von: Yuan, Tianshuo, et al.
Veröffentlicht: (2024)
von: Yuan, Tianshuo, et al.
Veröffentlicht: (2024)
SE-Agent: Self-Evolution Trajectory Optimization in Multi-Step Reasoning with LLM-Based Agents
von: Lin, Jiaye, et al.
Veröffentlicht: (2025)
von: Lin, Jiaye, et al.
Veröffentlicht: (2025)
LocateEdit-Bench: A Benchmark for Instruction-Based Editing Localization
von: Wu, Shiyu, et al.
Veröffentlicht: (2026)
von: Wu, Shiyu, et al.
Veröffentlicht: (2026)
ComplexBench-Edit: Benchmarking Complex Instruction-Driven Image Editing via Compositional Dependencies
von: Wang, Chenglin, et al.
Veröffentlicht: (2025)
von: Wang, Chenglin, et al.
Veröffentlicht: (2025)
ChordEdit: One-Step Low-Energy Transport for Image Editing
von: Lu, Liangsi, et al.
Veröffentlicht: (2026)
von: Lu, Liangsi, et al.
Veröffentlicht: (2026)
DreamBench++: A Human-Aligned Benchmark for Personalized Image Generation
von: Peng, Yuang, et al.
Veröffentlicht: (2024)
von: Peng, Yuang, et al.
Veröffentlicht: (2024)
EtCon: Edit-then-Consolidate for Reliable Knowledge Editing
von: Li, Ruilin, et al.
Veröffentlicht: (2025)
von: Li, Ruilin, et al.
Veröffentlicht: (2025)
SpeechEditBench: A Bilingual Multi-Attribute Benchmark for Instruction-Guided Speech Editing
von: Zhang, Hanlin, et al.
Veröffentlicht: (2026)
von: Zhang, Hanlin, et al.
Veröffentlicht: (2026)
Rethinking Where to Edit: Task-Aware Localization for Instruction-Based Image Editing
von: He, Jingxuan, et al.
Veröffentlicht: (2026)
von: He, Jingxuan, et al.
Veröffentlicht: (2026)
Text Embeddings by Weakly-Supervised Contrastive Pre-training
von: Wang, Liang, et al.
Veröffentlicht: (2022)
von: Wang, Liang, et al.
Veröffentlicht: (2022)
General OCR Theory: Towards OCR-2.0 via a Unified End-to-end Model
von: Wei, Haoran, et al.
Veröffentlicht: (2024)
von: Wei, Haoran, et al.
Veröffentlicht: (2024)
Random Policy Valuation is Enough for LLM Reasoning with Verifiable Rewards
von: He, Haoran, et al.
Veröffentlicht: (2025)
von: He, Haoran, et al.
Veröffentlicht: (2025)
Adam: Dense Retrieval Distillation with Adaptive Dark Examples
von: Tao, Chongyang, et al.
Veröffentlicht: (2022)
von: Tao, Chongyang, et al.
Veröffentlicht: (2022)
UniEdit-Flow: Unleashing Inversion and Editing in the Era of Flow Models
von: Jiao, Guanlong, et al.
Veröffentlicht: (2025)
von: Jiao, Guanlong, et al.
Veröffentlicht: (2025)
Open Vision Reasoner: Transferring Linguistic Cognitive Behavior for Visual Reasoning
von: Wei, Yana, et al.
Veröffentlicht: (2025)
von: Wei, Yana, et al.
Veröffentlicht: (2025)
ResetEdit: Precise Text-guided Editing of Generated Image via Resettable Starting Latent
von: Wang, Hanyi, et al.
Veröffentlicht: (2026)
von: Wang, Hanyi, et al.
Veröffentlicht: (2026)
RegionE: Adaptive Region-Aware Generation for Efficient Image Editing
von: Chen, Pengtao, et al.
Veröffentlicht: (2025)
von: Chen, Pengtao, et al.
Veröffentlicht: (2025)
WithAnyone: Towards Controllable and ID Consistent Image Generation
von: Xu, Hengyuan, et al.
Veröffentlicht: (2025)
von: Xu, Hengyuan, et al.
Veröffentlicht: (2025)
Learning to Compress: Unlocking the Potential of Large Language Models for Text Representation
von: Zhang, Yeqin, et al.
Veröffentlicht: (2025)
von: Zhang, Yeqin, et al.
Veröffentlicht: (2025)
HQ-Edit: A High-Quality Dataset for Instruction-based Image Editing
von: Hui, Mude, et al.
Veröffentlicht: (2024)
von: Hui, Mude, et al.
Veröffentlicht: (2024)
OmniSVG: A Unified Scalable Vector Graphics Generation Model
von: Yang, Yiying, et al.
Veröffentlicht: (2025)
von: Yang, Yiying, et al.
Veröffentlicht: (2025)
In-Context Learning with Unpaired Clips for Instruction-based Video Editing
von: Liao, Xinyao, et al.
Veröffentlicht: (2025)
von: Liao, Xinyao, et al.
Veröffentlicht: (2025)
FreeEdit: Mask-free Reference-based Image Editing with Multi-modal Instruction
von: He, Runze, et al.
Veröffentlicht: (2024)
von: He, Runze, et al.
Veröffentlicht: (2024)
AutoEdit: Automatic Hyperparameter Tuning for Image Editing
von: Pham, Chau, et al.
Veröffentlicht: (2025)
von: Pham, Chau, et al.
Veröffentlicht: (2025)
GEBench: Benchmarking Image Generation Models as GUI Environments
von: Li, Haodong, et al.
Veröffentlicht: (2026)
von: Li, Haodong, et al.
Veröffentlicht: (2026)
DirectEdit: Step-Level Accurate Inversion for Flow-Based Image Editing
von: Yang, Desong, et al.
Veröffentlicht: (2026)
von: Yang, Desong, et al.
Veröffentlicht: (2026)
TextEditBench: Evaluating Reasoning-aware Text Editing Beyond Rendering
von: Gui, Rui, et al.
Veröffentlicht: (2025)
von: Gui, Rui, et al.
Veröffentlicht: (2025)
InteractEdit: Zero-Shot Editing of Human-Object Interactions in Images
von: Hoe, Jiun Tian, et al.
Veröffentlicht: (2025)
von: Hoe, Jiun Tian, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
ReasonEdit: Towards Reasoning-Enhanced Image Editing Models
von: Yin, Fukun, et al.
Veröffentlicht: (2025) -
NextStep-1: Toward Autoregressive Image Generation with Continuous Tokens at Scale
von: NextStep Team, et al.
Veröffentlicht: (2025) -
Step-Audio-EditX Technical Report
von: Yan, Chao, et al.
Veröffentlicht: (2025) -
RealRestorer: Towards Generalizable Real-World Image Restoration with Large-Scale Image Editing Models
von: Yang, Yufeng, et al.
Veröffentlicht: (2026) -
DGAE: Diffusion-Guided Autoencoder for Efficient Latent Representation Learning
von: Liu, Dongxu, et al.
Veröffentlicht: (2025)