ControlEdit: A MultiModal Local Clothing Image Editing Method
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Cheng, Di, Shi, YingJie, Sun, ShiXin, Zhang, JiaFu, Wang, WeiJing, Liu, Yu |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Visual-Oriented Fine-Grained Knowledge Editing for MultiModal Large Language Models
von: Zeng, Zhen, et al.
Veröffentlicht: (2024)
von: Zeng, Zhen, et al.
Veröffentlicht: (2024)
Do Text Edits Generalize to Visual Generation? Benchmarking Cross-Modal Knowledge Editing in UMMs
von: Gao, Xin, et al.
Veröffentlicht: (2026)
von: Gao, Xin, et al.
Veröffentlicht: (2026)
MMCORE: MultiModal COnnection with Representation Aligned Latent Embeddings
von: Li, Zijie, et al.
Veröffentlicht: (2026)
von: Li, Zijie, et al.
Veröffentlicht: (2026)
SeedEdit: Align Image Re-Generation to Image Editing
von: Shi, Yichun, et al.
Veröffentlicht: (2024)
von: Shi, Yichun, et al.
Veröffentlicht: (2024)
MultiModal Action Conditioned Video Generation
von: Li, Yichen, et al.
Veröffentlicht: (2025)
von: Li, Yichen, et al.
Veröffentlicht: (2025)
MultiModal Fine-tuning with Synthetic Captions
von: Enomoto, Shohei, et al.
Veröffentlicht: (2026)
von: Enomoto, Shohei, et al.
Veröffentlicht: (2026)
ParallelEdits: Efficient Multi-object Image Editing
von: Huang, Mingzhen, et al.
Veröffentlicht: (2024)
von: Huang, Mingzhen, et al.
Veröffentlicht: (2024)
DesignEdit: Multi-Layered Latent Decomposition and Fusion for Unified & Accurate Image Editing
von: Jia, Yueru, et al.
Veröffentlicht: (2024)
von: Jia, Yueru, et al.
Veröffentlicht: (2024)
M3R: Localized Rainfall Nowcasting with Meteorology-Informed MultiModal Attention
von: Panta, Sanjeev, et al.
Veröffentlicht: (2026)
von: Panta, Sanjeev, et al.
Veröffentlicht: (2026)
MMA-Diffusion: MultiModal Attack on Diffusion Models
von: Yang, Yijun, et al.
Veröffentlicht: (2023)
von: Yang, Yijun, et al.
Veröffentlicht: (2023)
CharGen: High Accurate Character-Level Visual Text Generation Model with MultiModal Encoder
von: Ma, Lichen, et al.
Veröffentlicht: (2024)
von: Ma, Lichen, et al.
Veröffentlicht: (2024)
MoEdit: On Learning Quantity Perception for Multi-object Image Editing
von: Li, Yanfeng, et al.
Veröffentlicht: (2025)
von: Li, Yanfeng, et al.
Veröffentlicht: (2025)
MTPareto: A MultiModal Targeted Pareto Framework for Fake News Detection
von: Yan, Kaiying, et al.
Veröffentlicht: (2025)
von: Yan, Kaiying, et al.
Veröffentlicht: (2025)
MDE-Edit: Masked Dual-Editing for Multi-Object Image Editing via Diffusion Models
von: Zhu, Hongyang, et al.
Veröffentlicht: (2025)
von: Zhu, Hongyang, et al.
Veröffentlicht: (2025)
Lego-Edit: A General Image Editing Framework with Model-Level Bricks and MLLM Builder
von: Jia, Qifei, et al.
Veröffentlicht: (2025)
von: Jia, Qifei, et al.
Veröffentlicht: (2025)
Dementia Insights: A Context-Based MultiModal Approach
von: Mehdoui, Sahar Sinene, et al.
Veröffentlicht: (2025)
von: Mehdoui, Sahar Sinene, et al.
Veröffentlicht: (2025)
M3: 3D-Spatial MultiModal Memory
von: Zou, Xueyan, et al.
Veröffentlicht: (2025)
von: Zou, Xueyan, et al.
Veröffentlicht: (2025)
An Empirical Study on Parameter-Efficient Fine-Tuning for MultiModal Large Language Models
von: Zhou, Xiongtao, et al.
Veröffentlicht: (2024)
von: Zhou, Xiongtao, et al.
Veröffentlicht: (2024)
SeedEdit 3.0: Fast and High-Quality Generative Image Editing
von: Wang, Peng, et al.
Veröffentlicht: (2025)
von: Wang, Peng, et al.
Veröffentlicht: (2025)
ByteEdit: Boost, Comply and Accelerate Generative Image Editing
von: Ren, Yuxi, et al.
Veröffentlicht: (2024)
von: Ren, Yuxi, et al.
Veröffentlicht: (2024)
Edit2Perceive: Image Editing Diffusion Models Are Strong Dense Perceivers
von: Shi, Yiqing, et al.
Veröffentlicht: (2025)
von: Shi, Yiqing, et al.
Veröffentlicht: (2025)
TeamPath: Building MultiModal Pathology Experts with Reasoning AI Copilots
von: Liu, Tianyu, et al.
Veröffentlicht: (2025)
von: Liu, Tianyu, et al.
Veröffentlicht: (2025)
MLlm-DR: Towards Explainable Depression Recognition with MultiModal Large Language Models
von: Zhang, Wei, et al.
Veröffentlicht: (2025)
von: Zhang, Wei, et al.
Veröffentlicht: (2025)
Can MLLMs Reason in Multimodality? EMMA: An Enhanced MultiModal ReAsoning Benchmark
von: Hao, Yunzhuo, et al.
Veröffentlicht: (2025)
von: Hao, Yunzhuo, et al.
Veröffentlicht: (2025)
MultiModal-Learning for Predicting Molecular Properties: A Framework Based on Image and Graph Structures
von: Wang, Zhuoyuan, et al.
Veröffentlicht: (2023)
von: Wang, Zhuoyuan, et al.
Veröffentlicht: (2023)
EfficientEdit: Accelerating Code Editing via Edit-Oriented Speculative Decoding
von: Wang, Peiding, et al.
Veröffentlicht: (2025)
von: Wang, Peiding, et al.
Veröffentlicht: (2025)
LocateEdit-Bench: A Benchmark for Instruction-Based Editing Localization
von: Wu, Shiyu, et al.
Veröffentlicht: (2026)
von: Wu, Shiyu, et al.
Veröffentlicht: (2026)
MM-LLMs: Recent Advances in MultiModal Large Language Models
von: Zhang, Duzhen, et al.
Veröffentlicht: (2024)
von: Zhang, Duzhen, et al.
Veröffentlicht: (2024)
HAMMR: HierArchical MultiModal React agents for generic VQA
von: Castrejon, Lluis, et al.
Veröffentlicht: (2024)
von: Castrejon, Lluis, et al.
Veröffentlicht: (2024)
ChordEdit: One-Step Low-Energy Transport for Image Editing
von: Lu, Liangsi, et al.
Veröffentlicht: (2026)
von: Lu, Liangsi, et al.
Veröffentlicht: (2026)
MegaScale-Omni: A Hyper-Scale, Workload-Resilient System for MultiModal LLM Training in Production
von: Xue, Chunyu, et al.
Veröffentlicht: (2026)
von: Xue, Chunyu, et al.
Veröffentlicht: (2026)
MM-MoralBench: A MultiModal Moral Evaluation Benchmark for Large Vision-Language Models
von: Yan, Bei, et al.
Veröffentlicht: (2024)
von: Yan, Bei, et al.
Veröffentlicht: (2024)
SpotEdit: Evaluating Visually-Guided Image Editing Methods
von: Ghazanfari, Sara, et al.
Veröffentlicht: (2025)
von: Ghazanfari, Sara, et al.
Veröffentlicht: (2025)
Polarization Engineering of the Orbital Hall Conductivity in Two-dimensional Ferroelectric Higher-Order Topological Insulator Tl$_2$S and SnS
von: Hu, YingJie, et al.
Veröffentlicht: (2026)
von: Hu, YingJie, et al.
Veröffentlicht: (2026)
Edit-Compass & EditReward-Compass: A Unified Benchmark for Image Editing and Reward Modeling
von: Bai, Xuehai, et al.
Veröffentlicht: (2026)
von: Bai, Xuehai, et al.
Veröffentlicht: (2026)
Resolving UnderEdit & OverEdit with Iterative & Neighbor-Assisted Model Editing
von: Baghel, Bhiman Kumar, et al.
Veröffentlicht: (2025)
von: Baghel, Bhiman Kumar, et al.
Veröffentlicht: (2025)
UniRef-Image-Edit: Towards Scalable and Consistent Multi-Reference Image Editing
von: Wei, Hongyang, et al.
Veröffentlicht: (2026)
von: Wei, Hongyang, et al.
Veröffentlicht: (2026)
Re-M3Dr: Rebalanced MultiModal Mean Deviation Regression
von: Yin, Haojie, et al.
Veröffentlicht: (2026)
von: Yin, Haojie, et al.
Veröffentlicht: (2026)
On the Feasibility of Using MultiModal LLMs to Execute AR Social Engineering Attacks
von: Bi, Ting, et al.
Veröffentlicht: (2025)
von: Bi, Ting, et al.
Veröffentlicht: (2025)
FUSE : Failure-aware Usage of Subagent Evidence for MultiModal Search and Recommendation
von: Vatsa, Tushar, et al.
Veröffentlicht: (2025)
von: Vatsa, Tushar, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Visual-Oriented Fine-Grained Knowledge Editing for MultiModal Large Language Models
von: Zeng, Zhen, et al.
Veröffentlicht: (2024) -
Do Text Edits Generalize to Visual Generation? Benchmarking Cross-Modal Knowledge Editing in UMMs
von: Gao, Xin, et al.
Veröffentlicht: (2026) -
MMCORE: MultiModal COnnection with Representation Aligned Latent Embeddings
von: Li, Zijie, et al.
Veröffentlicht: (2026) -
SeedEdit: Align Image Re-Generation to Image Editing
von: Shi, Yichun, et al.
Veröffentlicht: (2024) -
MultiModal Action Conditioned Video Generation
von: Li, Yichen, et al.
Veröffentlicht: (2025)