SpatialEdit: Benchmarking Fine-Grained Image Spatial Editing
Fuente:
arXiv
Saved in:
| Main Authors: | Xiao, Yicheng, Zhang, Wenhu, Song, Lin, Chen, Yukang, Li, Wenbo, Jiang, Nan, Ren, Tianhe, Lin, Haokun, Huang, Wei, Huang, Haoyang, Li, Xiu, Duan, Nan, Qi, Xiaojuan |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
OpenSpatial: A Principled Data Engine for Empowering Spatial Intelligence
by: Liu, Jianhui, et al.
Published: (2026)
by: Liu, Jianhui, et al.
Published: (2026)
MindOmni: Unleashing Reasoning Generation in Vision Language Models with RGPO
by: Xiao, Yicheng, et al.
Published: (2025)
by: Xiao, Yicheng, et al.
Published: (2025)
JoyAI-Image: Awaking Spatial Intelligence in Unified Multimodal Understanding and Generation
by: Song, Lin, et al.
Published: (2026)
by: Song, Lin, et al.
Published: (2026)
Thinking with Novel Views: A Systematic Analysis of Generative-Augmented Spatial Intelligence
by: Zhang, Yanbing, et al.
Published: (2026)
by: Zhang, Yanbing, et al.
Published: (2026)
CogniEdit: Dense Gradient Flow Optimization for Fine-Grained Image Editing
by: Li, Yan, et al.
Published: (2025)
by: Li, Yan, et al.
Published: (2025)
FineEdit: Fine-Grained Image Edit with Bounding Box Guidance
by: Xu, Haohang, et al.
Published: (2026)
by: Xu, Haohang, et al.
Published: (2026)
PanguMotion: Continuous Driving Motion Forecasting with Pangu Transformers
by: Ren, Quanhao, et al.
Published: (2026)
by: Ren, Quanhao, et al.
Published: (2026)
Embodied3DBench: Benchmarking Low-Level Embodied Spatial Intelligence of Vision Language Models
by: Zhang, Jiyao, et al.
Published: (2026)
by: Zhang, Jiyao, et al.
Published: (2026)
See, Remember, Explore: A Benchmark and Baselines for Streaming Spatial Reasoning
by: Wei, Yuxi, et al.
Published: (2026)
by: Wei, Yuxi, et al.
Published: (2026)
UltraEdit: Instruction-based Fine-Grained Image Editing at Scale
by: Zhao, Haozhe, et al.
Published: (2024)
by: Zhao, Haozhe, et al.
Published: (2024)
ImgEdit: A Unified Image Editing Dataset and Benchmark
by: Ye, Yang, et al.
Published: (2025)
by: Ye, Yang, et al.
Published: (2025)
Optimality Deviation using the Koopman Operator
by: Lin, Yicheng, et al.
Published: (2025)
by: Lin, Yicheng, et al.
Published: (2025)
COVE: Unleashing the Diffusion Feature Correspondence for Consistent Video Editing
by: Wang, Jiangshan, et al.
Published: (2024)
by: Wang, Jiangshan, et al.
Published: (2024)
OmniEdit: Building Image Editing Generalist Models Through Specialist Supervision
by: Wei, Cong, et al.
Published: (2024)
by: Wei, Cong, et al.
Published: (2024)
Optimality Robustness in Koopman-Based Control
by: Lin, Yicheng, et al.
Published: (2026)
by: Lin, Yicheng, et al.
Published: (2026)
DBellQuant: Breaking the Bell with Double-Bell Transformation for LLMs Post Training Binarization
by: Ye, Zijian, et al.
Published: (2025)
by: Ye, Zijian, et al.
Published: (2025)
SliderEdit: Continuous Image Editing with Fine-Grained Instruction Control
by: Zarei, Arman, et al.
Published: (2025)
by: Zarei, Arman, et al.
Published: (2025)
FineMotion: A Dataset and Benchmark with both Spatial and Temporal Annotation for Fine-grained Motion Generation and Editing
by: Wu, Bizhu, et al.
Published: (2025)
by: Wu, Bizhu, et al.
Published: (2025)
DisProtEdit: Exploring Disentangled Representations for Multi-Attribute Protein Editing
by: Ku, Max, et al.
Published: (2025)
by: Ku, Max, et al.
Published: (2025)
EditFlow: Benchmarking and Optimizing Code Edit Recommendation Systems via Reconstruction of Developer Flows
by: Liu, Chenyan, et al.
Published: (2026)
by: Liu, Chenyan, et al.
Published: (2026)
MIKE: A New Benchmark for Fine-grained Multimodal Entity Knowledge Editing
by: Li, Jiaqi, et al.
Published: (2024)
by: Li, Jiaqi, et al.
Published: (2024)
SmartFreeEdit: Mask-Free Spatial-Aware Image Editing with Complex Instruction Understanding
by: Sun, Qianqian, et al.
Published: (2025)
by: Sun, Qianqian, et al.
Published: (2025)
SEDualVLN: A Spatially-Enhanced Dual-System for Vision-Language Navigation
by: Huang, Jingzhi, et al.
Published: (2026)
by: Huang, Jingzhi, et al.
Published: (2026)
LMM4Edit: Benchmarking and Evaluating Multimodal Image Editing with LMMs
by: Xu, Zitong, et al.
Published: (2025)
by: Xu, Zitong, et al.
Published: (2025)
Learning to Reason in 4D: Dynamic Spatial Understanding for Vision Language Models
by: Zhou, Shengchao, et al.
Published: (2025)
by: Zhou, Shengchao, et al.
Published: (2025)
PartialEdit: Identifying Partial Deepfakes in the Era of Neural Speech Editing
by: Zhang, You, et al.
Published: (2025)
by: Zhang, You, et al.
Published: (2025)
2D Instance Editing in 3D Space
by: Xie, Yuhuan, et al.
Published: (2025)
by: Xie, Yuhuan, et al.
Published: (2025)
Collu-Bench: A Benchmark for Predicting Language Model Hallucinations in Code
by: Jiang, Nan, et al.
Published: (2024)
by: Jiang, Nan, et al.
Published: (2024)
Bridging the Editing Gap in LLMs: FineEdit for Precise and Targeted Text Modifications
by: Zeng, Yiming, et al.
Published: (2025)
by: Zeng, Yiming, et al.
Published: (2025)
S$^2$Edit: Text-Guided Image Editing with Precise Semantic and Spatial Control
by: Liu, Xudong, et al.
Published: (2025)
by: Liu, Xudong, et al.
Published: (2025)
VidEdit: Zero-Shot and Spatially Aware Text-Driven Video Editing
by: Couairon, Paul, et al.
Published: (2023)
by: Couairon, Paul, et al.
Published: (2023)
FireSentry: A Multi-Modal Spatio-temporal Benchmark Dataset for Fine-Grained Wildfire Spread Forecasting
by: Zhou, Nan, et al.
Published: (2025)
by: Zhou, Nan, et al.
Published: (2025)
Spatial organization of phosphoinositide signaling
by: Siyu Lai, et al.
Published: (2025)
by: Siyu Lai, et al.
Published: (2025)
PartEdit: Fine-Grained Image Editing using Pre-Trained Diffusion Models
by: Cvejic, Aleksandar, et al.
Published: (2025)
by: Cvejic, Aleksandar, et al.
Published: (2025)
Spatial-DISE: A Unified Benchmark for Evaluating Spatial Reasoning in Vision-Language Models
by: Huang, Xinmiao, et al.
Published: (2025)
by: Huang, Xinmiao, et al.
Published: (2025)
SpatialReward: Verifiable Spatial Reward Modeling for Fine-Grained Spatial Consistency in Text-to-Image Generation
by: Zhou, Sashuai, et al.
Published: (2026)
by: Zhou, Sashuai, et al.
Published: (2026)
TAPTRv3: Spatial and Temporal Context Foster Robust Tracking of Any Point in Long Video
by: Qu, Jinyuan, et al.
Published: (2024)
by: Qu, Jinyuan, et al.
Published: (2024)
EfficientEdit: Accelerating Code Editing via Edit-Oriented Speculative Decoding
by: Wang, Peiding, et al.
Published: (2025)
by: Wang, Peiding, et al.
Published: (2025)
WorldEdit: Towards Open-World Image Editing with a Knowledge-Informed Benchmark
by: Lin, Wang, et al.
Published: (2026)
by: Lin, Wang, et al.
Published: (2026)
Facile Fabrication of Robust Supraparticles for Spatially Orthogonal Cascade Catalysis
by: Nan Xue, et al.
Published: (2025)
by: Nan Xue, et al.
Published: (2025)
Similar Items
-
OpenSpatial: A Principled Data Engine for Empowering Spatial Intelligence
by: Liu, Jianhui, et al.
Published: (2026) -
MindOmni: Unleashing Reasoning Generation in Vision Language Models with RGPO
by: Xiao, Yicheng, et al.
Published: (2025) -
JoyAI-Image: Awaking Spatial Intelligence in Unified Multimodal Understanding and Generation
by: Song, Lin, et al.
Published: (2026) -
Thinking with Novel Views: A Systematic Analysis of Generative-Augmented Spatial Intelligence
by: Zhang, Yanbing, et al.
Published: (2026) -
CogniEdit: Dense Gradient Flow Optimization for Fine-Grained Image Editing
by: Li, Yan, et al.
Published: (2025)