Repositioning the Subject within Image
Fuente:
arXiv
Saved in:
| Main Authors: | Wang, Yikai, Cao, Chenjie, Fan, Ke, Dong, Qiaole, Li, Yifan, Xue, Xiangyang, Fu, Yanwei |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
LeftRefill: Filling Right Canvas based on Left Reference through Generalized Text-to-Image Diffusion Model
by: Cao, Chenjie, et al.
Published: (2023)
by: Cao, Chenjie, et al.
Published: (2023)
Towards Enhanced Image Inpainting: Mitigating Unwanted Object Insertion and Preserving Color Consistency
by: Wang, Yikai, et al.
Published: (2023)
by: Wang, Yikai, et al.
Published: (2023)
Improving Neural Surface Reconstruction with Feature Priors from Multi-View Image
by: Ren, Xinlin, et al.
Published: (2024)
by: Ren, Xinlin, et al.
Published: (2024)
MemFlow: Optical Flow Estimation and Prediction with Memory
by: Dong, Qiaole, et al.
Published: (2024)
by: Dong, Qiaole, et al.
Published: (2024)
Online Dense Point Tracking with Streaming Memory
by: Dong, Qiaole, et al.
Published: (2025)
by: Dong, Qiaole, et al.
Published: (2025)
Aligned Stable Inpainting: Mitigating Unwanted Object Insertion and Preserving Color Consistency
by: Wang, Yikai, et al.
Published: (2026)
by: Wang, Yikai, et al.
Published: (2026)
Enhancing Video Inpainting with Aligned Frame Interval Guidance
by: Xie, Ming, et al.
Published: (2025)
by: Xie, Ming, et al.
Published: (2025)
MVInpainter: Learning Multi-View Consistent Inpainting to Bridge 2D and 3D Editing
by: Cao, Chenjie, et al.
Published: (2024)
by: Cao, Chenjie, et al.
Published: (2024)
MVGenMaster: Scaling Multi-View Generation from Any Image via 3D Priors Enhanced Diffusion Model
by: Cao, Chenjie, et al.
Published: (2024)
by: Cao, Chenjie, et al.
Published: (2024)
The Pictorial Cortex: Zero-Shot Cross-Subject fMRI-to-Image Reconstruction via Compositional Latent Modeling
by: Huo, Jingyang, et al.
Published: (2026)
by: Huo, Jingyang, et al.
Published: (2026)
Uni3C: Unifying Precisely 3D-Enhanced Camera and Human Motion Controls for Video Generation
by: Cao, Chenjie, et al.
Published: (2025)
by: Cao, Chenjie, et al.
Published: (2025)
AnyRefill: A Unified, Data-Efficient Framework for Left-Prompt-Guided Vision Tasks
by: Xie, Ming, et al.
Published: (2025)
by: Xie, Ming, et al.
Published: (2025)
MVSFormer++: Revealing the Devil in Transformer's Details for Multi-View Stereo
by: Cao, Chenjie, et al.
Published: (2024)
by: Cao, Chenjie, et al.
Published: (2024)
EgoSound: Benchmarking Sound Understanding in Egocentric Videos
by: Zhu, Bingwen, et al.
Published: (2026)
by: Zhu, Bingwen, et al.
Published: (2026)
Image-Text-Image Knowledge Transfer for Lifelong Person Re-Identification with Hybrid Clothing States
by: Wang, Qizao, et al.
Published: (2024)
by: Wang, Qizao, et al.
Published: (2024)
ReasonGrounder: LVLM-Guided Hierarchical Feature Splatting for Open-Vocabulary 3D Visual Grounding and Reasoning
by: Liu, Zhenyang, et al.
Published: (2025)
by: Liu, Zhenyang, et al.
Published: (2025)
Exploring Fine-Grained Representation and Recomposition for Cloth-Changing Person Re-Identification
by: Wang, Qizao, et al.
Published: (2023)
by: Wang, Qizao, et al.
Published: (2023)
Improving Multi-Subject Consistency in Open-Domain Image Generation with Isolation and Reposition Attention
by: He, Huiguo, et al.
Published: (2024)
by: He, Huiguo, et al.
Published: (2024)
PPMStereo: Pick-and-Play Memory Construction for Consistent Dynamic Stereo Matching
by: Wang, Yun, et al.
Published: (2025)
by: Wang, Yun, et al.
Published: (2025)
Content and Salient Semantics Collaboration for Cloth-Changing Person Re-Identification
by: Wang, Qizao, et al.
Published: (2024)
by: Wang, Qizao, et al.
Published: (2024)
CAP-Net: A Unified Network for 6D Pose and Size Estimation of Categorical Articulated Parts from a Single RGB-D Image
by: Huang, Jingshun, et al.
Published: (2025)
by: Huang, Jingshun, et al.
Published: (2025)
3D Skew-Normal Splatting
by: Wu, Xiangru, et al.
Published: (2026)
by: Wu, Xiangru, et al.
Published: (2026)
Beyond 'Templates': Category-Agnostic Object Pose, Size, and Shape Estimation from a Single View
by: Zhang, Jinyu, et al.
Published: (2025)
by: Zhang, Jinyu, et al.
Published: (2025)
DecoFuse: Decomposing and Fusing the "What", "Where", and "How" for Brain-Inspired fMRI-to-Video Decoding
by: Li, Chong, et al.
Published: (2025)
by: Li, Chong, et al.
Published: (2025)
OCRA: Object-Centric Learning with 3D and Tactile Priors for Human-to-Robot Action Transfer
by: Wang, Kuanning, et al.
Published: (2026)
by: Wang, Kuanning, et al.
Published: (2026)
Towards Global Optimal Visual In-Context Learning Prompt Selection
by: Xu, Chengming, et al.
Published: (2024)
by: Xu, Chengming, et al.
Published: (2024)
3D StreetUnveiler with Semantic-aware 2DGS -- a simple baseline
by: Xu, Jingwei, et al.
Published: (2024)
by: Xu, Jingwei, et al.
Published: (2024)
RAG-6DPose: Retrieval-Augmented 6D Pose Estimation via Leveraging CAD as Knowledge Base
by: Wang, Kuanning, et al.
Published: (2025)
by: Wang, Kuanning, et al.
Published: (2025)
NeuroPictor: Refining fMRI-to-Image Reconstruction via Multi-individual Pretraining and Multi-level Modulation
by: Huo, Jingyang, et al.
Published: (2024)
by: Huo, Jingyang, et al.
Published: (2024)
Unified Lexical Representation for Interpretable Visual-Language Alignment
by: Li, Yifan, et al.
Published: (2024)
by: Li, Yifan, et al.
Published: (2024)
You Only Estimate Once: Unified, One-stage, Real-Time Category-level Articulated Object 6D Pose Estimation for Robotic Grasping
by: Huang, Jingshun, et al.
Published: (2025)
by: Huang, Jingshun, et al.
Published: (2025)
SCOOP'D: Learning Mixed-Liquid-Solid Scooping via Sim2Real Generative Policy
by: Wang, Kuanning, et al.
Published: (2025)
by: Wang, Kuanning, et al.
Published: (2025)
Sub-Image Recapture for Multi-View 3D Reconstruction
by: Wang, Yanwei
Published: (2025)
by: Wang, Yanwei
Published: (2025)
VCD-Texture: Variance Alignment based 3D-2D Co-Denoising for Text-Guided Texturing
by: Liu, Shang, et al.
Published: (2024)
by: Liu, Shang, et al.
Published: (2024)
EAFormer: Scene Text Segmentation with Edge-Aware Transformers
by: Yu, Haiyang, et al.
Published: (2024)
by: Yu, Haiyang, et al.
Published: (2024)
A Neural Representation Framework with LLM-Driven Spatial Reasoning for Open-Vocabulary 3D Visual Grounding
by: Liu, Zhenyang, et al.
Published: (2025)
by: Liu, Zhenyang, et al.
Published: (2025)
CoLa: Chinese Character Decomposition with Compositional Latent Components
by: Shi, Fan, et al.
Published: (2025)
by: Shi, Fan, et al.
Published: (2025)
When Large Vision-Language Models Meet Person Re-Identification
by: Wang, Qizao, et al.
Published: (2024)
by: Wang, Qizao, et al.
Published: (2024)
TP-MDDN: Task-Preferenced Multi-Demand-Driven Navigation with Autonomous Decision-Making
by: Li, Shanshan, et al.
Published: (2025)
by: Li, Shanshan, et al.
Published: (2025)
Spatial Transport Optimization by Repositioning Attention Map for Training-Free Text-to-Image Synthesis
by: Han, Woojung, et al.
Published: (2025)
by: Han, Woojung, et al.
Published: (2025)
Similar Items
-
LeftRefill: Filling Right Canvas based on Left Reference through Generalized Text-to-Image Diffusion Model
by: Cao, Chenjie, et al.
Published: (2023) -
Towards Enhanced Image Inpainting: Mitigating Unwanted Object Insertion and Preserving Color Consistency
by: Wang, Yikai, et al.
Published: (2023) -
Improving Neural Surface Reconstruction with Feature Priors from Multi-View Image
by: Ren, Xinlin, et al.
Published: (2024) -
MemFlow: Optical Flow Estimation and Prediction with Memory
by: Dong, Qiaole, et al.
Published: (2024) -
Online Dense Point Tracking with Streaming Memory
by: Dong, Qiaole, et al.
Published: (2025)