RefineAnything: Multimodal Region-Specific Refinement for Perfect Local Details
Fuente:
arXiv
Saved in:
| Main Authors: | Zhou, Dewei, Li, You, Yang, Zongxin, Yang, Yi |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
MIGC++: Advanced Multi-Instance Generation Controller for Image Synthesis
by: Zhou, Dewei, et al.
Published: (2024)
by: Zhou, Dewei, et al.
Published: (2024)
DreamRenderer: Taming Multi-Instance Attribute Control in Large-Scale Text-to-Image Models
by: Zhou, Dewei, et al.
Published: (2025)
by: Zhou, Dewei, et al.
Published: (2025)
3DIS: Depth-Driven Decoupled Instance Synthesis for Text-to-Image Generation
by: Zhou, Dewei, et al.
Published: (2024)
by: Zhou, Dewei, et al.
Published: (2024)
3DIS-FLUX: simple and efficient multi-instance generation with DiT rendering
by: Zhou, Dewei, et al.
Published: (2025)
by: Zhou, Dewei, et al.
Published: (2025)
GD^2-NeRF: Generative Detail Compensation via GAN and Diffusion for One-shot Generalizable Neural Radiance Fields
by: Pan, Xiao, et al.
Published: (2024)
by: Pan, Xiao, et al.
Published: (2024)
SAMRefiner: Taming Segment Anything Model for Universal Mask Refinement
by: Lin, Yuqi, et al.
Published: (2025)
by: Lin, Yuqi, et al.
Published: (2025)
PixPerfect: Seamless Latent Diffusion Local Editing with Discriminative Pixel-Space Refinement
by: Zheng, Haitian, et al.
Published: (2025)
by: Zheng, Haitian, et al.
Published: (2025)
BideDPO: Conditional Image Generation with Simultaneous Text and Condition Alignment
by: Zhou, Dewei, et al.
Published: (2025)
by: Zhou, Dewei, et al.
Published: (2025)
Synergistic Multiscale Detail Refinement via Intrinsic Supervision for Underwater Image Enhancement
by: Zhang, Dehuan, et al.
Published: (2023)
by: Zhang, Dehuan, et al.
Published: (2023)
Zoom-Refine: Boosting High-Resolution Multimodal Understanding via Localized Zoom and Self-Refinement
by: Yu, Xuan, et al.
Published: (2025)
by: Yu, Xuan, et al.
Published: (2025)
OmniRefiner: Reinforcement-Guided Local Diffusion Refinement
by: Liu, Yaoli, et al.
Published: (2025)
by: Liu, Yaoli, et al.
Published: (2025)
From Prediction to Perfection: Introducing Refinement to Autoregressive Image Generation
by: Cheng, Cheng, et al.
Published: (2025)
by: Cheng, Cheng, et al.
Published: (2025)
MIGC: Multi-Instance Generation Controller for Text-to-Image Synthesis
by: Zhou, Dewei, et al.
Published: (2024)
by: Zhou, Dewei, et al.
Published: (2024)
Refinement via Regeneration: Enlarging Modification Space Boosts Image Refinement in Unified Multimodal Models
by: Guo, Jiayi, et al.
Published: (2026)
by: Guo, Jiayi, et al.
Published: (2026)
SIFU: Side-view Conditioned Implicit Function for Real-world Usable Clothed Human Reconstruction
by: Zhang, Zechuan, et al.
Published: (2023)
by: Zhang, Zechuan, et al.
Published: (2023)
Describe Anything: Detailed Localized Image and Video Captioning
by: Lian, Long, et al.
Published: (2025)
by: Lian, Long, et al.
Published: (2025)
AbsGS: Recovering Fine Details for 3D Gaussian Splatting
by: Ye, Zongxin, et al.
Published: (2024)
by: Ye, Zongxin, et al.
Published: (2024)
Region-Aware Text-to-Image Generation via Hard Binding and Soft Refinement
by: Chen, Zhennan, et al.
Published: (2024)
by: Chen, Zhennan, et al.
Published: (2024)
Enhancing Dataset Distillation via Label Inconsistency Elimination and Learning Pattern Refinement
by: Zhou, Chuhao, et al.
Published: (2024)
by: Zhou, Chuhao, et al.
Published: (2024)
VLM-R$^3$: Region Recognition, Reasoning, and Refinement for Enhanced Multimodal Chain-of-Thought
by: Jiang, Chaoya, et al.
Published: (2025)
by: Jiang, Chaoya, et al.
Published: (2025)
RefineStyle: Dynamic Convolution Refinement for StyleGAN
by: Xia, Siwei, et al.
Published: (2024)
by: Xia, Siwei, et al.
Published: (2024)
HeadStudio: Text to Animatable Head Avatars with 3D Gaussian Splatting
by: Zhou, Zhenglin, et al.
Published: (2024)
by: Zhou, Zhenglin, et al.
Published: (2024)
TRACE: High-Fidelity 3D Scene Editing via Tangible Reconstruction and Geometry-Aligned Contextual Video Masking
by: Hu, Jiyuan, et al.
Published: (2026)
by: Hu, Jiyuan, et al.
Published: (2026)
3D Object Manipulation in a Single Image using Generative Models
by: Zhao, Ruisi, et al.
Published: (2025)
by: Zhao, Ruisi, et al.
Published: (2025)
Are Image-to-Video Models Good Zero-Shot Image Editors?
by: Zhang, Zechuan, et al.
Published: (2025)
by: Zhang, Zechuan, et al.
Published: (2025)
FoleyDirector: Fine-Grained Temporal Steering for Video-to-Audio Generation via Structured Scripts
by: Li, You, et al.
Published: (2026)
by: Li, You, et al.
Published: (2026)
AnoRefiner: Anomaly-Aware Group-Wise Refinement for Zero-Shot Industrial Anomaly Detection
by: Huang, Dayou, et al.
Published: (2025)
by: Huang, Dayou, et al.
Published: (2025)
RS-SSM: Refining Forgotten Specifics in State Space Model for Video Semantic Segmentation
by: Zhu, Kai, et al.
Published: (2026)
by: Zhu, Kai, et al.
Published: (2026)
Temporal Action Localization with Cross Layer Task Decoupling and Refinement
by: Li, Qiang, et al.
Published: (2024)
by: Li, Qiang, et al.
Published: (2024)
Segment Anything Is Not Always Perfect: An Investigation of SAM on Different Real-world Applications
by: Ji, Wei, et al.
Published: (2023)
by: Ji, Wei, et al.
Published: (2023)
SmartRefine: A Scenario-Adaptive Refinement Framework for Efficient Motion Prediction
by: Zhou, Yang, et al.
Published: (2024)
by: Zhou, Yang, et al.
Published: (2024)
ContextGen: Contextual Layout Anchoring for Identity-Consistent Multi-Instance Generation
by: Xu, Ruihang, et al.
Published: (2025)
by: Xu, Ruihang, et al.
Published: (2025)
FIRE: A Dataset for Feedback Integration and Refinement Evaluation of Multimodal Models
by: Li, Pengxiang, et al.
Published: (2024)
by: Li, Pengxiang, et al.
Published: (2024)
DoraemonGPT: Toward Understanding Dynamic Scenes with Large Language Models (Exemplified as A Video Agent)
by: Yang, Zongxin, et al.
Published: (2024)
by: Yang, Zongxin, et al.
Published: (2024)
Enhancing Dataset Distillation via Non-Critical Region Refinement
by: Tran, Minh-Tuan, et al.
Published: (2025)
by: Tran, Minh-Tuan, et al.
Published: (2025)
Stroke3D: Lifting 2D strokes into rigged 3D model via latent diffusion models
by: Zhao, Ruisi, et al.
Published: (2026)
by: Zhao, Ruisi, et al.
Published: (2026)
In-Context Edit: Enabling Instructional Image Editing with In-Context Generation in Large Scale Diffusion Transformer
by: Zhang, Zechuan, et al.
Published: (2025)
by: Zhang, Zechuan, et al.
Published: (2025)
Multimodal OCR: Parse Anything from Documents
by: Zheng, Handong, et al.
Published: (2026)
by: Zheng, Handong, et al.
Published: (2026)
Controllable 3D Face Generation with Conditional Style Code Diffusion
by: Shen, Xiaolong, et al.
Published: (2023)
by: Shen, Xiaolong, et al.
Published: (2023)
IDPro: Flexible Interactive Video Object Segmentation by ID-queried Concurrent Propagation
by: Li, Kexin, et al.
Published: (2024)
by: Li, Kexin, et al.
Published: (2024)
Similar Items
-
MIGC++: Advanced Multi-Instance Generation Controller for Image Synthesis
by: Zhou, Dewei, et al.
Published: (2024) -
DreamRenderer: Taming Multi-Instance Attribute Control in Large-Scale Text-to-Image Models
by: Zhou, Dewei, et al.
Published: (2025) -
3DIS: Depth-Driven Decoupled Instance Synthesis for Text-to-Image Generation
by: Zhou, Dewei, et al.
Published: (2024) -
3DIS-FLUX: simple and efficient multi-instance generation with DiT rendering
by: Zhou, Dewei, et al.
Published: (2025) -
GD^2-NeRF: Generative Detail Compensation via GAN and Diffusion for One-shot Generalizable Neural Radiance Fields
by: Pan, Xiao, et al.
Published: (2024)