Text-Driven Image Editing via Learnable Regions
Fuente:
arXiv
Saved in:
| Main Authors: | Lin, Yuanze, Chen, Yi-Wen, Tsai, Yi-Hsuan, Jiang, Lu, Yang, Ming-Hsuan |
|---|---|
| Format: | Preprint |
| Published: |
2023
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
IllumiCraft: Unified Geometry and Illumination Diffusion for Controllable Video Generation
by: Lin, Yuanze, et al.
Published: (2025)
by: Lin, Yuanze, et al.
Published: (2025)
Chat-Edit-3D: Interactive 3D Scene Editing via Text Prompts
by: Fang, Shuangkang, et al.
Published: (2024)
by: Fang, Shuangkang, et al.
Published: (2024)
Self-training Room Layout Estimation via Geometry-aware Ray-casting
by: Solarte, Bolivar, et al.
Published: (2024)
by: Solarte, Bolivar, et al.
Published: (2024)
Delve into Visual Contrastive Decoding for Hallucination Mitigation of Large Vision-Language Models
by: Lee, Yi-Lun, et al.
Published: (2024)
by: Lee, Yi-Lun, et al.
Published: (2024)
Intra-Cluster Mixup: An Effective Data Augmentation Technique for Complementary-Label Learning
by: Mai, Tan-Ha, et al.
Published: (2025)
by: Mai, Tan-Ha, et al.
Published: (2025)
MRFD: Multi-Region Fusion Decoding with Self-Consistency for Mitigating Hallucinations in LVLMs
by: Ge, Haonan, et al.
Published: (2025)
by: Ge, Haonan, et al.
Published: (2025)
libcll: an Extendable Python Toolkit for Complementary-Label Learning
by: Ye, Nai-Xuan, et al.
Published: (2024)
by: Ye, Nai-Xuan, et al.
Published: (2024)
Self-Attention with State-Object Weighted Combination for Compositional Zero Shot Learning
by: Chang, Cheng-Hong, et al.
Published: (2025)
by: Chang, Cheng-Hong, et al.
Published: (2025)
Weakly Supervised 3D Object Detection via Multi-Level Visual Guidance
by: Huang, Kuan-Chih, et al.
Published: (2023)
by: Huang, Kuan-Chih, et al.
Published: (2023)
SuperEdit: Rectifying and Facilitating Supervision for Instruction-Based Image Editing
by: Li, Ming, et al.
Published: (2025)
by: Li, Ming, et al.
Published: (2025)
DreamPolisher: Towards High-Quality Text-to-3D Generation via Geometric Diffusion
by: Lin, Yuanze, et al.
Published: (2024)
by: Lin, Yuanze, et al.
Published: (2024)
Post-Disaster Affected Area Segmentation with a Vision Transformer (ViT)-based EVAP Model using Sentinel-2 and Formosat-5 Imagery
by: Chu, Yi-Shan, et al.
Published: (2025)
by: Chu, Yi-Shan, et al.
Published: (2025)
Who Brings the Frisbee: Probing Hidden Hallucination Factors in Large Vision-Language Model via Causality Analysis
by: Huang, Po-Hsuan, et al.
Published: (2024)
by: Huang, Po-Hsuan, et al.
Published: (2024)
Editing Massive Concepts in Text-to-Image Diffusion Models
by: Xiong, Tianwei, et al.
Published: (2024)
by: Xiong, Tianwei, et al.
Published: (2024)
Beyond Words: Multimodal LLM Knows When to Speak
by: Liao, Zikai, et al.
Published: (2025)
by: Liao, Zikai, et al.
Published: (2025)
Specify and Edit: Overcoming Ambiguity in Text-Based Image Editing
by: Iakovleva, Ekaterina, et al.
Published: (2024)
by: Iakovleva, Ekaterina, et al.
Published: (2024)
Harmonizing Generalization and Specialization: Uncertainty-Informed Collaborative Learning for Semi-supervised Medical Image Segmentation
by: Lu, Wenjing, et al.
Published: (2025)
by: Lu, Wenjing, et al.
Published: (2025)
AdaIR: Exploiting Underlying Similarities of Image Restoration Tasks with Adapters
by: Chen, Hao-Wei, et al.
Published: (2024)
by: Chen, Hao-Wei, et al.
Published: (2024)
MedEBench: Diagnosing Reliability in Text-Guided Medical Image Editing
by: Liu, Minghao, et al.
Published: (2025)
by: Liu, Minghao, et al.
Published: (2025)
CLIPDrag: Combining Text-based and Drag-based Instructions for Image Editing
by: Jiang, Ziqi, et al.
Published: (2024)
by: Jiang, Ziqi, et al.
Published: (2024)
Contrastive Denoising Score for Text-guided Latent Diffusion Image Editing
by: Nam, Hyelin, et al.
Published: (2023)
by: Nam, Hyelin, et al.
Published: (2023)
Text Speaks Louder than Vision: ASCII Art Reveals Textual Biases in Vision-Language Models
by: Wang, Zhaochen, et al.
Published: (2025)
by: Wang, Zhaochen, et al.
Published: (2025)
Pathways on the Image Manifold: Image Editing via Video Generation
by: Rotstein, Noam, et al.
Published: (2024)
by: Rotstein, Noam, et al.
Published: (2024)
SLA2: Sparse-Linear Attention with Learnable Routing and QAT
by: Zhang, Jintao, et al.
Published: (2026)
by: Zhang, Jintao, et al.
Published: (2026)
MCIE: Multimodal LLM-Driven Complex Instruction Image Editing with Spatial Guidance
by: Bai, Xuehai, et al.
Published: (2026)
by: Bai, Xuehai, et al.
Published: (2026)
Zero-Residual Concept Erasure via Progressive Alignment in Text-to-Image Model
by: Chen, Hongxu, et al.
Published: (2025)
by: Chen, Hongxu, et al.
Published: (2025)
TG-LLaVA: Text Guided LLaVA via Learnable Latent Embeddings
by: Yan, Dawei, et al.
Published: (2024)
by: Yan, Dawei, et al.
Published: (2024)
DIVE: Taming DINO for Subject-Driven Video Editing
by: Huang, Yi, et al.
Published: (2024)
by: Huang, Yi, et al.
Published: (2024)
Gaga: Group Any Gaussians via 3D-aware Memory Bank
by: Lyu, Weijie, et al.
Published: (2024)
by: Lyu, Weijie, et al.
Published: (2024)
Feedforward 3D Editing via Text-Steerable Image-to-3D
by: Ma, Ziqi, et al.
Published: (2025)
by: Ma, Ziqi, et al.
Published: (2025)
Toward Real-world BEV Perception: Depth Uncertainty Estimation via Gaussian Splatting
by: Lu, Shu-Wei, et al.
Published: (2025)
by: Lu, Shu-Wei, et al.
Published: (2025)
KnowVal: A Knowledge-Augmented and Value-Guided Autonomous Driving System
by: Xia, Zhongyu, et al.
Published: (2025)
by: Xia, Zhongyu, et al.
Published: (2025)
Action-slot: Visual Action-centric Representations for Multi-label Atomic Activity Recognition in Traffic Scenes
by: Kung, Chi-Hsi, et al.
Published: (2023)
by: Kung, Chi-Hsi, et al.
Published: (2023)
Ranking-aware adapter for text-driven image ordering with CLIP
by: Yu, Wei-Hsiang, et al.
Published: (2024)
by: Yu, Wei-Hsiang, et al.
Published: (2024)
LatentEditor: Text Driven Local Editing of 3D Scenes
by: Khalid, Umar, et al.
Published: (2023)
by: Khalid, Umar, et al.
Published: (2023)
Uniform Attention Maps: Boosting Image Fidelity in Reconstruction and Editing
by: Mo, Wenyi, et al.
Published: (2024)
by: Mo, Wenyi, et al.
Published: (2024)
Team NYCU at Defactify4: Robust Detection and Source Identification of AI-Generated Images Using CNN and CLIP-Based Models
by: Yang, Tsan-Tsung, et al.
Published: (2025)
by: Yang, Tsan-Tsung, et al.
Published: (2025)
RegionE: Adaptive Region-Aware Generation for Efficient Image Editing
by: Chen, Pengtao, et al.
Published: (2025)
by: Chen, Pengtao, et al.
Published: (2025)
Disentangling Regional Primitives for Image Generation
by: Chen, Zhengting, et al.
Published: (2024)
by: Chen, Zhengting, et al.
Published: (2024)
IE-Bench: Advancing the Measurement of Text-Driven Image Editing for Human Perception Alignment
by: Sun, Shangkun, et al.
Published: (2025)
by: Sun, Shangkun, et al.
Published: (2025)
Similar Items
-
IllumiCraft: Unified Geometry and Illumination Diffusion for Controllable Video Generation
by: Lin, Yuanze, et al.
Published: (2025) -
Chat-Edit-3D: Interactive 3D Scene Editing via Text Prompts
by: Fang, Shuangkang, et al.
Published: (2024) -
Self-training Room Layout Estimation via Geometry-aware Ray-casting
by: Solarte, Bolivar, et al.
Published: (2024) -
Delve into Visual Contrastive Decoding for Hallucination Mitigation of Large Vision-Language Models
by: Lee, Yi-Lun, et al.
Published: (2024) -
Intra-Cluster Mixup: An Effective Data Augmentation Technique for Complementary-Label Learning
by: Mai, Tan-Ha, et al.
Published: (2025)