Saved in:
| Main Authors: | Ryu, Suho, Kim, Kihyun, Baek, Eugene, Shin, Dongsoo, Lee, Joonseok |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2505.00502 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Towards Motion-aware Referring Image Segmentation
by: Kim, Chaeyun, et al.
Published: (2026)
by: Kim, Chaeyun, et al.
Published: (2026)
Geometry-Aware Image Flow Matching
by: Lee, Junho, et al.
Published: (2026)
by: Lee, Junho, et al.
Published: (2026)
Is There a Better Source Distribution than Gaussian? Exploring Source Distributions for Image Flow Matching
by: Lee, Junho, et al.
Published: (2025)
by: Lee, Junho, et al.
Published: (2025)
Text2HOI: Text-guided 3D Motion Generation for Hand-Object Interaction
by: Cha, Junuk, et al.
Published: (2024)
by: Cha, Junuk, et al.
Published: (2024)
Scalable Frame Sampling for Video Classification: A Semi-Optimal Policy Approach with Reduced Search Space
by: Lee, Junho, et al.
Published: (2024)
by: Lee, Junho, et al.
Published: (2024)
Modality-Aware Representation Learning for Zero-shot Sketch-based Image Retrieval
by: Lyou, Eunyi, et al.
Published: (2024)
by: Lyou, Eunyi, et al.
Published: (2024)
How Do Large Vision-Language Models See Text in Image? Unveiling the Distinctive Role of OCR Heads
by: Baek, Ingeol, et al.
Published: (2025)
by: Baek, Ingeol, et al.
Published: (2025)
Self-Guided Masked Autoencoder
by: Shin, Jeongwoo, et al.
Published: (2025)
by: Shin, Jeongwoo, et al.
Published: (2025)
Pixel-aligned RGB-NIR Stereo Imaging and Dataset for Robot Vision
by: Kim, Jinnyeong, et al.
Published: (2024)
by: Kim, Jinnyeong, et al.
Published: (2024)
Latent Diffusion Models with Masked AutoEncoders
by: Lee, Junho, et al.
Published: (2025)
by: Lee, Junho, et al.
Published: (2025)
DeltaSpace: A Semantic-aligned Feature Space for Flexible Text-guided Image Editing
by: Lyu, Yueming, et al.
Published: (2023)
by: Lyu, Yueming, et al.
Published: (2023)
Equivariant Latent Alignment via Flow Matching under Group Symmetries
by: Kim, Sunghyun, et al.
Published: (2026)
by: Kim, Sunghyun, et al.
Published: (2026)
Finding NeMo: Negative-mined Mosaic Augmentation for Referring Image Segmentation
by: Ha, Seongsu, et al.
Published: (2024)
by: Ha, Seongsu, et al.
Published: (2024)
Enhanced Detection of Tiny Objects in Aerial Images
by: Kim, Kihyun, et al.
Published: (2025)
by: Kim, Kihyun, et al.
Published: (2025)
CompBench: Benchmarking Complex Instruction-guided Image Editing
by: Jia, Bohan, et al.
Published: (2025)
by: Jia, Bohan, et al.
Published: (2025)
ArtSplat: Feed-Forward Articulated 3D Gaussian Splatting from Sparse Multi-State Uncalibrated Views
by: Lee, Inseo, et al.
Published: (2026)
by: Lee, Inseo, et al.
Published: (2026)
Preserve or Modify? Context-Aware Evaluation for Balancing Preservation and Modification in Text-Guided Image Editing
by: Kim, Yoonjeon, et al.
Published: (2024)
by: Kim, Yoonjeon, et al.
Published: (2024)
Efficient Generative Modeling beyond Memoryless Diffusion via Adjoint Schrödinger Bridge Matching
by: Shin, Jeongwoo, et al.
Published: (2026)
by: Shin, Jeongwoo, et al.
Published: (2026)
Minecraft-ify: Minecraft Style Image Generation with Text-guided Image Editing for In-Game Application
by: Kim, Bumsoo, et al.
Published: (2024)
by: Kim, Bumsoo, et al.
Published: (2024)
GaussianVideo: Efficient Video Representation and Compression by Gaussian Splatting
by: Lee, Inseo, et al.
Published: (2025)
by: Lee, Inseo, et al.
Published: (2025)
Safeguard Text-to-Image Diffusion Models with Human Feedback Inversion
by: Kim, Sanghyun, et al.
Published: (2024)
by: Kim, Sanghyun, et al.
Published: (2024)
Isometric Representation Learning for Disentangled Latent Space of Diffusion Models
by: Hahm, Jaehoon, et al.
Published: (2024)
by: Hahm, Jaehoon, et al.
Published: (2024)
StochCA: A Novel Approach for Exploiting Pretrained Models with Cross-Attention
by: Seo, Seungwon, et al.
Published: (2024)
by: Seo, Seungwon, et al.
Published: (2024)
CharDiff-LP: A Diffusion Model with Character-Level Guidance for License Plate Image Restoration
by: Na, Kihyun, et al.
Published: (2025)
by: Na, Kihyun, et al.
Published: (2025)
Latent Expression Generation for Referring Image Segmentation and Grounding
by: Yu, Seonghoon, et al.
Published: (2025)
by: Yu, Seonghoon, et al.
Published: (2025)
Negative-prompt Inversion: Fast Image Inversion for Editing with Text-guided Diffusion Models
by: Miyake, Daiki, et al.
Published: (2023)
by: Miyake, Daiki, et al.
Published: (2023)
B-RIGHT: Benchmark Re-evaluation for Integrity in Generalized Human-Object Interaction Testing
by: Jang, Yoojin, et al.
Published: (2025)
by: Jang, Yoojin, et al.
Published: (2025)
EVIDENT: Routing MLLM Adaptation through Entity-Grounded Visual Evidence for Cross-Domain Video Temporal Grounding
by: Ahn, Geo, et al.
Published: (2026)
by: Ahn, Geo, et al.
Published: (2026)
UniRef-Image-Edit: Towards Scalable and Consistent Multi-Reference Image Editing
by: Wei, Hongyang, et al.
Published: (2026)
by: Wei, Hongyang, et al.
Published: (2026)
ADP-DiT: Text-Guided Diffusion Transformer for Brain Image Generation in Alzheimer's Disease Progression
by: Lee, Juneyong, et al.
Published: (2026)
by: Lee, Juneyong, et al.
Published: (2026)
TextSculptor: Training and Benchmarking Scene Text Editing
by: Lin, Yiheng, et al.
Published: (2026)
by: Lin, Yiheng, et al.
Published: (2026)
ResetEdit: Precise Text-guided Editing of Generated Image via Resettable Starting Latent
by: Wang, Hanyi, et al.
Published: (2026)
by: Wang, Hanyi, et al.
Published: (2026)
Zero-shot Text-guided Infinite Image Synthesis with LLM guidance
by: Kwon, Soyeong, et al.
Published: (2024)
by: Kwon, Soyeong, et al.
Published: (2024)
TTD: Text-Tag Self-Distillation Enhancing Image-Text Alignment in CLIP to Alleviate Single Tag Bias
by: Jo, Sanghyun, et al.
Published: (2024)
by: Jo, Sanghyun, et al.
Published: (2024)
Revealing Directions for Text-guided 3D Face Editing
by: Chen, Zhuo, et al.
Published: (2024)
by: Chen, Zhuo, et al.
Published: (2024)
Neural Image Compression with Text-guided Encoding for both Pixel-level and Perceptual Fidelity
by: Lee, Hagyeong, et al.
Published: (2024)
by: Lee, Hagyeong, et al.
Published: (2024)
Towards Scalable and Consistent 3D Editing
by: Xia, Ruihao, et al.
Published: (2025)
by: Xia, Ruihao, et al.
Published: (2025)
Exploring Multimodal Diffusion Transformers for Enhanced Prompt-based Image Editing
by: Shin, Joonghyuk, et al.
Published: (2025)
by: Shin, Joonghyuk, et al.
Published: (2025)
GIE-Bench: Towards Grounded Evaluation for Text-Guided Image Editing
by: Qian, Yusu, et al.
Published: (2025)
by: Qian, Yusu, et al.
Published: (2025)
Image Clustering Conditioned on Text Criteria
by: Kwon, Sehyun, et al.
Published: (2023)
by: Kwon, Sehyun, et al.
Published: (2023)
Similar Items
-
Towards Motion-aware Referring Image Segmentation
by: Kim, Chaeyun, et al.
Published: (2026) -
Geometry-Aware Image Flow Matching
by: Lee, Junho, et al.
Published: (2026) -
Is There a Better Source Distribution than Gaussian? Exploring Source Distributions for Image Flow Matching
by: Lee, Junho, et al.
Published: (2025) -
Text2HOI: Text-guided 3D Motion Generation for Hand-Object Interaction
by: Cha, Junuk, et al.
Published: (2024) -
Scalable Frame Sampling for Video Classification: A Semi-Optimal Policy Approach with Reduced Search Space
by: Lee, Junho, et al.
Published: (2024)