OSPO: Object-Centric Self-Improving Preference Optimization for Text-to-Image Generation
Fuente:
arXiv
Guardado en:
| Autores principales: | Oh, Yoonjin, Kim, Yongjin, Kim, Hyomin, Chi, Donghwan, Kim, Sungwoong |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Slot-MLLM: Object-Centric Visual Tokenization for Multimodal LLM
por: Chi, Donghwan, et al.
Publicado: (2025)
por: Chi, Donghwan, et al.
Publicado: (2025)
MEVG: Multi-event Video Generation with Text-to-Video Models
por: Oh, Gyeongrok, et al.
Publicado: (2023)
por: Oh, Gyeongrok, et al.
Publicado: (2023)
Learning Object-Centric Representations in SAR Images with Multi-Level Feature Fusion
por: Jang, Oh-Tae, et al.
Publicado: (2025)
por: Jang, Oh-Tae, et al.
Publicado: (2025)
How Well Do Vision--Language Models Understand Cities? A Comparative Study on Spatial Reasoning from Street-View Images
por: Ro, Juneyoung, et al.
Publicado: (2025)
por: Ro, Juneyoung, et al.
Publicado: (2025)
Improving Text-to-Image Generation with Intrinsic Self-Confidence Rewards
por: Kim, Seungwook, et al.
Publicado: (2026)
por: Kim, Seungwook, et al.
Publicado: (2026)
LieHMR: Autoregressive Human Mesh Recovery with $SO(3)$ Diffusion
por: Kim, Donghwan, et al.
Publicado: (2025)
por: Kim, Donghwan, et al.
Publicado: (2025)
EraseLoRA: MLLM-Driven Foreground Exclusion and Background Subtype Aggregation for Dataset-Free Object Removal
por: Jo, Sanghyun, et al.
Publicado: (2025)
por: Jo, Sanghyun, et al.
Publicado: (2025)
Leveraging Image Augmentation for Object Manipulation: Towards Interpretable Controllability in Object-Centric Learning
por: Kim, Jinwoo, et al.
Publicado: (2023)
por: Kim, Jinwoo, et al.
Publicado: (2023)
Inlier-Centric Post-Training Quantization for Object Detection Models
por: Kim, Minsu, et al.
Publicado: (2026)
por: Kim, Minsu, et al.
Publicado: (2026)
Multi-hypotheses Conditioned Point Cloud Diffusion for 3D Human Reconstruction from Occluded Images
por: Kim, Donghwan, et al.
Publicado: (2024)
por: Kim, Donghwan, et al.
Publicado: (2024)
Task-Decoupled Image Inpainting Framework for Class-specific Object Remover
por: Oh, Changsuk, et al.
Publicado: (2024)
por: Oh, Changsuk, et al.
Publicado: (2024)
FastSTAR: Spatiotemporal Token Pruning for Efficient Autoregressive Video Synthesis
por: Yune, Sungwoong, et al.
Publicado: (2026)
por: Yune, Sungwoong, et al.
Publicado: (2026)
DeClotH: Decomposable 3D Cloth and Human Body Reconstruction from a Single Image
por: Nam, Hyeongjin, et al.
Publicado: (2025)
por: Nam, Hyeongjin, et al.
Publicado: (2025)
Introducing VaDA: Novel Image Segmentation Model for Maritime Object Segmentation Using New Dataset
por: Kim, Yongjin, et al.
Publicado: (2024)
por: Kim, Yongjin, et al.
Publicado: (2024)
AURA : Automatic Mask Generator using Randomized Input Sampling for Object Removal
por: Oh, Changsuk, et al.
Publicado: (2023)
por: Oh, Changsuk, et al.
Publicado: (2023)
PhysHanDI: Physics-Based Reconstruction of Hand-Deformable Object Interactions
por: Lee, Jihyun, et al.
Publicado: (2026)
por: Lee, Jihyun, et al.
Publicado: (2026)
OmniText: A Training-Free Generalist for Controllable Text-Image Manipulation
por: Gunawan, Agus, et al.
Publicado: (2025)
por: Gunawan, Agus, et al.
Publicado: (2025)
EVT: Efficient View Transformation for Multi-Modal 3D Object Detection
por: Lee, Yongjin, et al.
Publicado: (2024)
por: Lee, Yongjin, et al.
Publicado: (2024)
LPOI: Listwise Preference Optimization for Vision Language Models
por: Zadeh, Fatemeh Pesaran, et al.
Publicado: (2025)
por: Zadeh, Fatemeh Pesaran, et al.
Publicado: (2025)
Camera Splatting for Continuous View Optimization
por: Lee, Gahye, et al.
Publicado: (2025)
por: Lee, Gahye, et al.
Publicado: (2025)
Object Remover Performance Evaluation Methods using Class-wise Object Removal Images
por: Oh, Changsuk, et al.
Publicado: (2024)
por: Oh, Changsuk, et al.
Publicado: (2024)
TTD: Text-Tag Self-Distillation Enhancing Image-Text Alignment in CLIP to Alleviate Single Tag Bias
por: Jo, Sanghyun, et al.
Publicado: (2024)
por: Jo, Sanghyun, et al.
Publicado: (2024)
Scalable Ranked Preference Optimization for Text-to-Image Generation
por: Karthik, Shyamgopal, et al.
Publicado: (2024)
por: Karthik, Shyamgopal, et al.
Publicado: (2024)
Object-Centric Representation Learning for Enhanced 3D Semantic Scene Graph Prediction
por: Heo, KunHo, et al.
Publicado: (2025)
por: Heo, KunHo, et al.
Publicado: (2025)
PARTE: Part-Guided Texturing for 3D Human Reconstruction from a Single Image
por: Nam, Hyeongjin, et al.
Publicado: (2025)
por: Nam, Hyeongjin, et al.
Publicado: (2025)
TextBoost: Boosting Text Encoder for Personalized Text-to-Image Generation
por: Park, NaHyeon, et al.
Publicado: (2024)
por: Park, NaHyeon, et al.
Publicado: (2024)
Diffusion-driven GAN Inversion for Multi-Modal Face Image Generation
por: Kim, Jihyun, et al.
Publicado: (2024)
por: Kim, Jihyun, et al.
Publicado: (2024)
IRASNet: Improved Feature-Level Clutter Reduction for Domain Generalized SAR-ATR
por: Jang, Oh-Tae, et al.
Publicado: (2024)
por: Jang, Oh-Tae, et al.
Publicado: (2024)
MonoWAD: Weather-Adaptive Diffusion Model for Robust Monocular 3D Object Detection
por: Oh, Youngmin, et al.
Publicado: (2024)
por: Oh, Youngmin, et al.
Publicado: (2024)
MATANet: A Multi-context Attention and Taxonomy-Aware Network for Fine-Grained Underwater Recognition of Marine Species
por: Lee, Donghwan, et al.
Publicado: (2026)
por: Lee, Donghwan, et al.
Publicado: (2026)
TIV-Diffusion: Towards Object-Centric Movement for Text-driven Image to Video Generation
por: Wang, Xingrui, et al.
Publicado: (2024)
por: Wang, Xingrui, et al.
Publicado: (2024)
OFF-CLIP: Improving Normal Detection Confidence in Radiology CLIP with Simple Off-Diagonal Term Auto-Adjustment
por: Park, Junhyun, et al.
Publicado: (2025)
por: Park, Junhyun, et al.
Publicado: (2025)
Predict-Optimize-Distill: A Self-Improving Cycle for 4D Object Understanding
por: Wu, Mingxuan, et al.
Publicado: (2025)
por: Wu, Mingxuan, et al.
Publicado: (2025)
Model Agnostic Preference Optimization for Medical Image Segmentation
por: Nam, Yunseong, et al.
Publicado: (2025)
por: Nam, Yunseong, et al.
Publicado: (2025)
Dual Recursive Feedback on Generation and Appearance Latents for Pose-Robust Text-to-Image Diffusion
por: Kim, Jiwon, et al.
Publicado: (2025)
por: Kim, Jiwon, et al.
Publicado: (2025)
Prompt Augmentation for Self-supervised Text-guided Image Manipulation
por: Bodur, Rumeysa, et al.
Publicado: (2024)
por: Bodur, Rumeysa, et al.
Publicado: (2024)
MedROI: Codec-Agnostic Region of Interest-Centric Compression for Medical Images
por: Kim, Jiwon, et al.
Publicado: (2026)
por: Kim, Jiwon, et al.
Publicado: (2026)
Disentangled Object-Centric Image Representation for Robotic Manipulation
por: Emukpere, David, et al.
Publicado: (2025)
por: Emukpere, David, et al.
Publicado: (2025)
SlotVTG: Object-Centric Adapter for Generalizable Video Temporal Grounding
por: Han, Jiwook, et al.
Publicado: (2026)
por: Han, Jiwook, et al.
Publicado: (2026)
Neural Image Compression with Text-guided Encoding for both Pixel-level and Perceptual Fidelity
por: Lee, Hagyeong, et al.
Publicado: (2024)
por: Lee, Hagyeong, et al.
Publicado: (2024)
Ejemplares similares
-
Slot-MLLM: Object-Centric Visual Tokenization for Multimodal LLM
por: Chi, Donghwan, et al.
Publicado: (2025) -
MEVG: Multi-event Video Generation with Text-to-Video Models
por: Oh, Gyeongrok, et al.
Publicado: (2023) -
Learning Object-Centric Representations in SAR Images with Multi-Level Feature Fusion
por: Jang, Oh-Tae, et al.
Publicado: (2025) -
How Well Do Vision--Language Models Understand Cities? A Comparative Study on Spatial Reasoning from Street-View Images
por: Ro, Juneyoung, et al.
Publicado: (2025) -
Improving Text-to-Image Generation with Intrinsic Self-Confidence Rewards
por: Kim, Seungwook, et al.
Publicado: (2026)