Imagining the Unseen: Generative Location Modeling for Object Placement
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Yun, Jooyeol, Abati, Davide, Omran, Mohamed, Choo, Jaegul, Habibian, Amirhossein, Wiggers, Auke |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Scene-Aware Location Modeling for Data Augmentation in Automotive Object Detection
von: Petersen, Jens, et al.
Veröffentlicht: (2025)
von: Petersen, Jens, et al.
Veröffentlicht: (2025)
Controllable 3D Placement of Objects with Scene-Aware Diffusion Models
von: Omran, Mohamed, et al.
Veröffentlicht: (2025)
von: Omran, Mohamed, et al.
Veröffentlicht: (2025)
Gaussian Splatting is an Effective Data Generator for 3D Object Detection
von: Zanjani, Farhad G., et al.
Veröffentlicht: (2025)
von: Zanjani, Farhad G., et al.
Veröffentlicht: (2025)
Hybrid Gaussian Splatting for Novel Urban View Synthesis
von: Omran, Mohamed, et al.
Veröffentlicht: (2025)
von: Omran, Mohamed, et al.
Veröffentlicht: (2025)
Scaling Up Personalized Image Aesthetic Assessment via Task Vector Customization
von: Yun, Jooyeol, et al.
Veröffentlicht: (2024)
von: Yun, Jooyeol, et al.
Veröffentlicht: (2024)
Vector Prism: Animating Vector Graphics by Stratifying Semantic Structure
von: Yun, Jooyeol, et al.
Veröffentlicht: (2025)
von: Yun, Jooyeol, et al.
Veröffentlicht: (2025)
Devil is in the Detail: Towards Injecting Fine Details of Image Prompt in Image Generation via Conflict-free Guidance and Stratified Attention
von: Jo, Kyungmin, et al.
Veröffentlicht: (2025)
von: Jo, Kyungmin, et al.
Veröffentlicht: (2025)
Regularized Training with Generated Datasets for Name-Only Transfer of Vision-Language Models
von: Park, Minho, et al.
Veröffentlicht: (2024)
von: Park, Minho, et al.
Veröffentlicht: (2024)
Enabling Region-Specific Control via Lassos in Point-Based Colorization
von: Lee, Sanghyeon, et al.
Veröffentlicht: (2024)
von: Lee, Sanghyeon, et al.
Veröffentlicht: (2024)
SphereDiff: Tuning-free 360° Static and Dynamic Panorama Generation via Spherical Latent Representation
von: Park, Minho, et al.
Veröffentlicht: (2025)
von: Park, Minho, et al.
Veröffentlicht: (2025)
Object-Centric Diffusion for Efficient Video Editing
von: Kahatapitiya, Kumara, et al.
Veröffentlicht: (2024)
von: Kahatapitiya, Kumara, et al.
Veröffentlicht: (2024)
DesignLab: Designing Slides Through Iterative Detection and Correction
von: Yun, Jooyeol, et al.
Veröffentlicht: (2025)
von: Yun, Jooyeol, et al.
Veröffentlicht: (2025)
Skip-and-Play: Depth-Driven Pose-Preserved Image Generation for Any Objects
von: Jo, Kyungmin, et al.
Veröffentlicht: (2024)
von: Jo, Kyungmin, et al.
Veröffentlicht: (2024)
ReHyAt: Recurrent Hybrid Attention for Video Diffusion Transformers
von: Ghafoorian, Mohsen, et al.
Veröffentlicht: (2026)
von: Ghafoorian, Mohsen, et al.
Veröffentlicht: (2026)
Multi-Scale Local Speculative Decoding for Image Generation
von: Peruzzo, Elia, et al.
Veröffentlicht: (2026)
von: Peruzzo, Elia, et al.
Veröffentlicht: (2026)
Attention Surgery: An Efficient Recipe to Linearize Your Video Diffusion Transformer
von: Ghafoorian, Mohsen, et al.
Veröffentlicht: (2025)
von: Ghafoorian, Mohsen, et al.
Veröffentlicht: (2025)
Temporal In-Context Fine-Tuning with Temporal Reasoning for Versatile Control of Video Diffusion Models
von: Kim, Kinam, et al.
Veröffentlicht: (2025)
von: Kim, Kinam, et al.
Veröffentlicht: (2025)
Enhancing Novel View Synthesis via Geometry Grounded Set Diffusion
von: Zanjani, Farhad G., et al.
Veröffentlicht: (2026)
von: Zanjani, Farhad G., et al.
Veröffentlicht: (2026)
TV-LiVE: Training-Free, Text-Guided Video Editing via Layer Informed Vitality Exploitation
von: Kim, Min-Jung, et al.
Veröffentlicht: (2025)
von: Kim, Min-Jung, et al.
Veröffentlicht: (2025)
GaussianMotion: End-to-End Learning of Animatable Gaussian Avatars with Pose Guidance from Text
von: Shim, Gyumin, et al.
Veröffentlicht: (2025)
von: Shim, Gyumin, et al.
Veröffentlicht: (2025)
From Wardrobe to Canvas: Wardrobe Polyptych LoRA for Part-level Controllable Human Image Generation
von: Kim, Jeongho, et al.
Veröffentlicht: (2025)
von: Kim, Jeongho, et al.
Veröffentlicht: (2025)
PyramidalWan: On Making Pretrained Video Model Pyramidal for Efficient Inference
von: Korzhenkov, Denis, et al.
Veröffentlicht: (2026)
von: Korzhenkov, Denis, et al.
Veröffentlicht: (2026)
The 1st International Workshop on Disentangled Representation Learning for Controllable Generation (DRL4Real): Methods and Results
von: Chen, Qiuyu, et al.
Veröffentlicht: (2025)
von: Chen, Qiuyu, et al.
Veröffentlicht: (2025)
Learning to See What You Need: Gaze Attention for Multimodal Large Language Models
von: Song, Junha, et al.
Veröffentlicht: (2026)
von: Song, Junha, et al.
Veröffentlicht: (2026)
Clockwork Diffusion: Efficient Generation With Model-Step Distillation
von: Habibian, Amirhossein, et al.
Veröffentlicht: (2023)
von: Habibian, Amirhossein, et al.
Veröffentlicht: (2023)
LocPoseNet: Robust Location Prior for Unseen Object Pose Estimation
von: Zhao, Chen, et al.
Veröffentlicht: (2022)
von: Zhao, Chen, et al.
Veröffentlicht: (2022)
Bones Can't Be Triangles: Accurate and Efficient Vertebrae Keypoint Estimation through Collaborative Error Revision
von: Kim, Jinhee, et al.
Veröffentlicht: (2024)
von: Kim, Jinhee, et al.
Veröffentlicht: (2024)
Training Spatial-Frequency Visual Prompts and Probabilistic Clusters for Accurate Black-Box Transfer Learning
von: Cho, Wonwoo, et al.
Veröffentlicht: (2024)
von: Cho, Wonwoo, et al.
Veröffentlicht: (2024)
Enhancing Intrinsic Features for Debiasing via Investigating Class-Discerning Common Attributes in Bias-Contrastive Pair
von: Park, Jeonghoon, et al.
Veröffentlicht: (2024)
von: Park, Jeonghoon, et al.
Veröffentlicht: (2024)
What to Preserve and What to Transfer: Faithful, Identity-Preserving Diffusion-based Hairstyle Transfer
von: Chung, Chaeyeon, et al.
Veröffentlicht: (2024)
von: Chung, Chaeyeon, et al.
Veröffentlicht: (2024)
Zero-Shot Head Swapping in Real-World Scenarios
von: Kang, Taewoong, et al.
Veröffentlicht: (2025)
von: Kang, Taewoong, et al.
Veröffentlicht: (2025)
Seeing the Unseen: Visual Common Sense for Semantic Placement
von: Ramrakhya, Ram, et al.
Veröffentlicht: (2024)
von: Ramrakhya, Ram, et al.
Veröffentlicht: (2024)
Cross-Frame Representation Alignment for Fine-Tuning Video Diffusion Models
von: Hwang, Sungwon, et al.
Veröffentlicht: (2025)
von: Hwang, Sungwon, et al.
Veröffentlicht: (2025)
MoAlign: Motion-Centric Representation Alignment for Video Diffusion Models
von: Bhowmik, Aritra, et al.
Veröffentlicht: (2025)
von: Bhowmik, Aritra, et al.
Veröffentlicht: (2025)
Infinite-Homography as Robust Conditioning for Camera-Controlled Video Generation
von: Kim, Min-Jung, et al.
Veröffentlicht: (2025)
von: Kim, Min-Jung, et al.
Veröffentlicht: (2025)
PromptDresser: Improving the Quality and Controllability of Virtual Try-On via Generative Textual Prompt and Prompt-aware Mask
von: Kim, Jeongho, et al.
Veröffentlicht: (2024)
von: Kim, Jeongho, et al.
Veröffentlicht: (2024)
Gated Relational Alignment via Confidence-based Distillation for Efficient VLMs
von: Chen, Yanlong, et al.
Veröffentlicht: (2026)
von: Chen, Yanlong, et al.
Veröffentlicht: (2026)
An Empirical Study of the Generalization Ability of Lidar 3D Object Detectors to Unseen Domains
von: Eskandar, George, et al.
Veröffentlicht: (2024)
von: Eskandar, George, et al.
Veröffentlicht: (2024)
RL makes MLLMs see better than SFT
von: Song, Junha, et al.
Veröffentlicht: (2025)
von: Song, Junha, et al.
Veröffentlicht: (2025)
Adapting Segment Anything Model for Unseen Object Instance Segmentation
von: Cao, Rui, et al.
Veröffentlicht: (2024)
von: Cao, Rui, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Scene-Aware Location Modeling for Data Augmentation in Automotive Object Detection
von: Petersen, Jens, et al.
Veröffentlicht: (2025) -
Controllable 3D Placement of Objects with Scene-Aware Diffusion Models
von: Omran, Mohamed, et al.
Veröffentlicht: (2025) -
Gaussian Splatting is an Effective Data Generator for 3D Object Detection
von: Zanjani, Farhad G., et al.
Veröffentlicht: (2025) -
Hybrid Gaussian Splatting for Novel Urban View Synthesis
von: Omran, Mohamed, et al.
Veröffentlicht: (2025) -
Scaling Up Personalized Image Aesthetic Assessment via Task Vector Customization
von: Yun, Jooyeol, et al.
Veröffentlicht: (2024)