Finding 3D Scene Analogies with Multimodal Foundation Models
Fuente:
arXiv
Saved in:
| Main Authors: | Kim, Junho, Kim, Young Min |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Learning 3D Scene Analogies with Neural Contextual Scene Maps
by: Kim, Junho, et al.
Published: (2025)
by: Kim, Junho, et al.
Published: (2025)
Analogical Trajectory Transfer
by: Kim, Junho, et al.
Published: (2026)
by: Kim, Junho, et al.
Published: (2026)
LTGS: Long-Term Gaussian Scene Chronology From Sparse View Updates
by: Kim, Minkwan, et al.
Published: (2025)
by: Kim, Minkwan, et al.
Published: (2025)
Fully Geometric Panoramic Localization
by: Kim, Junho, et al.
Published: (2024)
by: Kim, Junho, et al.
Published: (2024)
Calibrating Panoramic Depth Estimation for Practical Localization and Mapping
by: Kim, Junho, et al.
Published: (2023)
by: Kim, Junho, et al.
Published: (2023)
RoEL: Robust Event-based 3D Line Reconstruction
by: Bae, Gwangtak, et al.
Published: (2026)
by: Bae, Gwangtak, et al.
Published: (2026)
PICCOLO: Point Cloud-Centric Omnidirectional Localization
by: Kim, Junho, et al.
Published: (2021)
by: Kim, Junho, et al.
Published: (2021)
CPO: Change Robust Panorama to Point Cloud Localization
by: Kim, Junho, et al.
Published: (2022)
by: Kim, Junho, et al.
Published: (2022)
Humans as a Calibration Pattern: Dynamic 3D Scene Reconstruction from Unsynchronized and Uncalibrated Videos
by: Choi, Changwoon, et al.
Published: (2024)
by: Choi, Changwoon, et al.
Published: (2024)
Let 2D Diffusion Model Know 3D-Consistency for Robust Text-to-3D Generation
by: Seo, Junyoung, et al.
Published: (2023)
by: Seo, Junyoung, et al.
Published: (2023)
Geometry-Aware Scene Configurations for Novel View Synthesis
by: Kim, Minkwan, et al.
Published: (2025)
by: Kim, Minkwan, et al.
Published: (2025)
SceneLinker: Compositional 3D Scene Generation via Semantic Scene Graph from RGB Sequences
by: Kim, Seok-Young, et al.
Published: (2026)
by: Kim, Seok-Young, et al.
Published: (2026)
DIP-R1: Deep Inspection and Perception with RL Looking Through and Understanding Complex Scenes
by: Park, Sungjune, et al.
Published: (2025)
by: Park, Sungjune, et al.
Published: (2025)
SceneMI: Motion In-betweening for Modeling Human-Scene Interactions
by: Hwang, Inwoo, et al.
Published: (2025)
by: Hwang, Inwoo, et al.
Published: (2025)
CapeLLM: Support-Free Category-Agnostic Pose Estimation with Multimodal Large Language Models
by: Kim, Junho, et al.
Published: (2024)
by: Kim, Junho, et al.
Published: (2024)
Event-Driven Storytelling with Multiple Lifelike Humans in a 3D Scene
by: Lim, Donggeun, et al.
Published: (2025)
by: Lim, Donggeun, et al.
Published: (2025)
Finding NeMo: Negative-mined Mosaic Augmentation for Referring Image Segmentation
by: Ha, Seongsu, et al.
Published: (2024)
by: Ha, Seongsu, et al.
Published: (2024)
ReMP: Reusable Motion Prior for Multi-domain 3D Human Pose Estimation and Motion Inbetweening
by: Jang, Hojun, et al.
Published: (2024)
by: Jang, Hojun, et al.
Published: (2024)
EggHand: A Multimodal Foundation Model for Egocentric Hand Pose Forecasting
by: Choi, Jaeyoung, et al.
Published: (2026)
by: Choi, Jaeyoung, et al.
Published: (2026)
CAPA: Contribution-Aware Pruning and FFN Approximation for Efficient Large Vision-Language Models
by: Jha, Samyak, et al.
Published: (2026)
by: Jha, Samyak, et al.
Published: (2026)
ExCellGen: Fast, Controllable, Photorealistic 3D Scene Generation from a Single Real-World Exemplar
by: Jambon, Clément, et al.
Published: (2024)
by: Jambon, Clément, et al.
Published: (2024)
3Doodle: Compact Abstraction of Objects with 3D Strokes
by: Choi, Changwoon, et al.
Published: (2024)
by: Choi, Changwoon, et al.
Published: (2024)
Privacy-Preserving Visual Localization with Event Cameras
by: Kim, Junho, et al.
Published: (2022)
by: Kim, Junho, et al.
Published: (2022)
Towards Holistic Surgical Scene Graph
by: Shin, Jongmin, et al.
Published: (2025)
by: Shin, Jongmin, et al.
Published: (2025)
Recovering Dynamic 3D Sketches from Videos
by: Lee, Jaeah, et al.
Published: (2025)
by: Lee, Jaeah, et al.
Published: (2025)
D-Cube: Exploiting Hyper-Features of Diffusion Model for Robust Medical Classification
by: Jang, Minhee, et al.
Published: (2024)
by: Jang, Minhee, et al.
Published: (2024)
VEGS: View Extrapolation of Urban Scenes in 3D Gaussian Splatting using Learned Priors
by: Hwang, Sungwon, et al.
Published: (2024)
by: Hwang, Sungwon, et al.
Published: (2024)
Gaussian Difference: Find Any Change Instance in 3D Scenes
by: Jiang, Binbin, et al.
Published: (2025)
by: Jiang, Binbin, et al.
Published: (2025)
OpenSU3D: Open World 3D Scene Understanding using Foundation Models
by: Mohiuddin, Rafay, et al.
Published: (2024)
by: Mohiuddin, Rafay, et al.
Published: (2024)
Aligned Novel View Image and Geometry Synthesis via Cross-modal Attention Instillation
by: Kwak, Min-Seop, et al.
Published: (2025)
by: Kwak, Min-Seop, et al.
Published: (2025)
D$^2$USt3R: Enhancing 3D Reconstruction for Dynamic Scenes
by: Han, Jisang, et al.
Published: (2025)
by: Han, Jisang, et al.
Published: (2025)
Integrating Meshes and 3D Gaussians for Indoor Scene Reconstruction with SAM Mask Guidance
by: Kim, Jiyeop, et al.
Published: (2024)
by: Kim, Jiyeop, et al.
Published: (2024)
CODE: Contrasting Self-generated Description to Combat Hallucination in Large Multi-modal Models
by: Kim, Junho, et al.
Published: (2024)
by: Kim, Junho, et al.
Published: (2024)
OnlineBEV: Recurrent Temporal Fusion in Bird's Eye View Representations for Multi-Camera 3D Perception
by: Koh, Junho, et al.
Published: (2025)
by: Koh, Junho, et al.
Published: (2025)
Difference Inversion: Interpolate and Isolate the Difference with Token Consistency for Image Analogy Generation
by: Kim, Hyunsoo, et al.
Published: (2025)
by: Kim, Hyunsoo, et al.
Published: (2025)
Is There a Better Source Distribution than Gaussian? Exploring Source Distributions for Image Flow Matching
by: Lee, Junho, et al.
Published: (2025)
by: Lee, Junho, et al.
Published: (2025)
Event-based Facial Keypoint Alignment via Cross-Modal Fusion Attention and Self-Supervised Multi-Event Representation Learning
by: Kang, Donghwa, et al.
Published: (2025)
by: Kang, Donghwa, et al.
Published: (2025)
Geometry-Aware Image Flow Matching
by: Lee, Junho, et al.
Published: (2026)
by: Lee, Junho, et al.
Published: (2026)
DepthFocus: Controllable Depth Estimation for See-Through Scenes
by: Min, Junhong, et al.
Published: (2025)
by: Min, Junhong, et al.
Published: (2025)
VG3T: Visual Geometry Grounded Gaussian Transformer
by: Kim, Junho, et al.
Published: (2025)
by: Kim, Junho, et al.
Published: (2025)
Similar Items
-
Learning 3D Scene Analogies with Neural Contextual Scene Maps
by: Kim, Junho, et al.
Published: (2025) -
Analogical Trajectory Transfer
by: Kim, Junho, et al.
Published: (2026) -
LTGS: Long-Term Gaussian Scene Chronology From Sparse View Updates
by: Kim, Minkwan, et al.
Published: (2025) -
Fully Geometric Panoramic Localization
by: Kim, Junho, et al.
Published: (2024) -
Calibrating Panoramic Depth Estimation for Practical Localization and Mapping
by: Kim, Junho, et al.
Published: (2023)