Grounding World Simulation Models in a Real-World Metropolis
Fuente:
arXiv
Guardado en:
| Autores principales: | Seo, Junyoung, Choi, Hyunwook, Kwon, Minkyung, Choi, Jinhyeok, Jin, Siyoon, Lee, Gayoung, Kim, Junho, Lee, JoungBin, Gu, Geonmo, Han, Dongyoon, Yun, Sangdoo, Kim, Seungryong, Kim, Jin-Hwa |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
CAMEO: Correspondence-Attention Alignment for Multi-View Diffusion Models
por: Kwon, Minkyung, et al.
Publicado: (2025)
por: Kwon, Minkyung, et al.
Publicado: (2025)
WorldCam: Interactive Autoregressive 3D Gaming Worlds with Camera Pose as a Unifying Geometric Representation
por: Nam, Jisu, et al.
Publicado: (2026)
por: Nam, Jisu, et al.
Publicado: (2026)
Aligned Novel View Image and Geometry Synthesis via Cross-modal Attention Instillation
por: Kwak, Min-Seop, et al.
Publicado: (2025)
por: Kwak, Min-Seop, et al.
Publicado: (2025)
MORPHOS: Autoregressive 4D Generation with Temporal Structured Latents
por: Kwon, Minkyung, et al.
Publicado: (2026)
por: Kwon, Minkyung, et al.
Publicado: (2026)
APPLE: Attribute-Preserving Pseudo-Labeling for Diffusion-Based Face Swapping
por: Kang, Jiwon, et al.
Publicado: (2026)
por: Kang, Jiwon, et al.
Publicado: (2026)
MATRIX: Mask Track Alignment for Interaction-aware Video Generation
por: Jin, Siyoon, et al.
Publicado: (2025)
por: Jin, Siyoon, et al.
Publicado: (2025)
Pygmalion Effect in Vision: Image-to-Clay Translation for Reflective Geometry Reconstruction
por: Lee, Gayoung, et al.
Publicado: (2025)
por: Lee, Gayoung, et al.
Publicado: (2025)
Projected Representation Conditioning for High-fidelity Novel View Synthesis
por: Kwak, Min-Seop, et al.
Publicado: (2026)
por: Kwak, Min-Seop, et al.
Publicado: (2026)
Pose-dIVE: Pose-Diversified Augmentation with Diffusion Model for Person Re-Identification
por: Kim, Inès Hyeonsu, et al.
Publicado: (2024)
por: Kim, Inès Hyeonsu, et al.
Publicado: (2024)
Match me if you can: Semi-Supervised Semantic Correspondence Learning with Unpaired Images
por: Kim, Jiwon, et al.
Publicado: (2023)
por: Kim, Jiwon, et al.
Publicado: (2023)
Direct Unlearning Optimization for Robust and Safe Text-to-Image Models
por: Park, Yong-Hyun, et al.
Publicado: (2024)
por: Park, Yong-Hyun, et al.
Publicado: (2024)
WorldKV: Efficient World Memory with World Retrieval and Compression
por: Yi, Jung, et al.
Publicado: (2026)
por: Yi, Jung, et al.
Publicado: (2026)
Let 2D Diffusion Model Know 3D-Consistency for Robust Text-to-3D Generation
por: Seo, Junyoung, et al.
Publicado: (2023)
por: Seo, Junyoung, et al.
Publicado: (2023)
Enhancing Creative Generation on Stable Diffusion-based Models
por: Han, Jiyeon, et al.
Publicado: (2025)
por: Han, Jiyeon, et al.
Publicado: (2025)
TimeChara: Evaluating Point-in-Time Character Hallucination of Role-Playing Large Language Models
por: Ahn, Jaewoo, et al.
Publicado: (2024)
por: Ahn, Jaewoo, et al.
Publicado: (2024)
3D Scene Prompting for Scene-Consistent Camera-Controllable Video Generation
por: Lee, JoungBin, et al.
Publicado: (2025)
por: Lee, JoungBin, et al.
Publicado: (2025)
Referring Video Object Segmentation via Language-aligned Track Selection
por: Kim, Seongchan, et al.
Publicado: (2024)
por: Kim, Seongchan, et al.
Publicado: (2024)
Repurposing Geometric Foundation Models for Multi-view Diffusion
por: Jang, Wooseok, et al.
Publicado: (2026)
por: Jang, Wooseok, et al.
Publicado: (2026)
Language-only Efficient Training of Zero-shot Composed Image Retrieval
por: Gu, Geonmo, et al.
Publicado: (2023)
por: Gu, Geonmo, et al.
Publicado: (2023)
DECOR:Decomposition and Projection of Text Embeddings for Text-to-Image Customization
por: Jang, Geonhui, et al.
Publicado: (2024)
por: Jang, Geonhui, et al.
Publicado: (2024)
Learning to See What You Need: Gaze Attention for Multimodal Large Language Models
por: Song, Junha, et al.
Publicado: (2026)
por: Song, Junha, et al.
Publicado: (2026)
DreamMatcher: Appearance Matching Self-Attention for Semantically-Consistent Text-to-Image Personalization
por: Nam, Jisu, et al.
Publicado: (2024)
por: Nam, Jisu, et al.
Publicado: (2024)
StyleKeeper: Prevent Content Leakage using Negative Visual Query Guidance
por: Jeong, Jaeseok, et al.
Publicado: (2025)
por: Jeong, Jaeseok, et al.
Publicado: (2025)
Visual Style Prompting with Swapping Self-Attention
por: Jeong, Jaeseok, et al.
Publicado: (2024)
por: Jeong, Jaeseok, et al.
Publicado: (2024)
Fast KVzip: Efficient and Accurate LLM Inference with Gated KV Eviction
por: Kim, Jang-Hyun, et al.
Publicado: (2026)
por: Kim, Jang-Hyun, et al.
Publicado: (2026)
Dexterous World Models
por: Kim, Byungjun, et al.
Publicado: (2025)
por: Kim, Byungjun, et al.
Publicado: (2025)
MoDiTalker: Motion-Disentangled Diffusion Model for High-Fidelity Talking Head Generation
por: Kim, Seyeon, et al.
Publicado: (2024)
por: Kim, Seyeon, et al.
Publicado: (2024)
MuCo: Multi-turn Contrastive Learning for Multimodal Embedding Model
por: Gu, Geonmo, et al.
Publicado: (2026)
por: Gu, Geonmo, et al.
Publicado: (2026)
Masking meets Supervision: A Strong Learning Alliance
por: Heo, Byeongho, et al.
Publicado: (2023)
por: Heo, Byeongho, et al.
Publicado: (2023)
HYPE: Hyperbolic Entailment Filtering for Underspecified Images and Texts
por: Kim, Wonjae, et al.
Publicado: (2024)
por: Kim, Wonjae, et al.
Publicado: (2024)
CompoDiff: Versatile Composed Image Retrieval With Latent Diffusion
por: Gu, Geonmo, et al.
Publicado: (2023)
por: Gu, Geonmo, et al.
Publicado: (2023)
VG3T: Visual Geometry Grounded Gaussian Transformer
por: Kim, Junho, et al.
Publicado: (2025)
por: Kim, Junho, et al.
Publicado: (2025)
KISS-IMU: Self-supervised Inertial Odometry with Motion-balanced Learning and Uncertainty-aware Inference
por: Choi, Jiwon, et al.
Publicado: (2026)
por: Choi, Jiwon, et al.
Publicado: (2026)
The Effects of Freewriting on L2 Writing Fluency, Emotions, and Perceptions Among Secondary EFL Learners
por: Yi Seul Choi, et al.
Publicado: (2025)
por: Yi Seul Choi, et al.
Publicado: (2025)
Emergent Temporal Correspondences from Video Diffusion Transformers
por: Nam, Jisu, et al.
Publicado: (2025)
por: Nam, Jisu, et al.
Publicado: (2025)
Example-Based Concept Analysis Framework for Deep Weather Forecast Models
por: Kim, Soyeon, et al.
Publicado: (2025)
por: Kim, Soyeon, et al.
Publicado: (2025)
MARS: Matching Attribute-aware Representations for Text-based Sequential Recommendation
por: Kim, Hyunsoo, et al.
Publicado: (2024)
por: Kim, Hyunsoo, et al.
Publicado: (2024)
InterRVOS: Interaction-aware Referring Video Object Segmentation
por: Jin, Woojeong, et al.
Publicado: (2025)
por: Jin, Woojeong, et al.
Publicado: (2025)
CoBA: Counterbias Text Augmentation for Mitigating Various Spurious Correlations via Semantic Triples
por: Jin, Kyohoon, et al.
Publicado: (2025)
por: Jin, Kyohoon, et al.
Publicado: (2025)
Vid-CamEdit: Video Camera Trajectory Editing with Generative Rendering from Estimated Geometry
por: Seo, Junyoung, et al.
Publicado: (2025)
por: Seo, Junyoung, et al.
Publicado: (2025)
Ejemplares similares
-
CAMEO: Correspondence-Attention Alignment for Multi-View Diffusion Models
por: Kwon, Minkyung, et al.
Publicado: (2025) -
WorldCam: Interactive Autoregressive 3D Gaming Worlds with Camera Pose as a Unifying Geometric Representation
por: Nam, Jisu, et al.
Publicado: (2026) -
Aligned Novel View Image and Geometry Synthesis via Cross-modal Attention Instillation
por: Kwak, Min-Seop, et al.
Publicado: (2025) -
MORPHOS: Autoregressive 4D Generation with Temporal Structured Latents
por: Kwon, Minkyung, et al.
Publicado: (2026) -
APPLE: Attribute-Preserving Pseudo-Labeling for Diffusion-Based Face Swapping
por: Kang, Jiwon, et al.
Publicado: (2026)