CityRAG: Stepping Into a City via Spatially-Grounded Video Generation
Fuente:
arXiv
Saved in:
| Main Authors: | Chou, Gene, Herrmann, Charles, Genova, Kyle, Deng, Boyang, Peng, Songyou, Hariharan, Bharath, Zhang, Jason Y., Snavely, Noah, Henzler, Philipp |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Visual Chronicles: Using Multimodal LLMs to Analyze Massive Collections of Images
by: Deng, Boyang, et al.
Published: (2025)
by: Deng, Boyang, et al.
Published: (2025)
FlashDepth: Real-time Streaming Video Depth Estimation at 2K Resolution
by: Chou, Gene, et al.
Published: (2025)
by: Chou, Gene, et al.
Published: (2025)
KFC-W: Generating 3D-Consistent Videos from Unposed Internet Photos
by: Chou, Gene, et al.
Published: (2024)
by: Chou, Gene, et al.
Published: (2024)
Can Generative Video Models Help Pose Estimation?
by: Cai, Ruojin, et al.
Published: (2024)
by: Cai, Ruojin, et al.
Published: (2024)
Learning Feature Descriptors using Camera Pose Supervision
by: Wang, Qianqian, et al.
Published: (2020)
by: Wang, Qianqian, et al.
Published: (2020)
MegaScenes: Scene-Level View Synthesis at Scale
by: Tung, Joseph, et al.
Published: (2024)
by: Tung, Joseph, et al.
Published: (2024)
C3Po: Cross-View Cross-Modality Correspondence by Pointmap Prediction
by: Huang, Kuan Wei, et al.
Published: (2025)
by: Huang, Kuan Wei, et al.
Published: (2025)
UFO-4D: Unposed Feedforward 4D Reconstruction from Two Images
by: Hur, Junhwa, et al.
Published: (2026)
by: Hur, Junhwa, et al.
Published: (2026)
MoMaps: Semantics-Aware Scene Motion Generation with Motion Maps
by: Lei, Jiahui, et al.
Published: (2025)
by: Lei, Jiahui, et al.
Published: (2025)
G3T Up! Gravity Aligned Coordinate Frames Simplify Pointmap Processing
by: Kani, Bharath Raj Nagoor, et al.
Published: (2026)
by: Kani, Bharath Raj Nagoor, et al.
Published: (2026)
SplatTalk: 3D VQA with Gaussian Splatting
by: Thai, Anh, et al.
Published: (2025)
by: Thai, Anh, et al.
Published: (2025)
Streetscapes: Large-scale Consistent Street View Generation Using Autoregressive Video Diffusion
by: Deng, Boyang, et al.
Published: (2024)
by: Deng, Boyang, et al.
Published: (2024)
ObjectCarver: Semi-automatic segmentation, reconstruction and separation of 3D objects
by: Hassena, Gemmechu, et al.
Published: (2024)
by: Hassena, Gemmechu, et al.
Published: (2024)
MOD-UV: Learning Mobile Object Detectors from Unlabeled Videos
by: Sun, Yihong, et al.
Published: (2024)
by: Sun, Yihong, et al.
Published: (2024)
Flat-Pack Bench: Evaluating Spatio-Temporal Understanding in Large Vision-Language Models through Furniture Assembly
by: Chetan, Aditya, et al.
Published: (2026)
by: Chetan, Aditya, et al.
Published: (2026)
GeoFlow: Enforcing Implicit Geometric Consistency in Video Generation
by: Ackermann, Jan, et al.
Published: (2026)
by: Ackermann, Jan, et al.
Published: (2026)
VLIC: Vision-Language Models As Perceptual Judges for Human-Aligned Image Compression
by: Sargent, Kyle, et al.
Published: (2025)
by: Sargent, Kyle, et al.
Published: (2025)
ArchSym: Detecting 3D-Grounded Architectural Symmetries in the Wild
by: Chen, Hanyu, et al.
Published: (2026)
by: Chen, Hanyu, et al.
Published: (2026)
Composing People Together: Iterative Pose-Image Generation for Multi-Person Interaction Scenes
by: Peng, Wenxuan, et al.
Published: (2026)
by: Peng, Wenxuan, et al.
Published: (2026)
Counter-Current Learning: A Biologically Plausible Dual Network Approach for Deep Learning
by: Kao, Chia-Hsiang, et al.
Published: (2024)
by: Kao, Chia-Hsiang, et al.
Published: (2024)
Four Steeples over the City Streets
by: Bulthuis, Kyle T.
Published: (2024)
by: Bulthuis, Kyle T.
Published: (2024)
Beyond the Frame: Generating 360 Panoramic Videos from Perspective Videos
by: Luo, Rundong, et al.
Published: (2025)
by: Luo, Rundong, et al.
Published: (2025)
The new era of Eurocapitalism
by: Henzler, Herbert
Published: (1992)
by: Henzler, Herbert
Published: (1992)
When the City Teaches the Car: Label-Free 3D Perception from Infrastructure
by: Xu, Zhen, et al.
Published: (2026)
by: Xu, Zhen, et al.
Published: (2026)
Wide-Baseline Relative Camera Pose Estimation with Directional Learning
by: Chen, Kefan, et al.
Published: (2021)
by: Chen, Kefan, et al.
Published: (2021)
Reinforcement Learning Integrated Agentic RAG for Software Test Cases Authoring
by: Hariharan, Mohanakrishnan
Published: (2025)
by: Hariharan, Mohanakrishnan
Published: (2025)
Live Interactive Training for Video Segmentation
by: Yang, Xinyu, et al.
Published: (2026)
by: Yang, Xinyu, et al.
Published: (2026)
MovieRecapsQA: A Multimodal Open-Ended Video Question-Answering Benchmark
by: Shaar, Shaden, et al.
Published: (2026)
by: Shaar, Shaden, et al.
Published: (2026)
Amman City, Jordan: Toward a Sustainable City from the Ground Up
by: Al-Msie'deen, Ra'Fat
Published: (2024)
by: Al-Msie'deen, Ra'Fat
Published: (2024)
ShadowDraw: From Any Object to Shadow-Drawing Compositional Art
by: Luo, Rundong, et al.
Published: (2025)
by: Luo, Rundong, et al.
Published: (2025)
Splat4D: Diffusion-Enhanced 4D Gaussian Splatting for Temporally and Spatially Consistent Content Creation
by: Yin, Minghao, et al.
Published: (2025)
by: Yin, Minghao, et al.
Published: (2025)
Urban Settings in the Competition among Cities
by: Philipp Klaus
Published: (2004)
by: Philipp Klaus
Published: (2004)
NeRF On-the-go: Exploiting Uncertainty for Distractor-free NeRFs in the Wild
by: Ren, Weining, et al.
Published: (2024)
by: Ren, Weining, et al.
Published: (2024)
Setting up a Library BBS: A Step-by-Step Guide.
by: Williams, Gene, et al.
Published: (1987)
by: Williams, Gene, et al.
Published: (1987)
Color Bind: Exploring Color Perception in Text-to-Image Models
by: Shomer-Chai, Shay, et al.
Published: (2025)
by: Shomer-Chai, Shay, et al.
Published: (2025)
WonderJourney: Going from Anywhere to Everywhere
by: Yu, Hong-Xing, et al.
Published: (2023)
by: Yu, Hong-Xing, et al.
Published: (2023)
Chapter Characterization of Atmospheric Mercury in the High-Altitude Background Station and Coastal Urban City in South Asia
by: R, Karthik, et al.
Published: (2024)
by: R, Karthik, et al.
Published: (2024)
Chapter Characterization of Atmospheric Mercury in the High-Altitude Background Station and Coastal Urban City in South Asia
by: Bharath, Karuppasamy, et al.
Published: (2021)
by: Bharath, Karuppasamy, et al.
Published: (2021)
CityGuessr: City-Level Video Geo-Localization on a Global Scale
by: Kulkarni, Parth Parag, et al.
Published: (2024)
by: Kulkarni, Parth Parag, et al.
Published: (2024)
Video, Cable Television and Public Libraries
by: Genova, B. K. L.
Published: (1978)
by: Genova, B. K. L.
Published: (1978)
Similar Items
-
Visual Chronicles: Using Multimodal LLMs to Analyze Massive Collections of Images
by: Deng, Boyang, et al.
Published: (2025) -
FlashDepth: Real-time Streaming Video Depth Estimation at 2K Resolution
by: Chou, Gene, et al.
Published: (2025) -
KFC-W: Generating 3D-Consistent Videos from Unposed Internet Photos
by: Chou, Gene, et al.
Published: (2024) -
Can Generative Video Models Help Pose Estimation?
by: Cai, Ruojin, et al.
Published: (2024) -
Learning Feature Descriptors using Camera Pose Supervision
by: Wang, Qianqian, et al.
Published: (2020)