ViSE: A Systematic Approach to Vision-Only Street-View Extrapolation
Fuente:
arXiv
Saved in:
| Main Authors: | Tan, Kaiyuan, Shen, Yingying, Sun, Haiyang, Wang, Bing, Chen, Guang, Ye, Hangjun |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
ExtraGS: Geometric-Aware Trajectory Extrapolation with Uncertainty-Guided Generative Priors
by: Tan, Kaiyuan, et al.
Published: (2025)
by: Tan, Kaiyuan, et al.
Published: (2025)
UFO: Unifying Feed-Forward and Optimization-based Methods for Large Driving Scene Modeling
by: Tan, Kaiyuan, et al.
Published: (2026)
by: Tan, Kaiyuan, et al.
Published: (2026)
Mirage: One-Step Video Diffusion for Photorealistic and Coherent Asset Editing in Driving Scenes
by: Wang, Shuyun, et al.
Published: (2025)
by: Wang, Shuyun, et al.
Published: (2025)
VGGDrive: Empowering Vision-Language Models with Cross-View Geometric Grounding for Autonomous Driving
by: Wang, Jie, et al.
Published: (2026)
by: Wang, Jie, et al.
Published: (2026)
Pixel-Perfect Visual Geometry Estimation
by: Xu, Gangwei, et al.
Published: (2026)
by: Xu, Gangwei, et al.
Published: (2026)
ParkGaussian: Surround-view 3D Gaussian Splatting for Autonomous Parking
by: Wei, Xiaobao, et al.
Published: (2026)
by: Wei, Xiaobao, et al.
Published: (2026)
Seeing through Satellite Images at Street Views
by: Qian, Ming, et al.
Published: (2025)
by: Qian, Ming, et al.
Published: (2025)
DriveExplorer: Images-Only Decoupled 4D Reconstruction with Progressive Restoration for Driving View Extrapolation
by: Jia, Yuang, et al.
Published: (2025)
by: Jia, Yuang, et al.
Published: (2025)
ViCA: Efficient Multimodal LLMs with Vision-Only Cross-Attention
by: Liu, Wenjie, et al.
Published: (2026)
by: Liu, Wenjie, et al.
Published: (2026)
WorldSplat: Gaussian-Centric Feed-Forward 4D Scene Generation for Autonomous Driving
by: Zhu, Ziyue, et al.
Published: (2025)
by: Zhu, Ziyue, et al.
Published: (2025)
UrbanCraft: Urban View Extrapolation via Hierarchical Sem-Geometric Priors
by: Wang, Tianhang, et al.
Published: (2025)
by: Wang, Tianhang, et al.
Published: (2025)
PointForward: Feedforward Driving Reconstruction through Point-Aligned Representations
by: Chi, Cheng, et al.
Published: (2026)
by: Chi, Cheng, et al.
Published: (2026)
DriveLaW:Unifying Planning and Video Generation in a Latent Driving World
by: Xia, Tianze, et al.
Published: (2025)
by: Xia, Tianze, et al.
Published: (2025)
LVDrive: Latent Visual Representation Enhanced Vision-Language-Action Autonomous Driving Model
by: Mei, Xiaodong, et al.
Published: (2026)
by: Mei, Xiaodong, et al.
Published: (2026)
Cross-View Geo-Localization with Street-View and VHR Satellite Imagery in Decentrality Settings
by: Xia, Panwang, et al.
Published: (2024)
by: Xia, Panwang, et al.
Published: (2024)
LaST-VLA: Thinking in Latent Spatio-Temporal Space for Vision-Language-Action in Autonomous Driving
by: Luo, Yuechen, et al.
Published: (2026)
by: Luo, Yuechen, et al.
Published: (2026)
Genesis: Multimodal Driving Scene Generation with Spatio-Temporal and Cross-Modal Consistency
by: Guo, Xiangyu, et al.
Published: (2025)
by: Guo, Xiangyu, et al.
Published: (2025)
Text2Street: Controllable Text-to-image Generation for Street Views
by: Su, Jinming, et al.
Published: (2024)
by: Su, Jinming, et al.
Published: (2024)
UltraViCo: Breaking Extrapolation Limits in Video Diffusion Transformers
by: Zhao, Min, et al.
Published: (2025)
by: Zhao, Min, et al.
Published: (2025)
Toward Physically Consistent Driving Video World Models under Challenging Trajectories
by: Zhou, Jiawei, et al.
Published: (2026)
by: Zhou, Jiawei, et al.
Published: (2026)
DGGT: Feedforward 4D Reconstruction of Dynamic Driving Scenes using Unposed Images
by: Chen, Xiaoxue, et al.
Published: (2025)
by: Chen, Xiaoxue, et al.
Published: (2025)
PerLDiff: Controllable Street View Synthesis Using Perspective-Layout Diffusion Models
by: Zhang, Jinhua, et al.
Published: (2024)
by: Zhang, Jinhua, et al.
Published: (2024)
UrbanVGGT: Scalable Sidewalk Width Estimation from Street View Images
by: Tan, Kaizhen, et al.
Published: (2026)
by: Tan, Kaizhen, et al.
Published: (2026)
ViTamin: Designing Scalable Vision Models in the Vision-Language Era
by: Chen, Jieneng, et al.
Published: (2024)
by: Chen, Jieneng, et al.
Published: (2024)
StreetCrafter: Street View Synthesis with Controllable Video Diffusion Models
by: Yan, Yunzhi, et al.
Published: (2024)
by: Yan, Yunzhi, et al.
Published: (2024)
CrossViewDiff: A Cross-View Diffusion Model for Satellite-to-Street View Synthesis
by: Li, Weijia, et al.
Published: (2024)
by: Li, Weijia, et al.
Published: (2024)
Recov-Vision: Linking Street View Imagery and Vision-Language Models for Post-Disaster Recovery
by: Xiao, Yiming, et al.
Published: (2025)
by: Xiao, Yiming, et al.
Published: (2025)
Pixel-Perfect Depth with Semantics-Prompted Diffusion Transformers
by: Xu, Gangwei, et al.
Published: (2025)
by: Xu, Gangwei, et al.
Published: (2025)
SGD: Street View Synthesis with Gaussian Splatting and Diffusion Prior
by: Yu, Zhongrui, et al.
Published: (2024)
by: Yu, Zhongrui, et al.
Published: (2024)
ViT-1.58b: Mobile Vision Transformers in the 1-bit Era
by: Yuan, Zhengqing, et al.
Published: (2024)
by: Yuan, Zhengqing, et al.
Published: (2024)
Spatially-Weighted CLIP for Street-View Geo-localization
by: Han, Ting, et al.
Published: (2026)
by: Han, Ting, et al.
Published: (2026)
UniDriveVLA: Unifying Understanding, Perception, and Action Planning for Autonomous Driving
by: Li, Yongkang, et al.
Published: (2026)
by: Li, Yongkang, et al.
Published: (2026)
Semantic-Aware Label Placement for Augmented Reality in Street View
by: Jia, Jianqing, et al.
Published: (2019)
by: Jia, Jianqing, et al.
Published: (2019)
Novel View Extrapolation with Video Diffusion Priors
by: Liu, Kunhao, et al.
Published: (2024)
by: Liu, Kunhao, et al.
Published: (2024)
GeoReasoner: Geo-localization with Reasoning in Street Views using a Large Vision-Language Model
by: Li, Ling, et al.
Published: (2024)
by: Li, Ling, et al.
Published: (2024)
NavAgent: Multi-scale Urban Street View Fusion For UAV Embodied Vision-and-Language Navigation
by: Liu, Youzhi, et al.
Published: (2024)
by: Liu, Youzhi, et al.
Published: (2024)
From Street View to Visual Network: Mapping the Visibility of Urban Landmarks with Vision-Language Models
by: Fan, Zicheng, et al.
Published: (2025)
by: Fan, Zicheng, et al.
Published: (2025)
Optimizing Vision-Language Interactions Through Decoder-Only Models
by: Tanaka, Kaito, et al.
Published: (2024)
by: Tanaka, Kaito, et al.
Published: (2024)
SVIA: A Street View Image Anonymization Framework for Self-Driving Applications
by: Liu, Dongyu, et al.
Published: (2025)
by: Liu, Dongyu, et al.
Published: (2025)
ReCogDrive: A Reinforced Cognitive Framework for End-to-End Autonomous Driving
by: Li, Yongkang, et al.
Published: (2025)
by: Li, Yongkang, et al.
Published: (2025)
Similar Items
-
ExtraGS: Geometric-Aware Trajectory Extrapolation with Uncertainty-Guided Generative Priors
by: Tan, Kaiyuan, et al.
Published: (2025) -
UFO: Unifying Feed-Forward and Optimization-based Methods for Large Driving Scene Modeling
by: Tan, Kaiyuan, et al.
Published: (2026) -
Mirage: One-Step Video Diffusion for Photorealistic and Coherent Asset Editing in Driving Scenes
by: Wang, Shuyun, et al.
Published: (2025) -
VGGDrive: Empowering Vision-Language Models with Cross-View Geometric Grounding for Autonomous Driving
by: Wang, Jie, et al.
Published: (2026) -
Pixel-Perfect Visual Geometry Estimation
by: Xu, Gangwei, et al.
Published: (2026)