UrbanVGGT: Scalable Sidewalk Width Estimation from Street View Images
Fuente:
arXiv
Saved in:
| Main Authors: | Tan, Kaizhen, Zhang, Fan |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Does Visual Token Pruning Improve Calibration? An Empirical Study on Confidence in MLLMs
by: Tan, Kaizhen
Published: (2026)
by: Tan, Kaizhen
Published: (2026)
VGGT-X: When VGGT Meets Dense Novel View Synthesis
by: Liu, Yang, et al.
Published: (2025)
by: Liu, Yang, et al.
Published: (2025)
From Street View to Visual Network: Mapping the Visibility of Urban Landmarks with Vision-Language Models
by: Fan, Zicheng, et al.
Published: (2025)
by: Fan, Zicheng, et al.
Published: (2025)
SegVGGT: Joint 3D Reconstruction and Instance Segmentation from Multi-View Images
by: Qu, Jinyuan, et al.
Published: (2026)
by: Qu, Jinyuan, et al.
Published: (2026)
Seeing through Satellite Images at Street Views
by: Qian, Ming, et al.
Published: (2025)
by: Qian, Ming, et al.
Published: (2025)
StreetSurfGS: Scalable Urban Street Surface Reconstruction with Planar-based Gaussian Splatting
by: Cui, Xiao, et al.
Published: (2024)
by: Cui, Xiao, et al.
Published: (2024)
VGGT-Segmentor: Geometry-Enhanced Cross-View Segmentation
by: Gao, Yulu, et al.
Published: (2026)
by: Gao, Yulu, et al.
Published: (2026)
Modeling Urban Food Insecurity with Google Street View Images
by: Li, David
Published: (2025)
by: Li, David
Published: (2025)
ZenSVI: An Open-Source Software for the Integrated Acquisition, Processing and Analysis of Street View Imagery Towards Scalable Urban Science
by: Ito, Koichi, et al.
Published: (2024)
by: Ito, Koichi, et al.
Published: (2024)
Urban Safety Perception Assessments via Integrating Multimodal Large Language Models with Street View Images
by: Zhang, Jiaxin, et al.
Published: (2024)
by: Zhang, Jiaxin, et al.
Published: (2024)
Eyes on the Streets: Leveraging Street-Level Imaging to Model Urban Crime Dynamics
by: Qi, Zhixuan, et al.
Published: (2024)
by: Qi, Zhixuan, et al.
Published: (2024)
VGGT-Det: Mining VGGT Internal Priors for Sensor-Geometry-Free Multi-View Indoor 3D Object Detection
by: Cao, Yang, et al.
Published: (2026)
by: Cao, Yang, et al.
Published: (2026)
ELEV-VISION: Automated Lowest Floor Elevation Estimation from Segmenting Street View Images
by: Ho, Yu-Hsuan, et al.
Published: (2023)
by: Ho, Yu-Hsuan, et al.
Published: (2023)
VGGT-World: Transforming VGGT into an Autoregressive Geometry World Model
by: Sun, Xiangyu, et al.
Published: (2026)
by: Sun, Xiangyu, et al.
Published: (2026)
VGGT-$Ω$
by: Wang, Jianyuan, et al.
Published: (2026)
by: Wang, Jianyuan, et al.
Published: (2026)
Decoding Tourist Perception in Historic Urban Quarters with Multimodal Social Media Data: An AI-Based Framework and Evidence from Shanghai
by: Tan, Kaizhen, et al.
Published: (2025)
by: Tan, Kaizhen, et al.
Published: (2025)
Street-View Image Generation from a Bird's-Eye View Layout
by: Swerdlow, Alexander, et al.
Published: (2023)
by: Swerdlow, Alexander, et al.
Published: (2023)
StreetView-Waste: A Multi-Task Dataset for Urban Waste Management
by: Paulo, Diogo J., et al.
Published: (2025)
by: Paulo, Diogo J., et al.
Published: (2025)
CREG: Compass Relational Evidence Graph for Characterizing Directional Structure in VLM Spatial-Reasoning Attribution
by: Tan, Kaizhen, et al.
Published: (2026)
by: Tan, Kaizhen, et al.
Published: (2026)
Semantic-Aware Label Placement for Augmented Reality in Street View
by: Jia, Jianqing, et al.
Published: (2019)
by: Jia, Jianqing, et al.
Published: (2019)
CityPulse: Fine-Grained Assessment of Urban Change with Street View Time Series
by: Huang, Tianyuan, et al.
Published: (2024)
by: Huang, Tianyuan, et al.
Published: (2024)
Multimodal Deep Learning for ATCO Command Lifecycle Modeling and Workload Prediction
by: Tan, Kaizhen
Published: (2025)
by: Tan, Kaizhen
Published: (2025)
S-VGGT: Structure-Aware Subscene Decomposition for Scalable 3D Foundation Models
by: Li, Xinze, et al.
Published: (2026)
by: Li, Xinze, et al.
Published: (2026)
FlashVGGT: Efficient and Scalable Visual Geometry Transformers with Compressed Descriptor Attention
by: Wang, Zipeng, et al.
Published: (2025)
by: Wang, Zipeng, et al.
Published: (2025)
Learning Street View Representations with Spatiotemporal Contrast
by: Li, Yong, et al.
Published: (2025)
by: Li, Yong, et al.
Published: (2025)
Precise and Robust Sidewalk Detection: Leveraging Ensemble Learning to Surpass LLM Limitations in Urban Environments
by: Shihab, Ibne Farabi, et al.
Published: (2024)
by: Shihab, Ibne Farabi, et al.
Published: (2024)
VGGT-MPR: VGGT-Enhanced Multimodal Place Recognition in Autonomous Driving Environments
by: Xu, Jingyi, et al.
Published: (2026)
by: Xu, Jingyi, et al.
Published: (2026)
FrameVGGT: Geometry-Aligned Frame-Level Memory for Bounded Streaming VGGT
by: Xu, Zhisong, et al.
Published: (2026)
by: Xu, Zhisong, et al.
Published: (2026)
VGGT-SLAM++
by: Mandal, Avilasha, et al.
Published: (2026)
by: Mandal, Avilasha, et al.
Published: (2026)
VGGT-HPE: Reframing Head Pose Estimation as Relative Pose Prediction
by: Vasileiou, Vasiliki, et al.
Published: (2026)
by: Vasileiou, Vasiliki, et al.
Published: (2026)
VGGT-360: Geometry-Consistent Zero-Shot Panoramic Depth Estimation
by: Yuan, Jiayi, et al.
Published: (2026)
by: Yuan, Jiayi, et al.
Published: (2026)
LiteVGGT: Boosting Vanilla VGGT via Geometry-aware Cached Token Merging
by: Shu, Zhijian, et al.
Published: (2025)
by: Shu, Zhijian, et al.
Published: (2025)
StreetCrafter: Street View Synthesis with Controllable Video Diffusion Models
by: Yan, Yunzhi, et al.
Published: (2024)
by: Yan, Yunzhi, et al.
Published: (2024)
Text2Street: Controllable Text-to-image Generation for Street Views
by: Su, Jinming, et al.
Published: (2024)
by: Su, Jinming, et al.
Published: (2024)
EndoVGGT: GNN-Enhanced Depth Estimation for Surgical 3D Reconstruction
by: Fan, Falong, et al.
Published: (2026)
by: Fan, Falong, et al.
Published: (2026)
AVGGT: Rethinking Global Attention for Accelerating VGGT
by: Sun, Xianbing, et al.
Published: (2025)
by: Sun, Xianbing, et al.
Published: (2025)
SVIA: A Street View Image Anonymization Framework for Self-Driving Applications
by: Liu, Dongyu, et al.
Published: (2025)
by: Liu, Dongyu, et al.
Published: (2025)
ViSE: A Systematic Approach to Vision-Only Street-View Extrapolation
by: Tan, Kaiyuan, et al.
Published: (2025)
by: Tan, Kaiyuan, et al.
Published: (2025)
VGGT-Long: Chunk it, Loop it, Align it -- Pushing VGGT's Limits on Kilometer-scale Long RGB Sequences
by: Deng, Kai, et al.
Published: (2025)
by: Deng, Kai, et al.
Published: (2025)
VGGT: Visual Geometry Grounded Transformer
by: Wang, Jianyuan, et al.
Published: (2025)
by: Wang, Jianyuan, et al.
Published: (2025)
Similar Items
-
Does Visual Token Pruning Improve Calibration? An Empirical Study on Confidence in MLLMs
by: Tan, Kaizhen
Published: (2026) -
VGGT-X: When VGGT Meets Dense Novel View Synthesis
by: Liu, Yang, et al.
Published: (2025) -
From Street View to Visual Network: Mapping the Visibility of Urban Landmarks with Vision-Language Models
by: Fan, Zicheng, et al.
Published: (2025) -
SegVGGT: Joint 3D Reconstruction and Instance Segmentation from Multi-View Images
by: Qu, Jinyuan, et al.
Published: (2026) -
Seeing through Satellite Images at Street Views
by: Qian, Ming, et al.
Published: (2025)