Built Environment Reasoning from Remote Sensing Imagery Using Large Vision--Language Models
Fuente:
arXiv
Saved in:
| Main Authors: | Wang, Dongdong, Balakrishnan, Deepak, Srinivasan, Ravi, Wang, Shenhao |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
X-Driver: Explainable Autonomous Driving with Vision-Language Models
by: Liu, Wei, et al.
Published: (2025)
by: Liu, Wei, et al.
Published: (2025)
TimeSpot: Benchmarking Geo-Temporal Understanding in Vision-Language Models in Real-World Settings
by: Wasi, Azmine Toushik, et al.
Published: (2026)
by: Wasi, Azmine Toushik, et al.
Published: (2026)
Extracting Object Heights From LiDAR & Aerial Imagery
by: Guerrero, Jesus
Published: (2024)
by: Guerrero, Jesus
Published: (2024)
VIPER Strike: Defeating Visual Reasoning CAPTCHAs via Structured Vision-Language Inference
by: Qi, Minfeng, et al.
Published: (2026)
by: Qi, Minfeng, et al.
Published: (2026)
Privacy of Groups in Dense Street Imagery
by: Franchi, Matt, et al.
Published: (2025)
by: Franchi, Matt, et al.
Published: (2025)
On Circuit-based Hybrid Quantum Neural Networks for Remote Sensing Imagery Classification
by: Sebastianelli, Alessandro, et al.
Published: (2021)
by: Sebastianelli, Alessandro, et al.
Published: (2021)
EATFormer: Improving Vision Transformer Inspired by Evolutionary Algorithm
by: Zhang, Jiangning, et al.
Published: (2022)
by: Zhang, Jiangning, et al.
Published: (2022)
CMAP: Cross-Modal Adaptive Prompting for Multi-Domain Task-Incremental Learning
by: Mandalika, Sriram
Published: (2026)
by: Mandalika, Sriram
Published: (2026)
Can Foundation Models Revolutionize Mobile AR Sparse Sensing?
by: Zhao, Yiqin, et al.
Published: (2025)
by: Zhao, Yiqin, et al.
Published: (2025)
VOLMO: Versatile and Open Large Models for Ophthalmology
by: Qin, Zhenyue, et al.
Published: (2026)
by: Qin, Zhenyue, et al.
Published: (2026)
PI-HMR: Towards Robust In-bed Temporal Human Shape Reconstruction with Contact Pressure Sensing
by: Wu, Ziyu, et al.
Published: (2025)
by: Wu, Ziyu, et al.
Published: (2025)
A Comprehensive Content Verification System for ensuring Digital Integrity in the Age of Deep Fakes
by: Kaja, RaviKanth
Published: (2024)
by: Kaja, RaviKanth
Published: (2024)
Self-evolving Embodied AI
by: Feng, Tongtong, et al.
Published: (2026)
by: Feng, Tongtong, et al.
Published: (2026)
Human-in-the-Loop: Quantitative Evaluation of 3D Models Generation by Large Language Models
by: Sadik, Ahmed R., et al.
Published: (2025)
by: Sadik, Ahmed R., et al.
Published: (2025)
Assistive Image Annotation Systems with Deep Learning and Natural Language Capabilities: A Review
by: Mots'oehli, Moseli
Published: (2024)
by: Mots'oehli, Moseli
Published: (2024)
Empirical Studies of Large Scale Environment Scanning by Consumer Electronics
by: Wang, Mengyuan, et al.
Published: (2025)
by: Wang, Mengyuan, et al.
Published: (2025)
Automated Facility Enumeration for Building Compliance Checking using Door Detection and Large Language Models
by: Zhang, Licheng, et al.
Published: (2025)
by: Zhang, Licheng, et al.
Published: (2025)
Spatio-Temporal Foundation Models: Vision, Challenges, and Opportunities
by: Goodge, Adam, et al.
Published: (2025)
by: Goodge, Adam, et al.
Published: (2025)
From Pixels to Nucleotides: End-to-End Token-Based Video Compression for DNA Storage
by: Ruan, Cihan, et al.
Published: (2026)
by: Ruan, Cihan, et al.
Published: (2026)
DoorDet: Semi-Automated Multi-Class Door Detection Dataset via Object Detection and Large Language Models
by: Zhang, Licheng, et al.
Published: (2025)
by: Zhang, Licheng, et al.
Published: (2025)
EgoPoseVR: Spatiotemporal Multi-Modal Reasoning for Egocentric Full-Body Pose in Virtual Reality
by: Cheng, Haojie, et al.
Published: (2026)
by: Cheng, Haojie, et al.
Published: (2026)
Pedestrian Intention Prediction via Vision-Language Foundation Models
by: Azarmi, Mohsen, et al.
Published: (2025)
by: Azarmi, Mohsen, et al.
Published: (2025)
Multi-Image Super Resolution Framework for Detection and Analysis of Plant Roots
by: Agarwal, Shubham, et al.
Published: (2026)
by: Agarwal, Shubham, et al.
Published: (2026)
Evaluating and Enhancing Trustworthiness of LLMs in Perception Tasks
by: Dona, Malsha Ashani Mahawatta, et al.
Published: (2024)
by: Dona, Malsha Ashani Mahawatta, et al.
Published: (2024)
Learned Display Radiance Fields with Lensless Cameras
by: Chen, Ziyang, et al.
Published: (2025)
by: Chen, Ziyang, et al.
Published: (2025)
Introducing Nylon Face Mask Attacks: A Dataset for Evaluating Generalised Face Presentation Attack Detection
by: Manasa, et al.
Published: (2025)
by: Manasa, et al.
Published: (2025)
SynSpill: Improved Industrial Spill Detection With Synthetic Data
by: Baranwal, Aaditya, et al.
Published: (2025)
by: Baranwal, Aaditya, et al.
Published: (2025)
Scrutinizing Data from Sky: An Examination of Its Veracity in Area Based Traffic Contexts
by: Ali, Yawar, et al.
Published: (2024)
by: Ali, Yawar, et al.
Published: (2024)
Fourier-based Action Recognition for Wildlife Behavior Quantification with Event Cameras
by: Hamann, Friedhelm, et al.
Published: (2024)
by: Hamann, Friedhelm, et al.
Published: (2024)
x-RAGE: eXtended Reality -- Action & Gesture Events Dataset
by: Parmar, Vivek, et al.
Published: (2024)
by: Parmar, Vivek, et al.
Published: (2024)
Towards Railway Domain Adaptation for LiDAR-based 3D Detection: Road-to-Rail and Sim-to-Real via SynDRA-BBox
by: Diaz, Xavier, et al.
Published: (2025)
by: Diaz, Xavier, et al.
Published: (2025)
Fast Quantum Convolutional Neural Networks for Low-Complexity Object Detection in Autonomous Driving Applications
by: Baek, Hankyul, et al.
Published: (2023)
by: Baek, Hankyul, et al.
Published: (2023)
Enhancing Autism Spectrum Disorder Early Detection with the Parent-Child Dyads Block-Play Protocol and an Attention-enhanced GCN-xLSTM Hybrid Deep Learning Framework
by: Li, Xiang, et al.
Published: (2024)
by: Li, Xiang, et al.
Published: (2024)
DashCam Video: A complementary low-cost data stream for on-demand forest-infrastructure system monitoring
by: Joshi, Durga, et al.
Published: (2025)
by: Joshi, Durga, et al.
Published: (2025)
Attention-based Generative Latent Replay: A Continual Learning Approach for WSI Analysis
by: Kumari, Pratibha, et al.
Published: (2025)
by: Kumari, Pratibha, et al.
Published: (2025)
Unlocking Comics: The AI4VA Dataset for Visual Understanding
by: Grönquist, Peter, et al.
Published: (2024)
by: Grönquist, Peter, et al.
Published: (2024)
SpecTrack: Learned Multi-Rotation Tracking via Speckle Imaging
by: Chen, Ziyang, et al.
Published: (2024)
by: Chen, Ziyang, et al.
Published: (2024)
Diff-GNSS: Diffusion-based Pseudorange Error Estimation
by: Zhu, Jiaqi, et al.
Published: (2025)
by: Zhu, Jiaqi, et al.
Published: (2025)
Probabilistic Online Event Downsampling
by: Girbau-Xalabarder, Andreu, et al.
Published: (2025)
by: Girbau-Xalabarder, Andreu, et al.
Published: (2025)
INSIGHT: Indoor Scene Intelligence from Geometric-Semantic Hierarchy Transfer for Public~Safety
by: Dimopoulos, Alexander Nikitas, et al.
Published: (2026)
by: Dimopoulos, Alexander Nikitas, et al.
Published: (2026)
Similar Items
-
X-Driver: Explainable Autonomous Driving with Vision-Language Models
by: Liu, Wei, et al.
Published: (2025) -
TimeSpot: Benchmarking Geo-Temporal Understanding in Vision-Language Models in Real-World Settings
by: Wasi, Azmine Toushik, et al.
Published: (2026) -
Extracting Object Heights From LiDAR & Aerial Imagery
by: Guerrero, Jesus
Published: (2024) -
VIPER Strike: Defeating Visual Reasoning CAPTCHAs via Structured Vision-Language Inference
by: Qi, Minfeng, et al.
Published: (2026) -
Privacy of Groups in Dense Street Imagery
by: Franchi, Matt, et al.
Published: (2025)