Streetscape Analysis with Generative AI (SAGAI): Vision-Language Assessment and Mapping of Urban Scenes
Fuente:
arXiv
Saved in:
| Main Authors: | Perez, Joan, Fusco, Giovanni |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
UVLM: A Universal Vision-Language Model Loader for Reproducible Multimodal Benchmarking
by: Perez, Joan, et al.
Published: (2026)
by: Perez, Joan, et al.
Published: (2026)
Predicting household socioeconomic position in Mozambique using satellite and household imagery
by: Milà, Carles, et al.
Published: (2024)
by: Milà, Carles, et al.
Published: (2024)
FOCUS on Contamination: Hydrology-Informed Noise-Aware Learning for Geospatial PFAS Mapping
by: Khan, Jowaria, et al.
Published: (2025)
by: Khan, Jowaria, et al.
Published: (2025)
VLSlice: Interactive Vision-and-Language Slice Discovery
by: Slyman, Eric, et al.
Published: (2023)
by: Slyman, Eric, et al.
Published: (2023)
Massively Multi-Person 3D Human Motion Forecasting with Scene Context
by: Mueller, Felix B, et al.
Published: (2024)
by: Mueller, Felix B, et al.
Published: (2024)
AI-Augmented Pollen Recognition in Optical and Holographic Microscopy for Veterinary Imaging
by: Warshaneyan, Swarn S., et al.
Published: (2025)
by: Warshaneyan, Swarn S., et al.
Published: (2025)
RAM-H1200: A Unified Evaluation and Dataset on Hand Radiographs for Rheumatoid Arthritis
by: Yang, Songxiao, et al.
Published: (2026)
by: Yang, Songxiao, et al.
Published: (2026)
Deep Learning with Self-Attention and Enhanced Preprocessing for Precise Diagnosis of Acute Lymphoblastic Leukemia from Bone Marrow Smears in Hemato-Oncology
by: Maruf, Md., et al.
Published: (2025)
by: Maruf, Md., et al.
Published: (2025)
DIsoN: Decentralized Isolation Networks for Out-of-Distribution Detection in Medical Imaging
by: Wagner, Felix, et al.
Published: (2025)
by: Wagner, Felix, et al.
Published: (2025)
Revealing an Unattractivity Bias in Mental Reconstruction of Occluded Faces using Generative Image Models
by: Riedmann, Frederik, et al.
Published: (2024)
by: Riedmann, Frederik, et al.
Published: (2024)
A Review of Pseudo-Labeling for Computer Vision
by: Kage, Patrick, et al.
Published: (2024)
by: Kage, Patrick, et al.
Published: (2024)
Do All Vision Transformers Need Registers? A Cross-Architectural Reassessment
by: Baxevanakis, Spiros, et al.
Published: (2026)
by: Baxevanakis, Spiros, et al.
Published: (2026)
Which Backbone to Use: A Resource-efficient Domain Specific Comparison for Computer Vision
by: Jeevan, Pranav, et al.
Published: (2024)
by: Jeevan, Pranav, et al.
Published: (2024)
PyCAT4: A Hierarchical Vision Transformer-based Framework for 3D Human Pose Estimation
by: Yang, Zongyou, et al.
Published: (2025)
by: Yang, Zongyou, et al.
Published: (2025)
Lost in Translation: How Language Re-Aligns Vision for Cross-Species Pathology
by: Arora, Ekansh
Published: (2026)
by: Arora, Ekansh
Published: (2026)
LeDiFlow: Learned Distribution-guided Flow Matching to Accelerate Image Generation
by: Zwick, Pascal, et al.
Published: (2025)
by: Zwick, Pascal, et al.
Published: (2025)
Automated Pollen Recognition in Optical and Holographic Microscopy Images
by: Warshaneyan, Swarn Singh, et al.
Published: (2025)
by: Warshaneyan, Swarn Singh, et al.
Published: (2025)
WaveMix: A Resource-efficient Neural Network for Image Analysis
by: Jeevan, Pranav, et al.
Published: (2022)
by: Jeevan, Pranav, et al.
Published: (2022)
Fine-Tuning Vision-Language Models for Understanding Current Damage and Scoring Priority with Quality Guard Agent
by: Yasuno, Takato
Published: (2026)
by: Yasuno, Takato
Published: (2026)
Explainable Classifier for Malignant Lymphoma Subtyping via Cell Graph and Image Fusion
by: Nishiyama, Daiki, et al.
Published: (2025)
by: Nishiyama, Daiki, et al.
Published: (2025)
High-Entropy Tokens as Multimodal Failure Points in Vision-Language Models
by: He, Mengqi, et al.
Published: (2025)
by: He, Mengqi, et al.
Published: (2025)
Decoding the Surgical Scene: A Scoping Review of Scene Graphs in Surgery
by: Henriques, Angelo, et al.
Published: (2025)
by: Henriques, Angelo, et al.
Published: (2025)
Automated Deep Learning Estimation of Anthropometric Measurements for Preparticipation Cardiovascular Screening
by: Mareque, Lucas R., et al.
Published: (2025)
by: Mareque, Lucas R., et al.
Published: (2025)
Seeing Roads Through Words: A Language-Guided Framework for RGB-T Driving Scene Segmentation
by: Reddy, Ruturaj, et al.
Published: (2026)
by: Reddy, Ruturaj, et al.
Published: (2026)
CC-SGG: Corner Case Scenario Generation using Learned Scene Graphs
by: Drayson, George, et al.
Published: (2023)
by: Drayson, George, et al.
Published: (2023)
Disentangling Generation and Regression in Stochastic Interpolants for Controllable Image Restoration
by: Liu, Yi, et al.
Published: (2026)
by: Liu, Yi, et al.
Published: (2026)
CytoDiff: AI-Driven Cytomorphology Image Synthesis for Medical Diagnostics
by: Boada, Jan Carreras, et al.
Published: (2025)
by: Boada, Jan Carreras, et al.
Published: (2025)
Self-Supervised Polyp Re-Identification in Colonoscopy
by: Intrator, Yotam, et al.
Published: (2023)
by: Intrator, Yotam, et al.
Published: (2023)
Unpaired Cross-Domain Calibration of DMSP to VIIRS Nighttime Light Data Based on CUT Network
by: Tong, Zhan, et al.
Published: (2026)
by: Tong, Zhan, et al.
Published: (2026)
Unpacking the Eye of the Beholder: Social Location, Identity, and the Moving Target of Political Perspectives
by: Sirotkina, Elena
Published: (2026)
by: Sirotkina, Elena
Published: (2026)
Towards Accurate and Efficient Waste Image Classification: A Hybrid Deep Learning and Machine Learning Approach
by: Nguyen, Ngoc-Bao-Quang, et al.
Published: (2025)
by: Nguyen, Ngoc-Bao-Quang, et al.
Published: (2025)
DIET-CP: Lightweight and Data Efficient Self Supervised Continued Pretraining
by: Rodas, Bryan, et al.
Published: (2025)
by: Rodas, Bryan, et al.
Published: (2025)
Quantized Vision-Language Models for Damage Assessment: A Comparative Study of LLaVA-1.5-7B Quantization Levels
by: Yasuno, Takato
Published: (2026)
by: Yasuno, Takato
Published: (2026)
PBSBench: A Multi-Level Vision-Language Framework and Benchmark for Hematopathology Whole Slide Image Interpretation
by: Wang, Yuanlong, et al.
Published: (2026)
by: Wang, Yuanlong, et al.
Published: (2026)
From Misclassifications to Outliers: Joint Reliability Assessment in Classification
by: Li, Yang, et al.
Published: (2026)
by: Li, Yang, et al.
Published: (2026)
Predictive Quality Assessment for Mobile Secure Graphics
by: Steigstra, Cas, et al.
Published: (2025)
by: Steigstra, Cas, et al.
Published: (2025)
Beyond Blur: A Fluid Perspective on Generative Diffusion Models
by: Gruszczynski, Grzegorz, et al.
Published: (2025)
by: Gruszczynski, Grzegorz, et al.
Published: (2025)
Deep Self-Supervised Disturbance Mapping with the OPERA Sentinel-1 Radiometric Terrain Corrected SAR Backscatter Product
by: Hardiman-Mostow, Harris, et al.
Published: (2025)
by: Hardiman-Mostow, Harris, et al.
Published: (2025)
Grounding Synthetic Data Generation With Vision and Language Models
by: Çağlar, Ümit Mert, et al.
Published: (2026)
by: Çağlar, Ümit Mert, et al.
Published: (2026)
The Uncanny Valley: A Comprehensive Analysis of Diffusion Models
by: Ghanem, Karam, et al.
Published: (2024)
by: Ghanem, Karam, et al.
Published: (2024)
Similar Items
-
UVLM: A Universal Vision-Language Model Loader for Reproducible Multimodal Benchmarking
by: Perez, Joan, et al.
Published: (2026) -
Predicting household socioeconomic position in Mozambique using satellite and household imagery
by: Milà, Carles, et al.
Published: (2024) -
FOCUS on Contamination: Hydrology-Informed Noise-Aware Learning for Geospatial PFAS Mapping
by: Khan, Jowaria, et al.
Published: (2025) -
VLSlice: Interactive Vision-and-Language Slice Discovery
by: Slyman, Eric, et al.
Published: (2023) -
Massively Multi-Person 3D Human Motion Forecasting with Scene Context
by: Mueller, Felix B, et al.
Published: (2024)