Scene Change Detection with Vision-Language Representation Learning
Fuente:
arXiv
Saved in:
| Main Authors: | Sheng, Diwei, Gohil, Vijayraj, Gaba, Satyam, Liu, Zihan, Hamilton-Fletcher, Giles, Rizzo, John-Ross, Liang, Yongqing, Feng, Chen |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Can Foundation Models Reliably Identify Spatial Hazards? A Case Study on Curb Segmentation
by: Sheng, Diwei, et al.
Published: (2024)
by: Sheng, Diwei, et al.
Published: (2024)
Improving Long-Tailed Object Detection with Balanced Group Softmax and Metric Learning
by: Gaba, Satyam
Published: (2025)
by: Gaba, Satyam
Published: (2025)
Generative AI for Enhanced Wildfire Detection: Bridging the Synthetic-Real Domain Gap
by: Gaba, Satyam
Published: (2025)
by: Gaba, Satyam
Published: (2025)
Robust Computer-Vision based Construction Site Detection for Assistive-Technology Applications
by: Feng, Junchi, et al.
Published: (2025)
by: Feng, Junchi, et al.
Published: (2025)
NYC-Indoor-VPR: A Long-Term Indoor Visual Place Recognition Dataset with Semi-Automatic Annotation
by: Sheng, Diwei, et al.
Published: (2024)
by: Sheng, Diwei, et al.
Published: (2024)
Does Embodiment Matter to Biomechanics and Function? A Comparative Analysis of Head-Mounted and Hand-Held Assistive Devices for Individuals with Blindness and Low Vision
by: Seth, Gaurav, et al.
Published: (2025)
by: Seth, Gaurav, et al.
Published: (2025)
Evaluating OCR Performance for Assistive Technology: Effects of Walking Speed, Camera Placement, and Camera Type
by: Feng, Junchi, et al.
Published: (2026)
by: Feng, Junchi, et al.
Published: (2026)
Exploring the Use of VLMs for Navigation Assistance for People with Blindness and Low Vision
by: Li, Yu, et al.
Published: (2026)
by: Li, Yu, et al.
Published: (2026)
LEO-VL: Efficient Scene Representation for Scalable 3D Vision-Language Learning
by: Huang, Jiangyong, et al.
Published: (2025)
by: Huang, Jiangyong, et al.
Published: (2025)
Weather-Robust Scene Semantics with Vision-Aligned 4D Radar
by: Hamilton, Kali, et al.
Published: (2026)
by: Hamilton, Kali, et al.
Published: (2026)
KRAST: Knowledge-Augmented Robotic Action Recognition with Structured Text for Vision-Language Models
by: Nguyen, Son Hai, et al.
Published: (2025)
by: Nguyen, Son Hai, et al.
Published: (2025)
VARS: Vision-based Assessment of Risk in Security Systems
by: Gupta, Pranav, et al.
Published: (2024)
by: Gupta, Pranav, et al.
Published: (2024)
Enhancing Gait Video Analysis in Neurodegenerative Diseases by Knowledge Augmentation in Vision Language Model
by: Wang, Diwei, et al.
Published: (2024)
by: Wang, Diwei, et al.
Published: (2024)
Dynamic Scene Understanding from Vision-Language Representations
by: Pruss, Shahaf, et al.
Published: (2025)
by: Pruss, Shahaf, et al.
Published: (2025)
Leveraging Geometric Priors for Unaligned Scene Change Detection
by: Liu, Ziling, et al.
Published: (2025)
by: Liu, Ziling, et al.
Published: (2025)
Perception Without Vision for Trajectory Prediction: Ego Vehicle Dynamics as Scene Representation for Efficient Active Learning in Autonomous Driving
by: Greer, Ross, et al.
Published: (2024)
by: Greer, Ross, et al.
Published: (2024)
A Vision-Language Framework for Multispectral Scene Representation Using Language-Grounded Features
by: Karanfil, Enes, et al.
Published: (2025)
by: Karanfil, Enes, et al.
Published: (2025)
Uncovering and Shaping the Latent Representation of 3D Scene Topology in Vision-Language Models
by: Wang, Haoming, et al.
Published: (2026)
by: Wang, Haoming, et al.
Published: (2026)
Zero-Shot Scene Change Detection
by: Cho, Kyusik, et al.
Published: (2024)
by: Cho, Kyusik, et al.
Published: (2024)
Vision and Language: Novel Representations and Artificial intelligence for Driving Scene Safety Assessment and Autonomous Vehicle Planning
by: Greer, Ross, et al.
Published: (2026)
by: Greer, Ross, et al.
Published: (2026)
Environmental Change Detection: Toward a Practical Task of Scene Change Detection
by: Cho, Kyusik, et al.
Published: (2025)
by: Cho, Kyusik, et al.
Published: (2025)
From Image Hashing to Scene Change Detection
by: Duong, Anh-Kiet, et al.
Published: (2026)
by: Duong, Anh-Kiet, et al.
Published: (2026)
SAMoE-VLA: A Scene Adaptive Mixture-of-Experts Vision-Language-Action Model for Autonomous Driving
by: You, Zihan, et al.
Published: (2026)
by: You, Zihan, et al.
Published: (2026)
Lookahead Exploration with Neural Radiance Representation for Continuous Vision-Language Navigation
by: Wang, Zihan, et al.
Published: (2024)
by: Wang, Zihan, et al.
Published: (2024)
Seeing Beyond the Scene: Enhancing Vision-Language Models with Interactional Reasoning
by: Liang, Dayong, et al.
Published: (2025)
by: Liang, Dayong, et al.
Published: (2025)
LV-OSD: Language-Vision-Complementary Open-Set Object Detection
by: Zhang, Yupeng, et al.
Published: (2026)
by: Zhang, Yupeng, et al.
Published: (2026)
AGIR: Assessing 3D Gait Impairment with Reasoning based on LLMs
by: Wang, Diwei, et al.
Published: (2025)
by: Wang, Diwei, et al.
Published: (2025)
doScenes: An Autonomous Driving Dataset with Natural Language Instruction for Human Interaction and Vision-Language Navigation
by: Roy, Parthib, et al.
Published: (2024)
by: Roy, Parthib, et al.
Published: (2024)
Seeing Across Views: Benchmarking Spatial Reasoning of Vision-Language Models in Robotic Scenes
by: Feng, Zhiyuan, et al.
Published: (2025)
by: Feng, Zhiyuan, et al.
Published: (2025)
Evaluating Cascaded Methods of Vision-Language Models for Zero-Shot Detection and Association of Hardhats for Increased Construction Safety
by: Choi, Lucas, et al.
Published: (2024)
by: Choi, Lucas, et al.
Published: (2024)
Towards Generalizable Scene Change Detection
by: Kim, Jaewoo, et al.
Published: (2024)
by: Kim, Jaewoo, et al.
Published: (2024)
CLAP: Concave Linear APproximation for Quadratic Graph Matching
by: Liang, Yongqing, et al.
Published: (2024)
by: Liang, Yongqing, et al.
Published: (2024)
EMPLACE: Self-Supervised Urban Scene Change Detection
by: Alpherts, Tim, et al.
Published: (2025)
by: Alpherts, Tim, et al.
Published: (2025)
Scenethesis: A Language and Vision Agentic Framework for 3D Scene Generation
by: Ling, Lu, et al.
Published: (2025)
by: Ling, Lu, et al.
Published: (2025)
Changes in Real Time: Online Scene Change Detection with Multi-View Fusion
by: Galappaththige, Chamuditha Jayanga, et al.
Published: (2025)
by: Galappaththige, Chamuditha Jayanga, et al.
Published: (2025)
Benchmarking Multi-Scene Fire and Smoke Detection
by: Han, Xiaoyi, et al.
Published: (2024)
by: Han, Xiaoyi, et al.
Published: (2024)
HMR3D: Hierarchical Multimodal Representation for 3D Scene Understanding with Large Vision-Language Model
by: Li, Chen, et al.
Published: (2025)
by: Li, Chen, et al.
Published: (2025)
SceneSplat: Gaussian Splatting-based Scene Understanding with Vision-Language Pretraining
by: Li, Yue, et al.
Published: (2025)
by: Li, Yue, et al.
Published: (2025)
Scene-VLM: Multimodal Video Scene Segmentation via Vision-Language Models
by: Berman, Nimrod, et al.
Published: (2025)
by: Berman, Nimrod, et al.
Published: (2025)
Shape2Scene: 3D Scene Representation Learning Through Pre-training on Shape Data
by: Feng, Tuo, et al.
Published: (2024)
by: Feng, Tuo, et al.
Published: (2024)
Similar Items
-
Can Foundation Models Reliably Identify Spatial Hazards? A Case Study on Curb Segmentation
by: Sheng, Diwei, et al.
Published: (2024) -
Improving Long-Tailed Object Detection with Balanced Group Softmax and Metric Learning
by: Gaba, Satyam
Published: (2025) -
Generative AI for Enhanced Wildfire Detection: Bridging the Synthetic-Real Domain Gap
by: Gaba, Satyam
Published: (2025) -
Robust Computer-Vision based Construction Site Detection for Assistive-Technology Applications
by: Feng, Junchi, et al.
Published: (2025) -
NYC-Indoor-VPR: A Long-Term Indoor Visual Place Recognition Dataset with Semi-Automatic Annotation
by: Sheng, Diwei, et al.
Published: (2024)