Urban Risk-Aware Navigation via VQA-Based Event Maps for People with Low Vision
Fuente:
arXiv
Guardado en:
| Autores principales: | Valls, Antoni, Sanchez-Riera, Jordi |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Uncertainty-Aware Gaussian Map for Vision-Language Navigation
por: Gao, Jianzhe, et al.
Publicado: (2026)
por: Gao, Jianzhe, et al.
Publicado: (2026)
Vision-Based Autonomous UAV Navigation and Landing for Urban Search and Rescue
por: Mittal, Mayank, et al.
Publicado: (2019)
por: Mittal, Mayank, et al.
Publicado: (2019)
Vision-Based Risk Aware Emergency Landing for UAVs in Complex Urban Environments
por: de la Torre-Vanegas, Julio, et al.
Publicado: (2025)
por: de la Torre-Vanegas, Julio, et al.
Publicado: (2025)
Exploring the Use of VLMs for Navigation Assistance for People with Blindness and Low Vision
por: Li, Yu, et al.
Publicado: (2026)
por: Li, Yu, et al.
Publicado: (2026)
Empowering Dynamic Urban Navigation with Stereo and Mid-Level Vision
por: Zhou, Wentao, et al.
Publicado: (2025)
por: Zhou, Wentao, et al.
Publicado: (2025)
Ask4VG: Risk-Aware Question Selection for Reducing Prior-Driven Answers in Medical VQA
por: Zhu, Xiaorong, et al.
Publicado: (2026)
por: Zhu, Xiaorong, et al.
Publicado: (2026)
Embodied Scene Understanding for Vision Language Models via MetaVQA
por: Wang, Weizhen, et al.
Publicado: (2025)
por: Wang, Weizhen, et al.
Publicado: (2025)
Verification Mirage: Mapping the Reliability Boundary of Self-Verification in Medical VQA
por: Jin, Ruinan, et al.
Publicado: (2026)
por: Jin, Ruinan, et al.
Publicado: (2026)
Uncertainty-Aware Vision-based Risk Object Identification via Conformal Risk Tube Prediction
por: Fu, Kai-Yu, et al.
Publicado: (2026)
por: Fu, Kai-Yu, et al.
Publicado: (2026)
VTAgent: Agentic Keyframe Anchoring for Evidence-Aware Video TextVQA
por: He, Haibin, et al.
Publicado: (2026)
por: He, Haibin, et al.
Publicado: (2026)
LookasideVLN: Direction-Aware Aerial Vision-and-Language Navigation
por: Ning, Yuwei, et al.
Publicado: (2026)
por: Ning, Yuwei, et al.
Publicado: (2026)
Identifying Crucial Objects in Blind and Low-Vision Individuals' Navigation
por: Islam, Md Touhidul, et al.
Publicado: (2024)
por: Islam, Md Touhidul, et al.
Publicado: (2024)
IRIS: Intent Resolution via Inference-time Saccades for Open-Ended VQA in Large Vision-Language Models
por: Madinei, Parsa, et al.
Publicado: (2026)
por: Madinei, Parsa, et al.
Publicado: (2026)
Do Vision-Language Models See Urban Scenes as People Do? An Urban Perception Benchmark
por: Mushkani, Rashid
Publicado: (2025)
por: Mushkani, Rashid
Publicado: (2025)
CC-VQA: Conflict- and Correlation-Aware Method for Mitigating Knowledge Conflict in Knowledge-Based Visual Question Answering
por: Hong, Yuyang, et al.
Publicado: (2026)
por: Hong, Yuyang, et al.
Publicado: (2026)
EventFlash: Towards Efficient MLLMs for Event-Based Vision
por: Liu, Shaoyu, et al.
Publicado: (2026)
por: Liu, Shaoyu, et al.
Publicado: (2026)
EduVQA: Towards Concept-Aware Assessment of Educational AI-Generated Videos
por: Chen, Baoliang, et al.
Publicado: (2026)
por: Chen, Baoliang, et al.
Publicado: (2026)
Look-Closer-Then-Diagnose: Confidence-Aware Ultrasound VQA via Active Zooming
por: Zhou, Yue, et al.
Publicado: (2026)
por: Zhou, Yue, et al.
Publicado: (2026)
Aerial Vision-and-Language Navigation with Grid-based View Selection and Map Construction
por: Zhao, Ganlong, et al.
Publicado: (2025)
por: Zhao, Ganlong, et al.
Publicado: (2025)
Dynamic Topology Awareness: Breaking the Granularity Rigidity in Vision-Language Navigation
por: Peng, Jiankun, et al.
Publicado: (2026)
por: Peng, Jiankun, et al.
Publicado: (2026)
Multi-View People Detection in Large Scenes via Supervised View-Wise Contribution Weighting
por: Zhang, Qi, et al.
Publicado: (2024)
por: Zhang, Qi, et al.
Publicado: (2024)
Vision-Language Navigation with Energy-Based Policy
por: Liu, Rui, et al.
Publicado: (2024)
por: Liu, Rui, et al.
Publicado: (2024)
SurgCheck: Do Vision-Language Models Really Look at Images in Surgical VQA?
por: Shin, Jongmin, et al.
Publicado: (2026)
por: Shin, Jongmin, et al.
Publicado: (2026)
Head-Aware Visual Cropping: Enhancing Fine-Grained VQA with Attention-Guided Subimage
por: Xie, Junfei, et al.
Publicado: (2026)
por: Xie, Junfei, et al.
Publicado: (2026)
InstantAvatar: Efficient 3D Head Reconstruction via Surface Rendering
por: Canela, Antonio, et al.
Publicado: (2023)
por: Canela, Antonio, et al.
Publicado: (2023)
3D Gaussian Map with Open-Set Semantic Grouping for Vision-Language Navigation
por: Gao, Jianzhe, et al.
Publicado: (2026)
por: Gao, Jianzhe, et al.
Publicado: (2026)
Enhancing Document VQA Models via Retrieval-Augmented Generation
por: López, Eric, et al.
Publicado: (2025)
por: López, Eric, et al.
Publicado: (2025)
Look to Locate: Vision-Based Multisensory Navigation with 3-D Digital Maps for GNSS-Challenged Environments
por: Elmaghraby, Ola, et al.
Publicado: (2025)
por: Elmaghraby, Ola, et al.
Publicado: (2025)
PASTS: Progress-Aware Spatio-Temporal Transformer Speaker For Vision-and-Language Navigation
por: Wang, Liuyi, et al.
Publicado: (2023)
por: Wang, Liuyi, et al.
Publicado: (2023)
Adaptive Event Stream Slicing for Open-Vocabulary Event-Based Object Detection via Vision-Language Knowledge Distillation
por: Zhang, Jinchang, et al.
Publicado: (2025)
por: Zhang, Jinchang, et al.
Publicado: (2025)
Interpreting Low-level Vision Models with Causal Effect Maps
por: Hu, Jinfan, et al.
Publicado: (2024)
por: Hu, Jinfan, et al.
Publicado: (2024)
UrbanSAM: Learning Invariance-Inspired Adapters for Segment Anything Models in Urban Construction
por: Li, Chenyu, et al.
Publicado: (2025)
por: Li, Chenyu, et al.
Publicado: (2025)
AwareVLN: Reasoning with Self-awareness for Vision-Language Navigation
por: Guo, Wenxuan, et al.
Publicado: (2026)
por: Guo, Wenxuan, et al.
Publicado: (2026)
From Street View to Visual Network: Mapping the Visibility of Urban Landmarks with Vision-Language Models
por: Fan, Zicheng, et al.
Publicado: (2025)
por: Fan, Zicheng, et al.
Publicado: (2025)
Vision-Based Localization in Dense Urban Environments: A Case Study of an Urban Village in China
por: Wu, Menglin, et al.
Publicado: (2026)
por: Wu, Menglin, et al.
Publicado: (2026)
SurgAnt-ViVQA: Learning to Anticipate Surgical Events through GRU-Driven Temporal Cross-Attention
por: Dhake, Shreyas C., et al.
Publicado: (2025)
por: Dhake, Shreyas C., et al.
Publicado: (2025)
Light-VQA+: A Video Quality Assessment Model for Exposure Correction with Vision-Language Guidance
por: Zhou, Xunchu, et al.
Publicado: (2024)
por: Zhou, Xunchu, et al.
Publicado: (2024)
IllusionVQA: A Challenging Optical Illusion Dataset for Vision Language Models
por: Shahgir, Haz Sameen, et al.
Publicado: (2024)
por: Shahgir, Haz Sameen, et al.
Publicado: (2024)
On the Role of Visual Grounding in VQA
por: Reich, Daniel, et al.
Publicado: (2024)
por: Reich, Daniel, et al.
Publicado: (2024)
RoboSurg-VQA: A Multimodal Benchmark for Surgical Segmentation-Aware Visual Question Answering
por: Zhang, Chengyi, et al.
Publicado: (2026)
por: Zhang, Chengyi, et al.
Publicado: (2026)
Ejemplares similares
-
Uncertainty-Aware Gaussian Map for Vision-Language Navigation
por: Gao, Jianzhe, et al.
Publicado: (2026) -
Vision-Based Autonomous UAV Navigation and Landing for Urban Search and Rescue
por: Mittal, Mayank, et al.
Publicado: (2019) -
Vision-Based Risk Aware Emergency Landing for UAVs in Complex Urban Environments
por: de la Torre-Vanegas, Julio, et al.
Publicado: (2025) -
Exploring the Use of VLMs for Navigation Assistance for People with Blindness and Low Vision
por: Li, Yu, et al.
Publicado: (2026) -
Empowering Dynamic Urban Navigation with Stereo and Mid-Level Vision
por: Zhou, Wentao, et al.
Publicado: (2025)