NaVIP: An Image-Centric Indoor Navigation Solution for Visually Impaired People
Fuente:
arXiv
Salvato in:
| Autori principali: | Yu, Jun, Zhang, Yifan, Aila, Badrinadh, Namboodiri, Vinod |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
WalkVLM:Aid Visually Impaired People Walking by Vision Language Model
di: Yuan, Zhiqiang, et al.
Pubblicazione: (2024)
di: Yuan, Zhiqiang, et al.
Pubblicazione: (2024)
Turn-by-Turn Indoor Navigation for the Visually Impaired
di: Srinivasaiah, Santosh, et al.
Pubblicazione: (2024)
di: Srinivasaiah, Santosh, et al.
Pubblicazione: (2024)
RoboVIP: Multi-View Video Generation with Visual Identity Prompting Augments Robot Manipulation
di: Wang, Boyang, et al.
Pubblicazione: (2026)
di: Wang, Boyang, et al.
Pubblicazione: (2026)
AltCanvas: A Tile-Based Image Editor with Generative AI for Blind or Visually Impaired People
di: Lee, Seonghee, et al.
Pubblicazione: (2024)
di: Lee, Seonghee, et al.
Pubblicazione: (2024)
A Large Vision-Language Model based Environment Perception System for Visually Impaired People
di: Chen, Zezhou, et al.
Pubblicazione: (2025)
di: Chen, Zezhou, et al.
Pubblicazione: (2025)
PinpointQA: A Dataset and Benchmark for Small Object-Centric Spatial Understanding in Indoor Videos
di: Zhou, Zhiyu, et al.
Pubblicazione: (2026)
di: Zhou, Zhiyu, et al.
Pubblicazione: (2026)
FloNa: Floor Plan Guided Embodied Visual Navigation
di: Li, Jiaxin, et al.
Pubblicazione: (2024)
di: Li, Jiaxin, et al.
Pubblicazione: (2024)
V.I.P. : Iterative Online Preference Distillation for Efficient Video Diffusion Models
di: Kim, Jisoo, et al.
Pubblicazione: (2025)
di: Kim, Jisoo, et al.
Pubblicazione: (2025)
Inside Knowledge: Graph-based Path Generation with Explainable Data Augmentation and Curriculum Learning for Visual Indoor Navigation
di: Airinei, Daniel, et al.
Pubblicazione: (2025)
di: Airinei, Daniel, et al.
Pubblicazione: (2025)
ROVER: Routing Object-Centric Visual Evidence for Grounded Multi-Image Reasoning
di: Lv, Guannan, et al.
Pubblicazione: (2026)
di: Lv, Guannan, et al.
Pubblicazione: (2026)
Dynamic Multi-Target Fusion for Efficient Audio-Visual Navigation
di: Yu, Yinfeng, et al.
Pubblicazione: (2025)
di: Yu, Yinfeng, et al.
Pubblicazione: (2025)
Navigating Beyond Dropout: An Intriguing Solution Towards Generalizable Image Super Resolution
di: Wang, Hongjun, et al.
Pubblicazione: (2024)
di: Wang, Hongjun, et al.
Pubblicazione: (2024)
\textsc{NaVIDA}: Vision-Language Navigation with Inverse Dynamics Augmentation
di: Zhu, Weiye, et al.
Pubblicazione: (2026)
di: Zhu, Weiye, et al.
Pubblicazione: (2026)
RAW: Robust Avatar Watermarking -- Benchmarking and Baseline
di: Parry, Jack, et al.
Pubblicazione: (2026)
di: Parry, Jack, et al.
Pubblicazione: (2026)
When the Prompt Becomes Visual: Vision-Centric Jailbreak Attacks for Large Image Editing Models
di: Hou, Jiacheng, et al.
Pubblicazione: (2026)
di: Hou, Jiacheng, et al.
Pubblicazione: (2026)
VIALM: A Survey and Benchmark of Visually Impaired Assistance with Large Models
di: Zhao, Yi, et al.
Pubblicazione: (2024)
di: Zhao, Yi, et al.
Pubblicazione: (2024)
AvatarShield: Visual Reinforcement Learning for Human-Centric Synthetic Video Detection
di: Xu, Zhipei, et al.
Pubblicazione: (2025)
di: Xu, Zhipei, et al.
Pubblicazione: (2025)
EIDT-V: Exploiting Intersections in Diffusion Trajectories for Model-Agnostic, Zero-Shot, Training-Free Text-to-Video Generation
di: Jagpal, Diljeet, et al.
Pubblicazione: (2025)
di: Jagpal, Diljeet, et al.
Pubblicazione: (2025)
ND-SDF: Learning Normal Deflection Fields for High-Fidelity Indoor Reconstruction
di: Tang, Ziyu, et al.
Pubblicazione: (2024)
di: Tang, Ziyu, et al.
Pubblicazione: (2024)
Exploring the Use of VLMs for Navigation Assistance for People with Blindness and Low Vision
di: Li, Yu, et al.
Pubblicazione: (2026)
di: Li, Yu, et al.
Pubblicazione: (2026)
Weaving Context Across Images: Improving Vision-Language Models through Focus-Centric Visual Chains
di: Zhang, Juntian, et al.
Pubblicazione: (2025)
di: Zhang, Juntian, et al.
Pubblicazione: (2025)
Transparent Visual Reasoning via Object-Centric Agent Collaboration
di: Teoh, Benjamin, et al.
Pubblicazione: (2025)
di: Teoh, Benjamin, et al.
Pubblicazione: (2025)
Slot-MLLM: Object-Centric Visual Tokenization for Multimodal LLM
di: Chi, Donghwan, et al.
Pubblicazione: (2025)
di: Chi, Donghwan, et al.
Pubblicazione: (2025)
Real-Time Pill Identification for the Visually Impaired Using Deep Learning
di: Dang, Bo, et al.
Pubblicazione: (2024)
di: Dang, Bo, et al.
Pubblicazione: (2024)
MotionEdit: Benchmarking and Learning Motion-Centric Image Editing
di: Wan, Yixin, et al.
Pubblicazione: (2025)
di: Wan, Yixin, et al.
Pubblicazione: (2025)
Fine-Tuning Vision-Language Models for Visual Navigation Assistance
di: Li, Xiao, et al.
Pubblicazione: (2025)
di: Li, Xiao, et al.
Pubblicazione: (2025)
Revisiting Salient Object Detection from an Observer-Centric Perspective
di: Zhang, Fuxi, et al.
Pubblicazione: (2026)
di: Zhang, Fuxi, et al.
Pubblicazione: (2026)
RVN-Bench: A Benchmark for Reactive Visual Navigation
di: Lee, Jaewon, et al.
Pubblicazione: (2026)
di: Lee, Jaewon, et al.
Pubblicazione: (2026)
Think Hierarchically, Act Dynamically: Hierarchical Multi-modal Fusion and Reasoning for Vision-and-Language Navigation
di: Yue, Junrong, et al.
Pubblicazione: (2025)
di: Yue, Junrong, et al.
Pubblicazione: (2025)
Vision-Based Localization and LLM-based Navigation for Indoor Environments
di: Rahimi, Keyan, et al.
Pubblicazione: (2025)
di: Rahimi, Keyan, et al.
Pubblicazione: (2025)
IVLMap: Instance-Aware Visual Language Grounding for Consumer Robot Navigation
di: Huang, Jiacui, et al.
Pubblicazione: (2024)
di: Huang, Jiacui, et al.
Pubblicazione: (2024)
Dual-View Visual Contextualization for Web Navigation
di: Kil, Jihyung, et al.
Pubblicazione: (2024)
di: Kil, Jihyung, et al.
Pubblicazione: (2024)
Residual Cross-Modal Fusion Networks for Audio-Visual Navigation
di: Wang, Yi, et al.
Pubblicazione: (2026)
di: Wang, Yi, et al.
Pubblicazione: (2026)
Fine-Grained Controllable Apparel Showcase Image Generation via Garment-Centric Outpainting
di: Zhang, Rong, et al.
Pubblicazione: (2025)
di: Zhang, Rong, et al.
Pubblicazione: (2025)
Learning Disentangled Representation in Object-Centric Models for Visual Dynamics Prediction via Transformers
di: Gandhi, Sanket, et al.
Pubblicazione: (2024)
di: Gandhi, Sanket, et al.
Pubblicazione: (2024)
Decision-based AI Visual Navigation for Cardiac Ultrasounds
di: Dimnaku, Andy, et al.
Pubblicazione: (2025)
di: Dimnaku, Andy, et al.
Pubblicazione: (2025)
Audio-Guided Visual Perception for Audio-Visual Navigation
di: Wang, Yi, et al.
Pubblicazione: (2025)
di: Wang, Yi, et al.
Pubblicazione: (2025)
ANAVI: Audio Noise Awareness using Visuals of Indoor environments for NAVIgation
di: Jain, Vidhi, et al.
Pubblicazione: (2024)
di: Jain, Vidhi, et al.
Pubblicazione: (2024)
Solution for CVPR 2024 UG2+ Challenge Track on All Weather Semantic Segmentation
di: Yu, Jun, et al.
Pubblicazione: (2024)
di: Yu, Jun, et al.
Pubblicazione: (2024)
3D MRI Image Pretraining via Controllable 2D Slice Navigation Task
di: Wang, Yu, et al.
Pubblicazione: (2026)
di: Wang, Yu, et al.
Pubblicazione: (2026)
Documenti analoghi
-
WalkVLM:Aid Visually Impaired People Walking by Vision Language Model
di: Yuan, Zhiqiang, et al.
Pubblicazione: (2024) -
Turn-by-Turn Indoor Navigation for the Visually Impaired
di: Srinivasaiah, Santosh, et al.
Pubblicazione: (2024) -
RoboVIP: Multi-View Video Generation with Visual Identity Prompting Augments Robot Manipulation
di: Wang, Boyang, et al.
Pubblicazione: (2026) -
AltCanvas: A Tile-Based Image Editor with Generative AI for Blind or Visually Impaired People
di: Lee, Seonghee, et al.
Pubblicazione: (2024) -
A Large Vision-Language Model based Environment Perception System for Visually Impaired People
di: Chen, Zezhou, et al.
Pubblicazione: (2025)