Vision-Based Localization and LLM-based Navigation for Indoor Environments
Fuente:
arXiv
Saved in:
| Main Authors: | Rahimi, Keyan, Haque, Md. Wasiul, Dasgupta, Sagar, Rahman, Mizanur |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Grid2Guide: A* Enabled Small Language Model for Indoor Navigation
by: Haque, Md. Wasiul, et al.
Published: (2025)
by: Haque, Md. Wasiul, et al.
Published: (2025)
AdaptViG: Adaptive Vision GNN with Exponential Decay Gating
by: Munir, Mustafa, et al.
Published: (2025)
by: Munir, Mustafa, et al.
Published: (2025)
Multi-Surrogate-Teacher Assistance for Representation Alignment in Fingerprint-based Indoor Localization
by: Nguyen, Son Minh, et al.
Published: (2024)
by: Nguyen, Son Minh, et al.
Published: (2024)
FusionEnsemble-Net: An Attention-Based Ensemble of Spatiotemporal Networks for Multimodal Sign Language Recognition
by: Islam, Md. Milon, et al.
Published: (2025)
by: Islam, Md. Milon, et al.
Published: (2025)
GreedyViG: Dynamic Axial Graph Construction for Efficient Vision GNNs
by: Munir, Mustafa, et al.
Published: (2024)
by: Munir, Mustafa, et al.
Published: (2024)
Adaptive Object Detection for Indoor Navigation Assistance: A Performance Evaluation of Real-Time Algorithms
by: Pratap, Abhinav, et al.
Published: (2025)
by: Pratap, Abhinav, et al.
Published: (2025)
LINE: LLM-based Iterative Neuron Explanations for Vision Models
by: Zaigrajew, Vladimir, et al.
Published: (2026)
by: Zaigrajew, Vladimir, et al.
Published: (2026)
Leaf-Based Plant Disease Detection and Explainable AI
by: Sagar, Saurav, et al.
Published: (2023)
by: Sagar, Saurav, et al.
Published: (2023)
Parameter-Efficient Quantized Mixture-of-Experts Meets Vision-Language Instruction Tuning for Semiconductor Electron Micrograph Analysis
by: Srinivas, Sakhinana Sagar, et al.
Published: (2024)
by: Srinivas, Sakhinana Sagar, et al.
Published: (2024)
Foundational Model for Electron Micrograph Analysis: Instruction-Tuning Small-Scale Language-and-Vision Assistant for Enterprise Adoption
by: Srinivas, Sakhinana Sagar, et al.
Published: (2024)
by: Srinivas, Sakhinana Sagar, et al.
Published: (2024)
TeaLeafVision: An Explainable and Robust Deep Learning Framework for Tea Leaf Disease Classification
by: Ahamed, Rafi, et al.
Published: (2026)
by: Ahamed, Rafi, et al.
Published: (2026)
Achieving Pareto Optimality using Efficient Parameter Reduction for DNNs in Resource-Constrained Edge Environment
by: Mih, Atah Nuh, et al.
Published: (2024)
by: Mih, Atah Nuh, et al.
Published: (2024)
ScoreMix: Synthetic Data Generation by Score Composition in Diffusion Models Improves Recognition
by: Rahimi, Parsa, et al.
Published: (2025)
by: Rahimi, Parsa, et al.
Published: (2025)
PlantVillageVQA: A Visual Question Answering Dataset for Benchmarking Vision-Language Models in Plant Science
by: Sakib, Syed Nazmus, et al.
Published: (2025)
by: Sakib, Syed Nazmus, et al.
Published: (2025)
AI and Vision based Autonomous Navigation of Nano-Drones in Partially-Known Environments
by: Sartori, Mattia, et al.
Published: (2025)
by: Sartori, Mattia, et al.
Published: (2025)
Towards Identifiable Unsupervised Domain Translation: A Diversified Distribution Matching Approach
by: Shrestha, Sagar, et al.
Published: (2024)
by: Shrestha, Sagar, et al.
Published: (2024)
Mamba in Vision: A Comprehensive Survey of Techniques and Applications
by: Rahman, Md Maklachur, et al.
Published: (2024)
by: Rahman, Md Maklachur, et al.
Published: (2024)
Lightweight Model for Poultry Disease Detection from Fecal Images Using Multi-Color Space Feature Optimization and Machine Learning
by: Islam, A. K. M. Shoriful, et al.
Published: (2025)
by: Islam, A. K. M. Shoriful, et al.
Published: (2025)
SPHINX: A Synthetic Environment for Visual Perception and Reasoning
by: Alam, Md Tanvirul, et al.
Published: (2025)
by: Alam, Md Tanvirul, et al.
Published: (2025)
Advancing AI-Powered Medical Image Synthesis: Insights from MedVQA-GI Challenge Using CLIP, Fine-Tuned Stable Diffusion, and Dream-Booth + LoRA
by: Peter, Ojonugwa Oluwafemi Ejiga, et al.
Published: (2025)
by: Peter, Ojonugwa Oluwafemi Ejiga, et al.
Published: (2025)
RAPID: Robust and Agile Planner Using Inverse Reinforcement Learning for Vision-Based Drone Navigation
by: Kim, Minwoo, et al.
Published: (2025)
by: Kim, Minwoo, et al.
Published: (2025)
Event-Based Eye Tracking. 2025 Event-based Vision Workshop
by: Chen, Qinyu, et al.
Published: (2025)
by: Chen, Qinyu, et al.
Published: (2025)
Intriguing Equivalence Structures of the Embedding Space of Vision Transformers
by: Salman, Shaeke, et al.
Published: (2024)
by: Salman, Shaeke, et al.
Published: (2024)
DarkQA: Benchmarking Vision-Language Models on Visual-Primitive Question Answering in Low-Light Indoor Scenes
by: Park, Yohan, et al.
Published: (2025)
by: Park, Yohan, et al.
Published: (2025)
A Bidirectional Siamese Recurrent Neural Network for Accurate Gait Recognition Using Body Landmarks
by: Progga, Proma Hossain, et al.
Published: (2024)
by: Progga, Proma Hossain, et al.
Published: (2024)
MobilePlantViT: A Mobile-friendly Hybrid ViT for Generalized Plant Disease Image Classification
by: Tonmoy, Moshiur Rahman, et al.
Published: (2025)
by: Tonmoy, Moshiur Rahman, et al.
Published: (2025)
Automated Toll Management System Using RFID and Image Processing
by: Ahmed, Raihan, et al.
Published: (2024)
by: Ahmed, Raihan, et al.
Published: (2024)
ComposableNav: Instruction-Following Navigation in Dynamic Environments via Composable Diffusion
by: Hu, Zichao, et al.
Published: (2025)
by: Hu, Zichao, et al.
Published: (2025)
AI-Powered Deepfake Detection Using CNN and Vision Transformer Architectures
by: Urmi, Sifatullah Sheikh, et al.
Published: (2026)
by: Urmi, Sifatullah Sheikh, et al.
Published: (2026)
Hierarchical Network Fusion for Multi-Modal Electron Micrograph Representation Learning with Foundational Large Language Models
by: Srinivas, Sakhinana Sagar, et al.
Published: (2024)
by: Srinivas, Sakhinana Sagar, et al.
Published: (2024)
Intriguing Differences Between Zero-Shot and Systematic Evaluations of Vision-Language Transformer Models
by: Salman, Shaeke, et al.
Published: (2024)
by: Salman, Shaeke, et al.
Published: (2024)
Autonomous Navigation in Complex Environments
by: Gerstenslager, Andrew, et al.
Published: (2024)
by: Gerstenslager, Andrew, et al.
Published: (2024)
A Signer-Invariant Conformer and Multi-Scale Fusion Transformer for Continuous Sign Language Recognition
by: Haque, Md Rezwanul, et al.
Published: (2025)
by: Haque, Md Rezwanul, et al.
Published: (2025)
An LLM-Empowered Low-Resolution Vision System for On-Device Human Behavior Understanding
by: Jiang, Siyang, et al.
Published: (2025)
by: Jiang, Siyang, et al.
Published: (2025)
Do We Need Large VLMs for Spotting Soccer Actions?
by: Chakraborty, Ritabrata, et al.
Published: (2025)
by: Chakraborty, Ritabrata, et al.
Published: (2025)
Attention over Scene Graphs: Indoor Scene Representations Toward CSAI Classification
by: Barros, Artur, et al.
Published: (2025)
by: Barros, Artur, et al.
Published: (2025)
MMFactory: A Universal Solution Search Engine for Vision-Language Tasks
by: Fan, Wan-Cyuan, et al.
Published: (2024)
by: Fan, Wan-Cyuan, et al.
Published: (2024)
Enhancement of Bengali OCR by Specialized Models and Advanced Techniques for Diverse Document Types
by: Rabby, AKM Shahariar Azad, et al.
Published: (2024)
by: Rabby, AKM Shahariar Azad, et al.
Published: (2024)
SliceVision-F2I: A Synthetic Feature-to-Image Dataset for Visual Pattern Representation on Network Slices
by: Rafi, Md. Abid Hasan, et al.
Published: (2025)
by: Rafi, Md. Abid Hasan, et al.
Published: (2025)
Deep Learning-Based Digitization of Overlapping ECG Images with Open-Source Python Code
by: Karbasi, Reza, et al.
Published: (2025)
by: Karbasi, Reza, et al.
Published: (2025)
Similar Items
-
Grid2Guide: A* Enabled Small Language Model for Indoor Navigation
by: Haque, Md. Wasiul, et al.
Published: (2025) -
AdaptViG: Adaptive Vision GNN with Exponential Decay Gating
by: Munir, Mustafa, et al.
Published: (2025) -
Multi-Surrogate-Teacher Assistance for Representation Alignment in Fingerprint-based Indoor Localization
by: Nguyen, Son Minh, et al.
Published: (2024) -
FusionEnsemble-Net: An Attention-Based Ensemble of Spatiotemporal Networks for Multimodal Sign Language Recognition
by: Islam, Md. Milon, et al.
Published: (2025) -
GreedyViG: Dynamic Axial Graph Construction for Efficient Vision GNNs
by: Munir, Mustafa, et al.
Published: (2024)