ANAVI: Audio Noise Awareness using Visuals of Indoor environments for NAVIgation
Fuente:
arXiv
Saved in:
| Main Authors: | Jain, Vidhi, Veerapaneni, Rishi, Bisk, Yonatan |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
MotIF: Motion Instruction Fine-tuning
by: Hwang, Minyoung, et al.
Published: (2024)
by: Hwang, Minyoung, et al.
Published: (2024)
HomeRobot: Open-Vocabulary Mobile Manipulation
by: Yenamandra, Sriram, et al.
Published: (2023)
by: Yenamandra, Sriram, et al.
Published: (2023)
RAVEN: Resilient Aerial Navigation via Open-Set Semantic Memory and Behavior Adaptation
by: Kim, Seungchan, et al.
Published: (2025)
by: Kim, Seungchan, et al.
Published: (2025)
Affordance RAG: Hierarchical Multimodal Retrieval with Affordance-Aware Embodied Memory for Mobile Manipulation
by: Korekata, Ryosuke, et al.
Published: (2025)
by: Korekata, Ryosuke, et al.
Published: (2025)
Residual Cross-Modal Fusion Networks for Audio-Visual Navigation
by: Wang, Yi, et al.
Published: (2026)
by: Wang, Yi, et al.
Published: (2026)
VISREAS: Complex Visual Reasoning with Unanswerable Questions
by: Akter, Syeda Nahida, et al.
Published: (2024)
by: Akter, Syeda Nahida, et al.
Published: (2024)
Enhancing Indoor Mobility with Connected Sensor Nodes: A Real-Time, Delay-Aware Cooperative Perception Approach
by: Ning, Minghao, et al.
Published: (2024)
by: Ning, Minghao, et al.
Published: (2024)
DegustaBot: Zero-Shot Visual Preference Estimation for Personalized Multi-Object Rearrangement
by: Newman, Benjamin A., et al.
Published: (2024)
by: Newman, Benjamin A., et al.
Published: (2024)
Toward General-Purpose Robots via Foundation Models: A Survey and Meta-Analysis
by: Hu, Yafei, et al.
Published: (2023)
by: Hu, Yafei, et al.
Published: (2023)
RobotArena $\infty$: Scalable Robot Benchmarking via Real-to-Sim Translation
by: Jangir, Yash, et al.
Published: (2025)
by: Jangir, Yash, et al.
Published: (2025)
Mapping High-level Semantic Regions in Indoor Environments without Object Recognition
by: Bigazzi, Roberto, et al.
Published: (2024)
by: Bigazzi, Roberto, et al.
Published: (2024)
AiSDF: Structure-aware Neural Signed Distance Fields in Indoor Scenes
by: Jang, Jaehoon, et al.
Published: (2024)
by: Jang, Jaehoon, et al.
Published: (2024)
Audio-3DVG: Unified Audio -- Point Cloud Fusion for 3D Visual Grounding
by: Cao-Dinh, Duc, et al.
Published: (2025)
by: Cao-Dinh, Duc, et al.
Published: (2025)
On the Application of Efficient Neural Mapping to Real-Time Indoor Localisation for Unmanned Ground Vehicles
by: Holder, Christopher J., et al.
Published: (2022)
by: Holder, Christopher J., et al.
Published: (2022)
RGB-only Active 3D Scene Graph Generation for Indoor Mobile Robots
by: Modi, Giorgia, et al.
Published: (2026)
by: Modi, Giorgia, et al.
Published: (2026)
EA-WM: Event-Aware Generative World Model with Structured Kinematic-to-Visual Action Fields
by: Yang, Zhaoyang, et al.
Published: (2026)
by: Yang, Zhaoyang, et al.
Published: (2026)
LatentPilot: Scene-Aware Vision-and-Language Navigation by Dreaming Ahead with Latent Visual Reasoning
by: Hao, Haihong, et al.
Published: (2026)
by: Hao, Haihong, et al.
Published: (2026)
GLEAM: Learning Generalizable Exploration Policy for Active Mapping in Complex 3D Indoor Scenes
by: Chen, Xiao, et al.
Published: (2025)
by: Chen, Xiao, et al.
Published: (2025)
REM: Evaluating LLM Embodied Spatial Reasoning through Multi-Frame Trajectories
by: Thompson, Jacob, et al.
Published: (2025)
by: Thompson, Jacob, et al.
Published: (2025)
Hearing Touch: Audio-Visual Pretraining for Contact-Rich Manipulation
by: Mejia, Jared, et al.
Published: (2024)
by: Mejia, Jared, et al.
Published: (2024)
Pair-VPR: Place-Aware Pre-training and Contrastive Pair Classification for Visual Place Recognition with Vision Transformers
by: Hausler, Stephen, et al.
Published: (2024)
by: Hausler, Stephen, et al.
Published: (2024)
SceneSmith: Agentic Generation of Simulation-Ready Indoor Scenes
by: Pfaff, Nicholas, et al.
Published: (2026)
by: Pfaff, Nicholas, et al.
Published: (2026)
DarkQA: Benchmarking Vision-Language Models on Visual-Primitive Question Answering in Low-Light Indoor Scenes
by: Park, Yohan, et al.
Published: (2025)
by: Park, Yohan, et al.
Published: (2025)
Ross3D: Reconstructive Visual Instruction Tuning with 3D-Awareness
by: Wang, Haochen, et al.
Published: (2025)
by: Wang, Haochen, et al.
Published: (2025)
INoD: Injected Noise Discriminator for Self-Supervised Representation Learning in Agricultural Fields
by: Hindel, Julia, et al.
Published: (2023)
by: Hindel, Julia, et al.
Published: (2023)
Empowering Embodied Visual Tracking with Visual Foundation Models and Offline RL
by: Zhong, Fangwei, et al.
Published: (2024)
by: Zhong, Fangwei, et al.
Published: (2024)
Robotic Visual Instruction
by: Li, Yanbang, et al.
Published: (2025)
by: Li, Yanbang, et al.
Published: (2025)
DINO-VO: A Feature-based Visual Odometry Leveraging a Visual Foundation Model
by: Azhari, Maulana Bisyir, et al.
Published: (2025)
by: Azhari, Maulana Bisyir, et al.
Published: (2025)
CRAFT: A Tendon-Driven Hand with Hybrid Hard-Soft Compliance
by: Lin, Leo, et al.
Published: (2026)
by: Lin, Leo, et al.
Published: (2026)
Embodied Uncertainty-Aware Object Segmentation
by: Fang, Xiaolin, et al.
Published: (2024)
by: Fang, Xiaolin, et al.
Published: (2024)
DVGT: Driving Visual Geometry Transformer
by: Zuo, Sicheng, et al.
Published: (2025)
by: Zuo, Sicheng, et al.
Published: (2025)
Subtask-Aware Visual Reward Learning from Segmented Demonstrations
by: Kim, Changyeon, et al.
Published: (2025)
by: Kim, Changyeon, et al.
Published: (2025)
Visual IRL for Human-Like Robotic Manipulation
by: Asali, Ehsan, et al.
Published: (2024)
by: Asali, Ehsan, et al.
Published: (2024)
Visual SLAMMOT Considering Multiple Motion Models
by: Tian, Peilin, et al.
Published: (2024)
by: Tian, Peilin, et al.
Published: (2024)
Language-Conditioned World Modeling for Visual Navigation
by: Dong, Yifei, et al.
Published: (2026)
by: Dong, Yifei, et al.
Published: (2026)
Robotic Manipulation by Imitating Generated Videos Without Physical Demonstrations
by: Patel, Shivansh, et al.
Published: (2025)
by: Patel, Shivansh, et al.
Published: (2025)
MemoNav: Working Memory Model for Visual Navigation
by: Li, Hongxin, et al.
Published: (2024)
by: Li, Hongxin, et al.
Published: (2024)
Spatially Visual Perception for End-to-End Robotic Learning
by: Davies, Travis, et al.
Published: (2024)
by: Davies, Travis, et al.
Published: (2024)
Open-Vocabulary Action Localization with Iterative Visual Prompting
by: Wake, Naoki, et al.
Published: (2024)
by: Wake, Naoki, et al.
Published: (2024)
Sparse Imagination for Efficient Visual World Model Planning
by: Chun, Junha, et al.
Published: (2025)
by: Chun, Junha, et al.
Published: (2025)
Similar Items
-
MotIF: Motion Instruction Fine-tuning
by: Hwang, Minyoung, et al.
Published: (2024) -
HomeRobot: Open-Vocabulary Mobile Manipulation
by: Yenamandra, Sriram, et al.
Published: (2023) -
RAVEN: Resilient Aerial Navigation via Open-Set Semantic Memory and Behavior Adaptation
by: Kim, Seungchan, et al.
Published: (2025) -
Affordance RAG: Hierarchical Multimodal Retrieval with Affordance-Aware Embodied Memory for Mobile Manipulation
by: Korekata, Ryosuke, et al.
Published: (2025) -
Residual Cross-Modal Fusion Networks for Audio-Visual Navigation
by: Wang, Yi, et al.
Published: (2026)