DialNav: Multi-turn Dialog Navigation with a Remote Guide
Fuente:
arXiv
Salvato in:
| Autori principali: | Han, Leekyeung, Min, Hyunji, Hwangbo, Gyeom, Choi, Jonghyun, Seo, Paul Hongsuck |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Robust Image Self-Recovery against Tampering using Watermark Generation with Pixel Shuffling
di: Kim, Minyoung, et al.
Pubblicazione: (2025)
di: Kim, Minyoung, et al.
Pubblicazione: (2025)
Attention Misses Visual Risk: Risk-Adaptive Steering for Multimodal Safety Alignment
di: Park, Jonghyun, et al.
Pubblicazione: (2025)
di: Park, Jonghyun, et al.
Pubblicazione: (2025)
Image Diffusion Models Exhibit Emergent Temporal Propagation in Videos
di: Kim, Youngseo, et al.
Pubblicazione: (2025)
di: Kim, Youngseo, et al.
Pubblicazione: (2025)
V$^2$Dial: Unification of Video and Visual Dialog via Multimodal Experts
di: Abdessaied, Adnen, et al.
Pubblicazione: (2025)
di: Abdessaied, Adnen, et al.
Pubblicazione: (2025)
Multi-Granularity Video Object Segmentation
di: Lim, Sangbeom, et al.
Pubblicazione: (2024)
di: Lim, Sangbeom, et al.
Pubblicazione: (2024)
GaussNav: Gaussian Splatting for Visual Navigation
di: Lei, Xiaohan, et al.
Pubblicazione: (2024)
di: Lei, Xiaohan, et al.
Pubblicazione: (2024)
Pseudo-RIS: Distinctive Pseudo-supervision Generation for Referring Image Segmentation
di: Yu, Seonghoon, et al.
Pubblicazione: (2024)
di: Yu, Seonghoon, et al.
Pubblicazione: (2024)
Learning Correlation Structures for Vision Transformers
di: Kim, Manjin, et al.
Pubblicazione: (2024)
di: Kim, Manjin, et al.
Pubblicazione: (2024)
DialogGen: Multi-modal Interactive Dialogue System for Multi-turn Text-to-Image Generation
di: Huang, Minbin, et al.
Pubblicazione: (2024)
di: Huang, Minbin, et al.
Pubblicazione: (2024)
Bridging Audio and Vision: Zero-Shot Audiovisual Segmentation by Connecting Pretrained Models
di: Lee, Seung-jae, et al.
Pubblicazione: (2025)
di: Lee, Seung-jae, et al.
Pubblicazione: (2025)
Direct Diffusion Score Preference Optimization via Stepwise Contrastive Policy-Pair Supervision
di: Kim, Dohyun, et al.
Pubblicazione: (2025)
di: Kim, Dohyun, et al.
Pubblicazione: (2025)
Budgeted Online Continual Learning by Adaptive Layer Freezing and Frequency-based Sampling
di: Seo, Minhyuk, et al.
Pubblicazione: (2024)
di: Seo, Minhyuk, et al.
Pubblicazione: (2024)
OASIS: Online Sample Selection for Continual Visual Instruction Tuning
di: Lee, Minjae, et al.
Pubblicazione: (2025)
di: Lee, Minjae, et al.
Pubblicazione: (2025)
CoNav: Collaborative Cross-Modal Reasoning for Embodied Navigation
di: Hao, Haihong, et al.
Pubblicazione: (2025)
di: Hao, Haihong, et al.
Pubblicazione: (2025)
CAT-Seg: Cost Aggregation for Open-Vocabulary Semantic Segmentation
di: Cho, Seokju, et al.
Pubblicazione: (2023)
di: Cho, Seokju, et al.
Pubblicazione: (2023)
Spectral-Adaptive Modulation Networks for Visual Perception
di: Yun, Guhnoo, et al.
Pubblicazione: (2025)
di: Yun, Guhnoo, et al.
Pubblicazione: (2025)
UrbanNav: Learning Language-Guided Urban Navigation from Web-Scale Human Trajectories
di: Mei, Yanghong, et al.
Pubblicazione: (2025)
di: Mei, Yanghong, et al.
Pubblicazione: (2025)
Breaking the Visual Shortcuts in Multimodal Knowledge-Based Visual Question Answering
di: Lee, Dosung, et al.
Pubblicazione: (2025)
di: Lee, Dosung, et al.
Pubblicazione: (2025)
Towards Open-Vocabulary Semantic Segmentation Without Semantic Labels
di: Shin, Heeseong, et al.
Pubblicazione: (2024)
di: Shin, Heeseong, et al.
Pubblicazione: (2024)
Nav-R1: Reasoning and Navigation in Embodied Scenes
di: Liu, Qingxiang, et al.
Pubblicazione: (2025)
di: Liu, Qingxiang, et al.
Pubblicazione: (2025)
DialogCC: An Automated Pipeline for Creating High-Quality Multi-Modal Dialogue Dataset
di: Lee, Young-Jun, et al.
Pubblicazione: (2022)
di: Lee, Young-Jun, et al.
Pubblicazione: (2022)
MM-Nav: Multi-View VLA Model for Robust Visual Navigation via Multi-Expert Learning
di: Xu, Tianyu, et al.
Pubblicazione: (2025)
di: Xu, Tianyu, et al.
Pubblicazione: (2025)
PlatonicNav: Unveiling Semantic Correspondence in Navigation with Platonic Topological Maps
di: Long, Junlin, et al.
Pubblicazione: (2026)
di: Long, Junlin, et al.
Pubblicazione: (2026)
NavBench: Probing Multimodal Large Language Models for Embodied Navigation
di: Qiao, Yanyuan, et al.
Pubblicazione: (2025)
di: Qiao, Yanyuan, et al.
Pubblicazione: (2025)
EarthDial: Turning Multi-sensory Earth Observations to Interactive Dialogues
di: Soni, Sagar, et al.
Pubblicazione: (2024)
di: Soni, Sagar, et al.
Pubblicazione: (2024)
ShadowNav: Autonomous Global Localization for Lunar Navigation in Darkness
di: Atha, Deegan, et al.
Pubblicazione: (2024)
di: Atha, Deegan, et al.
Pubblicazione: (2024)
Hyp2Nav: Hyperbolic Planning and Curiosity for Crowd Navigation
di: di Melendugno, Guido Maria D'Amely, et al.
Pubblicazione: (2024)
di: di Melendugno, Guido Maria D'Amely, et al.
Pubblicazione: (2024)
CoNav: A Benchmark for Human-Centered Collaborative Navigation
di: Li, Changhao, et al.
Pubblicazione: (2024)
di: Li, Changhao, et al.
Pubblicazione: (2024)
FOM-Nav: Frontier-Object Maps for Object Goal Navigation
di: Chabal, Thomas, et al.
Pubblicazione: (2025)
di: Chabal, Thomas, et al.
Pubblicazione: (2025)
SR-Nav: Spatial Relationships Matter for Zero-shot Object Goal Navigation
di: Fang, Leyuan, et al.
Pubblicazione: (2026)
di: Fang, Leyuan, et al.
Pubblicazione: (2026)
WarNav: An Autonomous Driving Benchmark for Segmentation of Navigable Zones in War Scenes
di: Graviers, Marc-Emmanuel Coupvent des, et al.
Pubblicazione: (2025)
di: Graviers, Marc-Emmanuel Coupvent des, et al.
Pubblicazione: (2025)
GuideNav: User-Informed Development of a Vision-Only Robotic Navigation Assistant For Blind Travelers
di: Hwang, Hochul, et al.
Pubblicazione: (2025)
di: Hwang, Hochul, et al.
Pubblicazione: (2025)
LongNav-R1: Horizon-Adaptive Multi-Turn RL for Long-Horizon VLA Navigation
di: Hu, Yue, et al.
Pubblicazione: (2026)
di: Hu, Yue, et al.
Pubblicazione: (2026)
NavAgent: Multi-scale Urban Street View Fusion For UAV Embodied Vision-and-Language Navigation
di: Liu, Youzhi, et al.
Pubblicazione: (2024)
di: Liu, Youzhi, et al.
Pubblicazione: (2024)
Seg4Diff: Unveiling Open-Vocabulary Segmentation in Text-to-Image Diffusion Transformers
di: Kim, Chaehyun, et al.
Pubblicazione: (2025)
di: Kim, Chaehyun, et al.
Pubblicazione: (2025)
NavQ: Learning a Q-Model for Foresighted Vision-and-Language Navigation
di: Xu, Peiran, et al.
Pubblicazione: (2025)
di: Xu, Peiran, et al.
Pubblicazione: (2025)
DynaNav: Dynamic Feature and Layer Selection for Efficient Visual Navigation
di: Wang, Jiahui, et al.
Pubblicazione: (2025)
di: Wang, Jiahui, et al.
Pubblicazione: (2025)
PIG-Nav: Key Insights for Pretrained Image Goal Navigation Models
di: Wan, Jiansong, et al.
Pubblicazione: (2025)
di: Wan, Jiansong, et al.
Pubblicazione: (2025)
Multi-Level Knowledge Distillation and Dynamic Self-Supervised Learning for Continual Learning
di: Kim, Taeheon, et al.
Pubblicazione: (2025)
di: Kim, Taeheon, et al.
Pubblicazione: (2025)
OctoNav: Towards Generalist Embodied Navigation
di: Gao, Chen, et al.
Pubblicazione: (2025)
di: Gao, Chen, et al.
Pubblicazione: (2025)
Documenti analoghi
-
Robust Image Self-Recovery against Tampering using Watermark Generation with Pixel Shuffling
di: Kim, Minyoung, et al.
Pubblicazione: (2025) -
Attention Misses Visual Risk: Risk-Adaptive Steering for Multimodal Safety Alignment
di: Park, Jonghyun, et al.
Pubblicazione: (2025) -
Image Diffusion Models Exhibit Emergent Temporal Propagation in Videos
di: Kim, Youngseo, et al.
Pubblicazione: (2025) -
V$^2$Dial: Unification of Video and Visual Dialog via Multimodal Experts
di: Abdessaied, Adnen, et al.
Pubblicazione: (2025) -
Multi-Granularity Video Object Segmentation
di: Lim, Sangbeom, et al.
Pubblicazione: (2024)