Continual Vision-and-Language Navigation
Fuente:
arXiv
Saved in:
| Main Authors: | Jeong, Seongjun, Kang, Gi-Cheon, Choi, Seongho, Kim, Joochan, Zhang, Byoung-Tak |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Zero-Shot Vision-and-Language Navigation with Collision Mitigation in Continuous Environment
by: Jeong, Seongjun, et al.
Published: (2024)
by: Jeong, Seongjun, et al.
Published: (2024)
Exploring Ordinal Bias in Action Recognition for Instructional Videos
by: Kim, Joochan, et al.
Published: (2025)
by: Kim, Joochan, et al.
Published: (2025)
Background-aware Moment Detection for Video Moment Retrieval
by: Jung, Minjoon, et al.
Published: (2023)
by: Jung, Minjoon, et al.
Published: (2023)
CogME: A Cognition-Inspired Multi-Dimensional Evaluation Metric for Story Understanding
by: Shin, Minjung, et al.
Published: (2021)
by: Shin, Minjung, et al.
Published: (2021)
Learning Coordinate-based Convolutional Kernels for Continuous SE(3) Equivariant and Efficient Point Cloud Analysis
by: Kim, Jaein, et al.
Published: (2026)
by: Kim, Jaein, et al.
Published: (2026)
Socratic Planner: Self-QA-Based Zero-Shot Planning for Embodied Instruction Following
by: Shin, Suyeon, et al.
Published: (2024)
by: Shin, Suyeon, et al.
Published: (2024)
OCK: Unsupervised Dynamic Video Prediction with Object-Centric Kinematics
by: Song, Yeon-Ji, et al.
Published: (2024)
by: Song, Yeon-Ji, et al.
Published: (2024)
Surface-Based Visibility-Guided Uncertainty for Continuous Active 3D Neural Reconstruction
by: Kim, Hyunseo, et al.
Published: (2024)
by: Kim, Hyunseo, et al.
Published: (2024)
OV-MAP : Open-Vocabulary Zero-Shot 3D Instance Segmentation Map for Robots
by: Kim, Juno, et al.
Published: (2025)
by: Kim, Juno, et al.
Published: (2025)
EgoExo-Con: Exploring View-Invariant Video Temporal Understanding
by: Jung, Minjoon, et al.
Published: (2025)
by: Jung, Minjoon, et al.
Published: (2025)
Demystifying KAN for Vision Tasks: The RepKAN Approach
by: Cheon, Minjong
Published: (2026)
by: Cheon, Minjong
Published: (2026)
OBSER: Object-Based Sub-Environment Recognition for Zero-Shot Environmental Inference
by: Choi, Won-Seok, et al.
Published: (2025)
by: Choi, Won-Seok, et al.
Published: (2025)
UnitedVLN: Generalizable Gaussian Splatting for Continuous Vision-Language Navigation
by: Dai, Guangzhao, et al.
Published: (2024)
by: Dai, Guangzhao, et al.
Published: (2024)
PGA: Personalizing Grasping Agents with Single Human-Robot Interaction
by: Kim, Junghyun, et al.
Published: (2023)
by: Kim, Junghyun, et al.
Published: (2023)
Instruction-tuned Self-Questioning Framework for Multimodal Reasoning
by: Jang, You-Won, et al.
Published: (2025)
by: Jang, You-Won, et al.
Published: (2025)
3D-Aware Vision-Language Models Fine-Tuning with Geometric Distillation
by: Lee, Seonho, et al.
Published: (2025)
by: Lee, Seonho, et al.
Published: (2025)
Don't Miss the Forest for the Trees: Attentional Vision Calibration for Large Vision Language Models
by: Woo, Sangmin, et al.
Published: (2024)
by: Woo, Sangmin, et al.
Published: (2024)
Learning to Retrieve Navigable Candidates for Efficient Vision-and-Language Navigation
by: Gu, Shutian, et al.
Published: (2026)
by: Gu, Shutian, et al.
Published: (2026)
SIA: Enhancing Safety via Intent Awareness for Vision-Language Models
by: Na, Youngjin, et al.
Published: (2025)
by: Na, Youngjin, et al.
Published: (2025)
Demonstrating the Efficacy of Kolmogorov-Arnold Networks in Vision Tasks
by: Cheon, Minjong
Published: (2024)
by: Cheon, Minjong
Published: (2024)
SiNGER: A Clearer Voice Distills Vision Transformers Further
by: Yu, Geunhyeok, et al.
Published: (2025)
by: Yu, Geunhyeok, et al.
Published: (2025)
Multimodal Distribution Matching for Vision-Language Dataset Distillation
by: Jeong, Jongoh, et al.
Published: (2026)
by: Jeong, Jongoh, et al.
Published: (2026)
Fine-Tuning Vision-Language Models for Visual Navigation Assistance
by: Li, Xiao, et al.
Published: (2025)
by: Li, Xiao, et al.
Published: (2025)
VLM-HOI: Vision Language Models for Interpretable Human-Object Interaction Analysis
by: Kang, Donggoo, et al.
Published: (2024)
by: Kang, Donggoo, et al.
Published: (2024)
LightZeroNav: Zero-Shot Vision Language Navigation in Continuous Environments Based on Lightweight VLMs
by: Luo, Kun, et al.
Published: (2026)
by: Luo, Kun, et al.
Published: (2026)
Ground-level Viewpoint Vision-and-Language Navigation in Continuous Environments
by: Li, Zerui, et al.
Published: (2025)
by: Li, Zerui, et al.
Published: (2025)
Vision-and-Language Navigation via Causal Learning
by: Wang, Liuyi, et al.
Published: (2024)
by: Wang, Liuyi, et al.
Published: (2024)
In-situ and Non-contact Etch Depth Prediction in Plasma Etching via Machine Learning (ANN & BNN) and Digital Image Colorimetry
by: Kang, Minji, et al.
Published: (2025)
by: Kang, Minji, et al.
Published: (2025)
Are Large Vision-Language Models Ready to Guide Blind and Low-Vision Individuals?
by: Kim, Eunki, et al.
Published: (2025)
by: Kim, Eunki, et al.
Published: (2025)
KAN-CL: Per-Knot Importance Regularization for Continual Learning with Kolmogorov-Arnold Networks
by: Cheon, Minjong
Published: (2026)
by: Cheon, Minjong
Published: (2026)
Advancing Meteorological Forecasting: AI-based Approach to Synoptic Weather Map Analysis
by: Choi, Yo-Hwan, et al.
Published: (2024)
by: Choi, Yo-Hwan, et al.
Published: (2024)
\textsc{NaVIDA}: Vision-Language Navigation with Inverse Dynamics Augmentation
by: Zhu, Weiye, et al.
Published: (2026)
by: Zhu, Weiye, et al.
Published: (2026)
Thinking Diffusion: Penalize and Guide Visual-Grounded Reasoning in Diffusion Multimodal Language Models
by: Kim, Keuntae, et al.
Published: (2026)
by: Kim, Keuntae, et al.
Published: (2026)
How Blind and Low-Vision Individuals Prefer Large Vision-Language Model-Generated Scene Descriptions
by: An, Na Min, et al.
Published: (2025)
by: An, Na Min, et al.
Published: (2025)
Vision-and-Language Navigation with Analogical Textual Descriptions in LLMs
by: Zhang, Yue, et al.
Published: (2025)
by: Zhang, Yue, et al.
Published: (2025)
RITUAL: Random Image Transformations as a Universal Anti-hallucination Lever in Large Vision Language Models
by: Woo, Sangmin, et al.
Published: (2024)
by: Woo, Sangmin, et al.
Published: (2024)
To Ask or Not to Ask? Detecting Absence of Information in Vision and Language Navigation
by: Abraham, Savitha Sam, et al.
Published: (2024)
by: Abraham, Savitha Sam, et al.
Published: (2024)
Actional Atomic-Concept Learning for Demystifying Vision-Language Navigation
by: Lin, Bingqian, et al.
Published: (2023)
by: Lin, Bingqian, et al.
Published: (2023)
Dual-Anchoring: Addressing State Drift in Vision-Language Navigation
by: Wu, Kangyi, et al.
Published: (2026)
by: Wu, Kangyi, et al.
Published: (2026)
Your Large Vision-Language Model Only Needs A Few Attention Heads For Visual Grounding
by: Kang, Seil, et al.
Published: (2025)
by: Kang, Seil, et al.
Published: (2025)
Similar Items
-
Zero-Shot Vision-and-Language Navigation with Collision Mitigation in Continuous Environment
by: Jeong, Seongjun, et al.
Published: (2024) -
Exploring Ordinal Bias in Action Recognition for Instructional Videos
by: Kim, Joochan, et al.
Published: (2025) -
Background-aware Moment Detection for Video Moment Retrieval
by: Jung, Minjoon, et al.
Published: (2023) -
CogME: A Cognition-Inspired Multi-Dimensional Evaluation Metric for Story Understanding
by: Shin, Minjung, et al.
Published: (2021) -
Learning Coordinate-based Convolutional Kernels for Continuous SE(3) Equivariant and Efficient Point Cloud Analysis
by: Kim, Jaein, et al.
Published: (2026)