ESceme: Vision-and-Language Navigation with Episodic Scene Memory
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Zheng, Qi, Liu, Daqing, Wang, Chaoyue, Zhang, Jing, Wang, Dadong, Tao, Dacheng |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2023
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Visual Superordinate Abstraction for Robust Concept Learning
von: Zheng, Qi, et al.
Veröffentlicht: (2022)
von: Zheng, Qi, et al.
Veröffentlicht: (2022)
Planning from Imagination: Episodic Simulation and Episodic Memory for Vision-and-Language Navigation
von: Pan, Yiyuan, et al.
Veröffentlicht: (2024)
von: Pan, Yiyuan, et al.
Veröffentlicht: (2024)
Cocktail: Mixing Multi-Modality Controls for Text-Conditional Image Generation
von: Hu, Minghui, et al.
Veröffentlicht: (2023)
von: Hu, Minghui, et al.
Veröffentlicht: (2023)
HandRefiner: Refining Malformed Hands in Generated Images by Diffusion-based Conditional Inpainting
von: Lu, Wenquan, et al.
Veröffentlicht: (2023)
von: Lu, Wenquan, et al.
Veröffentlicht: (2023)
General Scene Adaptation for Vision-and-Language Navigation
von: Hong, Haodong, et al.
Veröffentlicht: (2025)
von: Hong, Haodong, et al.
Veröffentlicht: (2025)
Towards Modality-agnostic Label-efficient Segmentation with Entropy-Regularized Distribution Alignment
von: Tang, Liyao, et al.
Veröffentlicht: (2024)
von: Tang, Liyao, et al.
Veröffentlicht: (2024)
COSMO: Combination of Selective Memorization for Low-cost Vision-and-Language Navigation
von: Zhang, Siqi, et al.
Veröffentlicht: (2025)
von: Zhang, Siqi, et al.
Veröffentlicht: (2025)
MiniVLN: Efficient Vision-and-Language Navigation by Progressive Knowledge Distillation
von: Zhu, Junyou, et al.
Veröffentlicht: (2024)
von: Zhu, Junyou, et al.
Veröffentlicht: (2024)
EAGLE: Episodic Appearance- and Geometry-aware Memory for Unified 2D-3D Visual Query Localization in Egocentric Vision
von: Cao, Yifei, et al.
Veröffentlicht: (2025)
von: Cao, Yifei, et al.
Veröffentlicht: (2025)
Trajectory Consistency Distillation: Improved Latent Consistency Distillation by Semi-Linear Consistency Function with Trajectory Mapping
von: Zheng, Jianbin, et al.
Veröffentlicht: (2024)
von: Zheng, Jianbin, et al.
Veröffentlicht: (2024)
GPT4Scene: Understand 3D Scenes from Videos with Vision-Language Models
von: Qi, Zhangyang, et al.
Veröffentlicht: (2025)
von: Qi, Zhangyang, et al.
Veröffentlicht: (2025)
When ControlNet Meets Inexplicit Masks: A Case Study of ControlNet on its Contour-following Ability
von: Xuan, Wenjie, et al.
Veröffentlicht: (2024)
von: Xuan, Wenjie, et al.
Veröffentlicht: (2024)
Scene-aware Human Motion Forecasting via Mutual Distance Prediction
von: Xing, Chaoyue, et al.
Veröffentlicht: (2023)
von: Xing, Chaoyue, et al.
Veröffentlicht: (2023)
SpatialNav: Leveraging Spatial Scene Graphs for Zero-Shot Vision-and-Language Navigation
von: Zhang, Jiwen, et al.
Veröffentlicht: (2026)
von: Zhang, Jiwen, et al.
Veröffentlicht: (2026)
Navigating Beyond Instructions: Vision-and-Language Navigation in Obstructed Environments
von: Hong, Haodong, et al.
Veröffentlicht: (2024)
von: Hong, Haodong, et al.
Veröffentlicht: (2024)
InterPhys: Physics-aware Human Motion Synthesis in a Dynamic Scene
von: Xing, Chaoyue, et al.
Veröffentlicht: (2026)
von: Xing, Chaoyue, et al.
Veröffentlicht: (2026)
MSNav: Zero-Shot Vision-and-Language Navigation with Dynamic Memory and LLM Spatial Reasoning
von: Liu, Chenghao, et al.
Veröffentlicht: (2025)
von: Liu, Chenghao, et al.
Veröffentlicht: (2025)
EGOSTREAM: A Diagnostic Benchmark for Streaming Episodic Memory in Egocentric Vision
von: Forte, Rosario, et al.
Veröffentlicht: (2026)
von: Forte, Rosario, et al.
Veröffentlicht: (2026)
Advances in Radiance Field for Dynamic Scene: From Neural Field to Gaussian Field
von: Fan, Jinlong, et al.
Veröffentlicht: (2025)
von: Fan, Jinlong, et al.
Veröffentlicht: (2025)
LAB-Det: Language as a Domain-Invariant Bridge for Training-Free One-Shot Domain Generalization in Object Detection
von: Zhang, Xu, et al.
Veröffentlicht: (2026)
von: Zhang, Xu, et al.
Veröffentlicht: (2026)
Volumetric Environment Representation for Vision-Language Navigation
von: Liu, Rui, et al.
Veröffentlicht: (2024)
von: Liu, Rui, et al.
Veröffentlicht: (2024)
Vision-Language Navigation with Energy-Based Policy
von: Liu, Rui, et al.
Veröffentlicht: (2024)
von: Liu, Rui, et al.
Veröffentlicht: (2024)
On Geometry-Enhanced Parameter-Efficient Fine-Tuning for 3D Scene Segmentation
von: Tang, Liyao, et al.
Veröffentlicht: (2025)
von: Tang, Liyao, et al.
Veröffentlicht: (2025)
DivScene: Towards Open-Vocabulary Object Navigation with Large Vision Language Models in Diverse Scenes
von: Wang, Zhaowei, et al.
Veröffentlicht: (2024)
von: Wang, Zhaowei, et al.
Veröffentlicht: (2024)
Natural Reflection Backdoor Attack on Vision Language Model for Autonomous Driving
von: Liu, Ming, et al.
Veröffentlicht: (2025)
von: Liu, Ming, et al.
Veröffentlicht: (2025)
VLN-R1: Vision-Language Navigation via Reinforcement Fine-Tuning
von: Qi, Zhangyang, et al.
Veröffentlicht: (2025)
von: Qi, Zhangyang, et al.
Veröffentlicht: (2025)
FlexVLN: Flexible Adaptation for Diverse Vision-and-Language Navigation Tasks
von: Zhang, Siqi, et al.
Veröffentlicht: (2025)
von: Zhang, Siqi, et al.
Veröffentlicht: (2025)
Interpretable Zero-shot Referring Expression Comprehension with Query-driven Scene Graphs
von: Wu, Yike, et al.
Veröffentlicht: (2026)
von: Wu, Yike, et al.
Veröffentlicht: (2026)
Perceptual Taxonomy: Evaluating and Guiding Hierarchical Scene Reasoning in Vision-Language Models
von: Lee, Jonathan, et al.
Veröffentlicht: (2025)
von: Lee, Jonathan, et al.
Veröffentlicht: (2025)
World-Consistent Data Generation for Vision-and-Language Navigation
von: Zhong, Yu, et al.
Veröffentlicht: (2024)
von: Zhong, Yu, et al.
Veröffentlicht: (2024)
Review of Hallucination Understanding in Large Language and Vision Models
von: Ho, Zhengyi, et al.
Veröffentlicht: (2025)
von: Ho, Zhengyi, et al.
Veröffentlicht: (2025)
GeoZero: Incentivizing Reasoning from Scratch on Geospatial Scenes
von: Wang, Di, et al.
Veröffentlicht: (2025)
von: Wang, Di, et al.
Veröffentlicht: (2025)
Deep Learning for Event-based Vision: A Comprehensive Survey and Benchmarks
von: Zheng, Xu, et al.
Veröffentlicht: (2023)
von: Zheng, Xu, et al.
Veröffentlicht: (2023)
Efficient-VLN: A Training-Efficient Vision-Language Navigation Model
von: Zheng, Duo, et al.
Veröffentlicht: (2025)
von: Zheng, Duo, et al.
Veröffentlicht: (2025)
Uncertainty-Aware Gaussian Map for Vision-Language Navigation
von: Gao, Jianzhe, et al.
Veröffentlicht: (2026)
von: Gao, Jianzhe, et al.
Veröffentlicht: (2026)
Interactive Episodic Memory with User Feedback
von: Subedi, Nikesh, et al.
Veröffentlicht: (2026)
von: Subedi, Nikesh, et al.
Veröffentlicht: (2026)
Predicate Debiasing in Vision-Language Models Integration for Scene Graph Generation Enhancement
von: Wang, Yuxuan, et al.
Veröffentlicht: (2024)
von: Wang, Yuxuan, et al.
Veröffentlicht: (2024)
Vision-Language Memory for Spatial Reasoning
von: Liu, Zuntao, et al.
Veröffentlicht: (2025)
von: Liu, Zuntao, et al.
Veröffentlicht: (2025)
The Essence of Balance for Self-Improving Agents in Vision-and-Language Navigation
von: Liu, Zhen, et al.
Veröffentlicht: (2026)
von: Liu, Zhen, et al.
Veröffentlicht: (2026)
OSGNet @ Ego4D Episodic Memory Challenge 2025
von: Feng, Yisen, et al.
Veröffentlicht: (2025)
von: Feng, Yisen, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Visual Superordinate Abstraction for Robust Concept Learning
von: Zheng, Qi, et al.
Veröffentlicht: (2022) -
Planning from Imagination: Episodic Simulation and Episodic Memory for Vision-and-Language Navigation
von: Pan, Yiyuan, et al.
Veröffentlicht: (2024) -
Cocktail: Mixing Multi-Modality Controls for Text-Conditional Image Generation
von: Hu, Minghui, et al.
Veröffentlicht: (2023) -
HandRefiner: Refining Malformed Hands in Generated Images by Diffusion-based Conditional Inpainting
von: Lu, Wenquan, et al.
Veröffentlicht: (2023) -
General Scene Adaptation for Vision-and-Language Navigation
von: Hong, Haodong, et al.
Veröffentlicht: (2025)