TINA: Think, Interaction, and Action Framework for Zero-Shot Vision Language Navigation
Fuente:
arXiv
Salvato in:
| Autori principali: | Li, Dingbang, Chen, Wenzhou, Lin, Xin |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Think, Act, Build: An Agentic Framework with Vision Language Models for Zero-Shot 3D Visual Grounding
di: Wang, Haibo, et al.
Pubblicazione: (2026)
di: Wang, Haibo, et al.
Pubblicazione: (2026)
Zero-Shot Vision-and-Language Navigation with Collision Mitigation in Continuous Environment
di: Jeong, Seongjun, et al.
Pubblicazione: (2024)
di: Jeong, Seongjun, et al.
Pubblicazione: (2024)
LightZeroNav: Zero-Shot Vision Language Navigation in Continuous Environments Based on Lightweight VLMs
di: Luo, Kun, et al.
Pubblicazione: (2026)
di: Luo, Kun, et al.
Pubblicazione: (2026)
SpatialNav: Leveraging Spatial Scene Graphs for Zero-Shot Vision-and-Language Navigation
di: Zhang, Jiwen, et al.
Pubblicazione: (2026)
di: Zhang, Jiwen, et al.
Pubblicazione: (2026)
DreamNav: A Trajectory-Based Imaginative Framework for Zero-Shot Vision-and-Language Navigation
di: Wang, Yunheng, et al.
Pubblicazione: (2025)
di: Wang, Yunheng, et al.
Pubblicazione: (2025)
Navigating the Trade-off: A Synthesis of Defensive Strategies for Zero-Shot Adversarial Robustness in Vision-Language Models
di: Xu, Zane, et al.
Pubblicazione: (2025)
di: Xu, Zane, et al.
Pubblicazione: (2025)
CorNav: Autonomous Agent with Self-Corrected Planning for Zero-Shot Vision-and-Language Navigation
di: Liang, Xiwen, et al.
Pubblicazione: (2023)
di: Liang, Xiwen, et al.
Pubblicazione: (2023)
ZSPAPrune: Zero-Shot Prompt-Aware Token Pruning for Vision-Language Models
di: Zhang, Pu, et al.
Pubblicazione: (2025)
di: Zhang, Pu, et al.
Pubblicazione: (2025)
P2DNav: Panorama-to-Downview Reasoning for Zero-shot Vision-and-Language Navigation
di: Sheng, Kai, et al.
Pubblicazione: (2026)
di: Sheng, Kai, et al.
Pubblicazione: (2026)
Spatio-Temporal Context Prompting for Zero-Shot Action Detection
di: Huang, Wei-Jhe, et al.
Pubblicazione: (2024)
di: Huang, Wei-Jhe, et al.
Pubblicazione: (2024)
Benchmarking Zero-Shot Recognition with Vision-Language Models: Challenges on Granularity and Specificity
di: Xu, Zhenlin, et al.
Pubblicazione: (2023)
di: Xu, Zhenlin, et al.
Pubblicazione: (2023)
Binary Verification for Zero-Shot Vision
di: Hu, Rongbin, et al.
Pubblicazione: (2025)
di: Hu, Rongbin, et al.
Pubblicazione: (2025)
Think Hierarchically, Act Dynamically: Hierarchical Multi-modal Fusion and Reasoning for Vision-and-Language Navigation
di: Yue, Junrong, et al.
Pubblicazione: (2025)
di: Yue, Junrong, et al.
Pubblicazione: (2025)
VGAS: Value-Guided Action-Chunk Selection for Few-Shot Vision-Language-Action Adaptation
di: Xu, Changhua, et al.
Pubblicazione: (2026)
di: Xu, Changhua, et al.
Pubblicazione: (2026)
To Move or Not to Move: Constraint-based Planning Enables Zero-Shot Generalization for Interactive Navigation
di: Vashisth, Apoorva, et al.
Pubblicazione: (2026)
di: Vashisth, Apoorva, et al.
Pubblicazione: (2026)
VLM4VLA: Revisiting Vision-Language-Models in Vision-Language-Action Models
di: Zhang, Jianke, et al.
Pubblicazione: (2026)
di: Zhang, Jianke, et al.
Pubblicazione: (2026)
LLM meets Vision-Language Models for Zero-Shot One-Class Classification
di: Bendou, Yassir, et al.
Pubblicazione: (2024)
di: Bendou, Yassir, et al.
Pubblicazione: (2024)
Frequency-Semantic Enhanced Variational Autoencoder for Zero-Shot Skeleton-based Action Recognition
di: Wu, Wenhan, et al.
Pubblicazione: (2025)
di: Wu, Wenhan, et al.
Pubblicazione: (2025)
Schrödinger's Navigator: Imagining an Ensemble of Futures for Zero-Shot Object Navigation
di: He, Yu, et al.
Pubblicazione: (2025)
di: He, Yu, et al.
Pubblicazione: (2025)
TransVLM: A Vision-Language Framework and Benchmark for Detecting Any Shot Transitions
di: Chen, Ce, et al.
Pubblicazione: (2026)
di: Chen, Ce, et al.
Pubblicazione: (2026)
Dual-Anchoring: Addressing State Drift in Vision-Language Navigation
di: Wu, Kangyi, et al.
Pubblicazione: (2026)
di: Wu, Kangyi, et al.
Pubblicazione: (2026)
Bad Seeing or Bad Thinking? Rewarding Perception for Vision-Language Reasoning
di: Wang, Haozhe, et al.
Pubblicazione: (2026)
di: Wang, Haozhe, et al.
Pubblicazione: (2026)
Frequency-Enhanced Diffusion Models: Curriculum-Guided Semantic Alignment for Zero-Shot Skeleton Action Recognition
di: Zhou, Yuxi, et al.
Pubblicazione: (2026)
di: Zhou, Yuxi, et al.
Pubblicazione: (2026)
Actional Atomic-Concept Learning for Demystifying Vision-Language Navigation
di: Lin, Bingqian, et al.
Pubblicazione: (2023)
di: Lin, Bingqian, et al.
Pubblicazione: (2023)
VisualThink-VLA: Visual Intermediate Reasoning for Effective and Low-Latency Vision-Language-Action Policies
di: Gao, Mingjian, et al.
Pubblicazione: (2026)
di: Gao, Mingjian, et al.
Pubblicazione: (2026)
MAGIC: Meta-Ability Guided Interactive Chain-of-Distillation for Effective-and-Efficient Vision-and-Language Navigation
di: Wang, Liuyi, et al.
Pubblicazione: (2024)
di: Wang, Liuyi, et al.
Pubblicazione: (2024)
Text-Guided Attention is All You Need for Zero-Shot Robustness in Vision-Language Models
di: Yu, Lu, et al.
Pubblicazione: (2024)
di: Yu, Lu, et al.
Pubblicazione: (2024)
Evaluating Vision-Language Models for Zero-Shot Detection, Classification, and Association of Motorcycles, Passengers, and Helmets
di: Choi, Lucas, et al.
Pubblicazione: (2024)
di: Choi, Lucas, et al.
Pubblicazione: (2024)
VLAgeBench: Benchmarking Large Vision-Language Models for Zero-Shot Human Age Estimation
di: Sajib, Rakib Hossain, et al.
Pubblicazione: (2026)
di: Sajib, Rakib Hossain, et al.
Pubblicazione: (2026)
VLA-Thinker: Boosting Vision-Language-Action Models through Thinking-with-Image Reasoning
di: Wang, Chaoyang, et al.
Pubblicazione: (2026)
di: Wang, Chaoyang, et al.
Pubblicazione: (2026)
Continual Vision-and-Language Navigation
di: Jeong, Seongjun, et al.
Pubblicazione: (2024)
di: Jeong, Seongjun, et al.
Pubblicazione: (2024)
Think or Not? Selective Reasoning via Reinforcement Learning for Vision-Language Models
di: Wang, Jiaqi, et al.
Pubblicazione: (2025)
di: Wang, Jiaqi, et al.
Pubblicazione: (2025)
A Navigation Framework Utilizing Vision-Language Models
di: Duan, Yicheng, et al.
Pubblicazione: (2025)
di: Duan, Yicheng, et al.
Pubblicazione: (2025)
TinyVLM: Zero-Shot Object Detection on Microcontrollers via Vision-Language Distillation with Matryoshka Embeddings
di: Wilson, Bibin
Pubblicazione: (2026)
di: Wilson, Bibin
Pubblicazione: (2026)
HeatPrompt: Zero-Shot Vision-Language Modeling of Urban Heat Demand from Satellite Images
di: Thota, Kundan, et al.
Pubblicazione: (2026)
di: Thota, Kundan, et al.
Pubblicazione: (2026)
Vision-and-Language Navigation via Causal Learning
di: Wang, Liuyi, et al.
Pubblicazione: (2024)
di: Wang, Liuyi, et al.
Pubblicazione: (2024)
Zero-Shot Action Generalization with Limited Observations
di: Alchihabi, Abdullah, et al.
Pubblicazione: (2025)
di: Alchihabi, Abdullah, et al.
Pubblicazione: (2025)
doScenes: An Autonomous Driving Dataset with Natural Language Instruction for Human Interaction and Vision-Language Navigation
di: Roy, Parthib, et al.
Pubblicazione: (2024)
di: Roy, Parthib, et al.
Pubblicazione: (2024)
Human-Aware Vision-and-Language Navigation: Bridging Simulation to Reality with Dynamic Human Interactions
di: Li, Heng, et al.
Pubblicazione: (2024)
di: Li, Heng, et al.
Pubblicazione: (2024)
Learning to Retrieve Navigable Candidates for Efficient Vision-and-Language Navigation
di: Gu, Shutian, et al.
Pubblicazione: (2026)
di: Gu, Shutian, et al.
Pubblicazione: (2026)
Documenti analoghi
-
Think, Act, Build: An Agentic Framework with Vision Language Models for Zero-Shot 3D Visual Grounding
di: Wang, Haibo, et al.
Pubblicazione: (2026) -
Zero-Shot Vision-and-Language Navigation with Collision Mitigation in Continuous Environment
di: Jeong, Seongjun, et al.
Pubblicazione: (2024) -
LightZeroNav: Zero-Shot Vision Language Navigation in Continuous Environments Based on Lightweight VLMs
di: Luo, Kun, et al.
Pubblicazione: (2026) -
SpatialNav: Leveraging Spatial Scene Graphs for Zero-Shot Vision-and-Language Navigation
di: Zhang, Jiwen, et al.
Pubblicazione: (2026) -
DreamNav: A Trajectory-Based Imaginative Framework for Zero-Shot Vision-and-Language Navigation
di: Wang, Yunheng, et al.
Pubblicazione: (2025)