VELMA: Verbalization Embodiment of LLM Agents for Vision and Language Navigation in Street View
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Schumann, Raphael, Zhu, Wanrong, Feng, Weixi, Fu, Tsu-Jui, Riezler, Stefan, Wang, William Yang |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2023
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Text-to-OverpassQL: A Natural Language Interface for Complex Geodata Querying of OpenStreetMap
von: Staniek, Michael, et al.
Veröffentlicht: (2023)
von: Staniek, Michael, et al.
Veröffentlicht: (2023)
Boosting Process-Correct CoT Reasoning by Modeling Solvability of Multiple-Choice QA
von: Schumann, Raphael, et al.
Veröffentlicht: (2025)
von: Schumann, Raphael, et al.
Veröffentlicht: (2025)
Mango: Multi-Agent Web Navigation via Global-View Optimization
von: Tong, Weixi, et al.
Veröffentlicht: (2026)
von: Tong, Weixi, et al.
Veröffentlicht: (2026)
TC-Bench: Benchmarking Temporal Compositionality in Text-to-Video and Image-to-Video Generation
von: Feng, Weixi, et al.
Veröffentlicht: (2024)
von: Feng, Weixi, et al.
Veröffentlicht: (2024)
Prompting Large Language Models with Human Error Markings for Self-Correcting Machine Translation
von: Berger, Nathaniel, et al.
Veröffentlicht: (2024)
von: Berger, Nathaniel, et al.
Veröffentlicht: (2024)
Training and Evaluation of Guideline-Based Medical Reasoning in LLMs
von: Staniek, Michael, et al.
Veröffentlicht: (2025)
von: Staniek, Michael, et al.
Veröffentlicht: (2025)
From Text to Pixel: Advancing Long-Context Understanding in MLLMs
von: Lu, Yujie, et al.
Veröffentlicht: (2024)
von: Lu, Yujie, et al.
Veröffentlicht: (2024)
Counterfactual Vision-and-Language Navigation via Adversarial Path Sampling
von: Fu, Tsu-Jui, et al.
Veröffentlicht: (2019)
von: Fu, Tsu-Jui, et al.
Veröffentlicht: (2019)
Verbal Process Supervision Elicits Better Coding Agents
von: Chen, Hao-Yuan, et al.
Veröffentlicht: (2025)
von: Chen, Hao-Yuan, et al.
Veröffentlicht: (2025)
Do Visual Imaginations Improve Vision-and-Language Navigation Agents?
von: Perincherry, Akhil, et al.
Veröffentlicht: (2025)
von: Perincherry, Akhil, et al.
Veröffentlicht: (2025)
Discffusion: Discriminative Diffusion Models as Few-shot Vision and Language Learners
von: He, Xuehai, et al.
Veröffentlicht: (2023)
von: He, Xuehai, et al.
Veröffentlicht: (2023)
The Effects of Embodiment and Personality Expression on Learning in LLM-based Educational Agents
von: Sonlu, Sinan, et al.
Veröffentlicht: (2024)
von: Sonlu, Sinan, et al.
Veröffentlicht: (2024)
Are LLM Decisions Faithful to Verbal Confidence?
von: Wang, Jiawei, et al.
Veröffentlicht: (2026)
von: Wang, Jiawei, et al.
Veröffentlicht: (2026)
Learning to Translate Ambiguous Terminology by Preference Optimization on Post-Edits
von: Berger, Nathaniel, et al.
Veröffentlicht: (2025)
von: Berger, Nathaniel, et al.
Veröffentlicht: (2025)
GROKE: Vision-Free Navigation Instruction Evaluation via Graph Reasoning on OpenStreetMap
von: Shami, Farzad, et al.
Veröffentlicht: (2026)
von: Shami, Farzad, et al.
Veröffentlicht: (2026)
NavAgent: Multi-scale Urban Street View Fusion For UAV Embodied Vision-and-Language Navigation
von: Liu, Youzhi, et al.
Veröffentlicht: (2024)
von: Liu, Youzhi, et al.
Veröffentlicht: (2024)
Post-edits Are Preferences Too
von: Berger, Nathaniel, et al.
Veröffentlicht: (2024)
von: Berger, Nathaniel, et al.
Veröffentlicht: (2024)
NavHint: Vision and Language Navigation Agent with a Hint Generator
von: Zhang, Yue, et al.
Veröffentlicht: (2024)
von: Zhang, Yue, et al.
Veröffentlicht: (2024)
Direct Confidence Alignment: Aligning Verbalized Confidence with Internal Confidence In Large Language Models
von: Zhang, Glenn, et al.
Veröffentlicht: (2025)
von: Zhang, Glenn, et al.
Veröffentlicht: (2025)
T2V-Turbo: Breaking the Quality Bottleneck of Video Consistency Model with Mixed Reward Feedback
von: Li, Jiachen, et al.
Veröffentlicht: (2024)
von: Li, Jiachen, et al.
Veröffentlicht: (2024)
Generative Digital Twins: Vision-Language Simulation Models for Executable Industrial Systems
von: Hsu, YuChe, et al.
Veröffentlicht: (2025)
von: Hsu, YuChe, et al.
Veröffentlicht: (2025)
Large Language Models Are Latent Variable Models: Explaining and Finding Good Demonstrations for In-Context Learning
von: Wang, Xinyi, et al.
Veröffentlicht: (2023)
von: Wang, Xinyi, et al.
Veröffentlicht: (2023)
Overconfidence is Key: Verbalized Uncertainty Evaluation in Large Language and Vision-Language Models
von: Groot, Tobias, et al.
Veröffentlicht: (2024)
von: Groot, Tobias, et al.
Veröffentlicht: (2024)
Automatic Layout Planning for Visually-Rich Documents with Instruction-Following Models
von: Zhu, Wanrong, et al.
Veröffentlicht: (2024)
von: Zhu, Wanrong, et al.
Veröffentlicht: (2024)
On the Robustness of Verbal Confidence of LLMs in Adversarial Attacks
von: Obadinma, Stephen, et al.
Veröffentlicht: (2025)
von: Obadinma, Stephen, et al.
Veröffentlicht: (2025)
Qwen-VLA: Unifying Vision-Language-Action Modeling across Tasks, Environments, and Robot Embodiments
von: Wang, Qiuyue, et al.
Veröffentlicht: (2026)
von: Wang, Qiuyue, et al.
Veröffentlicht: (2026)
FinCon: A Synthesized LLM Multi-Agent System with Conceptual Verbal Reinforcement for Enhanced Financial Decision Making
von: Yu, Yangyang, et al.
Veröffentlicht: (2024)
von: Yu, Yangyang, et al.
Veröffentlicht: (2024)
What Limits Vision-and-Language Navigation ?
von: Wang, Yunheng, et al.
Veröffentlicht: (2026)
von: Wang, Yunheng, et al.
Veröffentlicht: (2026)
On Verbalized Confidence Scores for LLMs
von: Yang, Daniel, et al.
Veröffentlicht: (2024)
von: Yang, Daniel, et al.
Veröffentlicht: (2024)
LASER: LLM Agent with State-Space Exploration for Web Navigation
von: Ma, Kaixin, et al.
Veröffentlicht: (2023)
von: Ma, Kaixin, et al.
Veröffentlicht: (2023)
CodeJudge: Evaluating Code Generation with Large Language Models
von: Tong, Weixi, et al.
Veröffentlicht: (2024)
von: Tong, Weixi, et al.
Veröffentlicht: (2024)
Navigating Beyond Instructions: Vision-and-Language Navigation in Obstructed Environments
von: Hong, Haodong, et al.
Veröffentlicht: (2024)
von: Hong, Haodong, et al.
Veröffentlicht: (2024)
Probing the Category of Verbal Aspect in Transformer Language Models
von: Katinskaia, Anisia, et al.
Veröffentlicht: (2024)
von: Katinskaia, Anisia, et al.
Veröffentlicht: (2024)
MMWorld: Towards Multi-discipline Multi-faceted World Model Evaluation in Videos
von: He, Xuehai, et al.
Veröffentlicht: (2024)
von: He, Xuehai, et al.
Veröffentlicht: (2024)
What if LLMs Have Different World Views: Simulating Alien Civilizations with LLM-based Agents
von: Xue, Zhaoqian, et al.
Veröffentlicht: (2024)
von: Xue, Zhaoqian, et al.
Veröffentlicht: (2024)
VTS-LLM: Domain-Adaptive LLM Agent for Enhancing Awareness in Vessel Traffic Services through Natural Language
von: Sun, Sijin, et al.
Veröffentlicht: (2025)
von: Sun, Sijin, et al.
Veröffentlicht: (2025)
Breaking Down and Building Up: Mixture of Skill-Based Vision-and-Language Navigation Agents
von: Ma, Tianyi, et al.
Veröffentlicht: (2025)
von: Ma, Tianyi, et al.
Veröffentlicht: (2025)
Attacking Vision-Language Computer Agents via Pop-ups
von: Zhang, Yanzhe, et al.
Veröffentlicht: (2024)
von: Zhang, Yanzhe, et al.
Veröffentlicht: (2024)
Are Large Language Models More Honest in Their Probabilistic or Verbalized Confidence?
von: Ni, Shiyu, et al.
Veröffentlicht: (2024)
von: Ni, Shiyu, et al.
Veröffentlicht: (2024)
Leveraging Language Models and Machine Learning in Verbal Autopsy Analysis
von: Chu, Yue
Veröffentlicht: (2025)
von: Chu, Yue
Veröffentlicht: (2025)
Ähnliche Einträge
-
Text-to-OverpassQL: A Natural Language Interface for Complex Geodata Querying of OpenStreetMap
von: Staniek, Michael, et al.
Veröffentlicht: (2023) -
Boosting Process-Correct CoT Reasoning by Modeling Solvability of Multiple-Choice QA
von: Schumann, Raphael, et al.
Veröffentlicht: (2025) -
Mango: Multi-Agent Web Navigation via Global-View Optimization
von: Tong, Weixi, et al.
Veröffentlicht: (2026) -
TC-Bench: Benchmarking Temporal Compositionality in Text-to-Video and Image-to-Video Generation
von: Feng, Weixi, et al.
Veröffentlicht: (2024) -
Prompting Large Language Models with Human Error Markings for Self-Correcting Machine Translation
von: Berger, Nathaniel, et al.
Veröffentlicht: (2024)