World-Consistent Data Generation for Vision-and-Language Navigation
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Zhong, Yu, Zhang, Rui, Zhang, Zihao, Wang, Shuo, Fang, Chuan, Zhang, Xishan, Guo, Jiaming, Peng, Shaohui, Huang, Di, Yan, Yanyang, Hu, Xing, Guo, Qi |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Run, Ruminate, and Regulate: A Dual-process Thinking System for Vision-and-Language Navigation
von: Zhong, Yu, et al.
Veröffentlicht: (2025)
von: Zhong, Yu, et al.
Veröffentlicht: (2025)
Ex3: Automatic Novel Writing by Extracting, Excelsior and Expanding
von: Huang, Lei, et al.
Veröffentlicht: (2024)
von: Huang, Lei, et al.
Veröffentlicht: (2024)
Efficient Diffusion Planning with Temporal Diffusion
von: Guo, Jiaming, et al.
Veröffentlicht: (2025)
von: Guo, Jiaming, et al.
Veröffentlicht: (2025)
QiMeng-CodeV-SVA: Training Specialized LLMs for Hardware Assertion Generation via RTL-Grounded Bidirectional Data Synthesis
von: Wu, Yutong, et al.
Veröffentlicht: (2026)
von: Wu, Yutong, et al.
Veröffentlicht: (2026)
Code Driven Planning with Domain-Adaptive Critic
von: Tian, Zikang, et al.
Veröffentlicht: (2025)
von: Tian, Zikang, et al.
Veröffentlicht: (2025)
Luban: Building Open-Ended Creative Agents via Autonomous Embodied Verification
von: Guo, Yuxuan, et al.
Veröffentlicht: (2024)
von: Guo, Yuxuan, et al.
Veröffentlicht: (2024)
Assessing and Understanding Creativity in Large Language Models
von: Zhao, Yunpu, et al.
Veröffentlicht: (2024)
von: Zhao, Yunpu, et al.
Veröffentlicht: (2024)
PrimitiveVLA: Learning Reusable Motion Primitives for Efficient and Generalizable Robotic Manipulation
von: Li, Yutai, et al.
Veröffentlicht: (2026)
von: Li, Yutai, et al.
Veröffentlicht: (2026)
Object-Level Verbalized Confidence Calibration in Vision-Language Models via Semantic Perturbation
von: Zhao, Yunpu, et al.
Veröffentlicht: (2025)
von: Zhao, Yunpu, et al.
Veröffentlicht: (2025)
NavHint: Vision and Language Navigation Agent with a Hint Generator
von: Zhang, Yue, et al.
Veröffentlicht: (2024)
von: Zhang, Yue, et al.
Veröffentlicht: (2024)
Sparse Video Generation Propels Real-World Beyond-the-View Vision-Language Navigation
von: Zhang, Hai, et al.
Veröffentlicht: (2026)
von: Zhang, Hai, et al.
Veröffentlicht: (2026)
WMNav: Integrating Vision-Language Models into World Models for Object Goal Navigation
von: Nie, Dujun, et al.
Veröffentlicht: (2025)
von: Nie, Dujun, et al.
Veröffentlicht: (2025)
Efficient-VLN: A Training-Efficient Vision-Language Navigation Model
von: Zheng, Duo, et al.
Veröffentlicht: (2025)
von: Zheng, Duo, et al.
Veröffentlicht: (2025)
Prompt-based Visual Alignment for Zero-shot Policy Transfer
von: Gao, Haihan, et al.
Veröffentlicht: (2024)
von: Gao, Haihan, et al.
Veröffentlicht: (2024)
IDCNet: Guided Video Diffusion for Metric-Consistent RGBD Scene Generation with Precise Camera Control
von: Liu, Lijuan, et al.
Veröffentlicht: (2025)
von: Liu, Lijuan, et al.
Veröffentlicht: (2025)
GTPC-SSCD: Gate-guided Two-level Perturbation Consistency-based Semi-Supervised Change Detection
von: Xing, Yan, et al.
Veröffentlicht: (2024)
von: Xing, Yan, et al.
Veröffentlicht: (2024)
Towards Open Environments and Instructions: General Vision-Language Navigation via Fast-Slow Interactive Reasoning
von: Li, Yang, et al.
Veröffentlicht: (2026)
von: Li, Yang, et al.
Veröffentlicht: (2026)
QiMeng-CRUX: Narrowing the Gap Between Natural Language and Verilog via Core Refined Understanding eXpression for Circuit Design
von: Huang, Lei, et al.
Veröffentlicht: (2025)
von: Huang, Lei, et al.
Veröffentlicht: (2025)
SceneTAP: Scene-Coherent Typographic Adversarial Planner against Vision-Language Models in Real-World Environments
von: Cao, Yue, et al.
Veröffentlicht: (2024)
von: Cao, Yue, et al.
Veröffentlicht: (2024)
LCGNav: Local Candidate-Aware Geometric Enhancement for General Topological Planning in Vision-Language Navigation
von: Peng, Jiankun, et al.
Veröffentlicht: (2026)
von: Peng, Jiankun, et al.
Veröffentlicht: (2026)
QiMeng-PRepair: Precise Code Repair via Edit-Aware Reward Optimization
von: Ke, Changxin, et al.
Veröffentlicht: (2026)
von: Ke, Changxin, et al.
Veröffentlicht: (2026)
WorldVLN: Autoregressive World Action Model for Aerial Vision-Language Navigation
von: Zhao, Baining, et al.
Veröffentlicht: (2026)
von: Zhao, Baining, et al.
Veröffentlicht: (2026)
CodeV: Empowering LLMs with HDL Generation through Multi-Level Summarization
von: Zhao, Yang, et al.
Veröffentlicht: (2024)
von: Zhao, Yang, et al.
Veröffentlicht: (2024)
CAVALRY-V: A Large-Scale Generator Framework for Adversarial Attacks on Video MLLMs
von: Zhang, Jiaming, et al.
Veröffentlicht: (2025)
von: Zhang, Jiaming, et al.
Veröffentlicht: (2025)
QiMeng-SALV: Signal-Aware Learning for Verilog Code Generation
von: Zhang, Yang, et al.
Veröffentlicht: (2025)
von: Zhang, Yang, et al.
Veröffentlicht: (2025)
Linguistics-Vision Monotonic Consistent Network for Sign Language Production
von: Wang, Xu, et al.
Veröffentlicht: (2024)
von: Wang, Xu, et al.
Veröffentlicht: (2024)
Uncertainty-Aware Gaussian Map for Vision-Language Navigation
von: Gao, Jianzhe, et al.
Veröffentlicht: (2026)
von: Gao, Jianzhe, et al.
Veröffentlicht: (2026)
LKV: End-to-End Learning of Head-wise Budgets and Token Selection for LLM KV Cache Eviction
von: Zhou, Enshuai, et al.
Veröffentlicht: (2026)
von: Zhou, Enshuai, et al.
Veröffentlicht: (2026)
FeatNavigator: Automatic Feature Augmentation on Tabular Data
von: Liang, Jiaming, et al.
Veröffentlicht: (2024)
von: Liang, Jiaming, et al.
Veröffentlicht: (2024)
RealBench: Benchmarking Verilog Generation Models with Real-World IP Designs
von: Jin, Pengwei, et al.
Veröffentlicht: (2025)
von: Jin, Pengwei, et al.
Veröffentlicht: (2025)
Unveiling the Tapestry of Consistency in Large Vision-Language Models
von: Zhang, Yuan, et al.
Veröffentlicht: (2024)
von: Zhang, Yuan, et al.
Veröffentlicht: (2024)
UniGeoSeg: Towards Unified Open-World Segmentation for Geospatial Scenes
von: Ni, Shuo, et al.
Veröffentlicht: (2025)
von: Ni, Shuo, et al.
Veröffentlicht: (2025)
Stream-Omni: Simultaneous Multimodal Interactions with Large Language-Vision-Speech Model
von: Zhang, Shaolei, et al.
Veröffentlicht: (2025)
von: Zhang, Shaolei, et al.
Veröffentlicht: (2025)
SPAN-Nav: Generalized Spatial Awareness for Versatile Vision-Language Navigation
von: Liu, Jiahang, et al.
Veröffentlicht: (2026)
von: Liu, Jiahang, et al.
Veröffentlicht: (2026)
FlexVLN: Flexible Adaptation for Diverse Vision-and-Language Navigation Tasks
von: Zhang, Siqi, et al.
Veröffentlicht: (2025)
von: Zhang, Siqi, et al.
Veröffentlicht: (2025)
Disrupting Vision-Language Model-Driven Navigation Services via Adversarial Object Fusion
von: Xie, Chunlong, et al.
Veröffentlicht: (2025)
von: Xie, Chunlong, et al.
Veröffentlicht: (2025)
DiffZOO: A Purely Query-Based Black-Box Attack for Red-teaming Text-to-Image Generative Model via Zeroth Order Optimization
von: Dang, Pucheng, et al.
Veröffentlicht: (2024)
von: Dang, Pucheng, et al.
Veröffentlicht: (2024)
DynaMath: A Dynamic Visual Benchmark for Evaluating Mathematical Reasoning Robustness of Vision Language Models
von: Zou, Chengke, et al.
Veröffentlicht: (2024)
von: Zou, Chengke, et al.
Veröffentlicht: (2024)
Dynamic Topology Awareness: Breaking the Granularity Rigidity in Vision-Language Navigation
von: Peng, Jiankun, et al.
Veröffentlicht: (2026)
von: Peng, Jiankun, et al.
Veröffentlicht: (2026)
An Improved Recursive Algorithm for V-BLAST to Save Memories without Sacrificing Speed
von: Zhu, Hufei, et al.
Veröffentlicht: (2023)
von: Zhu, Hufei, et al.
Veröffentlicht: (2023)
Ähnliche Einträge
-
Run, Ruminate, and Regulate: A Dual-process Thinking System for Vision-and-Language Navigation
von: Zhong, Yu, et al.
Veröffentlicht: (2025) -
Ex3: Automatic Novel Writing by Extracting, Excelsior and Expanding
von: Huang, Lei, et al.
Veröffentlicht: (2024) -
Efficient Diffusion Planning with Temporal Diffusion
von: Guo, Jiaming, et al.
Veröffentlicht: (2025) -
QiMeng-CodeV-SVA: Training Specialized LLMs for Hardware Assertion Generation via RTL-Grounded Bidirectional Data Synthesis
von: Wu, Yutong, et al.
Veröffentlicht: (2026) -
Code Driven Planning with Domain-Adaptive Critic
von: Tian, Zikang, et al.
Veröffentlicht: (2025)