PanoGen++: Domain-Adapted Text-Guided Panoramic Environment Generation for Vision-and-Language Navigation
Fuente:
arXiv
Saved in:
| Main Authors: | Wang, Sen, Zhou, Dongliang, Xie, Liang, Xu, Chao, Yan, Ye, Yin, Erwei |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Safe-VLN: Collision Avoidance for Vision-and-Language Navigation of Autonomous Robots Operating in Continuous Environments
by: Yue, Lu, et al.
Published: (2023)
by: Yue, Lu, et al.
Published: (2023)
LipGen: Viseme-Guided Lip Video Generation for Enhancing Visual Speech Recognition
by: Hao, Bowen, et al.
Published: (2025)
by: Hao, Bowen, et al.
Published: (2025)
ST-Booster: An Iterative SpatioTemporal Perception Booster for Vision-and-Language Navigation in Continuous Environments
by: Yue, Lu, et al.
Published: (2025)
by: Yue, Lu, et al.
Published: (2025)
SIG-Chat: Spatial Intent-Guided Conversational Gesture Generation Involving How, When and Where
by: Huang, Yiheng, et al.
Published: (2025)
by: Huang, Yiheng, et al.
Published: (2025)
DA-PTQ: Drift-Aware Post-Training Quantization for Efficient Vision-Language-Action Models
by: Xu, Siyuan, et al.
Published: (2026)
by: Xu, Siyuan, et al.
Published: (2026)
AC^2-VLA: Action-Context-Aware Adaptive Computation in Vision-Language-Action Models for Efficient Robotic Manipulation
by: Yu, Wenda, et al.
Published: (2026)
by: Yu, Wenda, et al.
Published: (2026)
WaterVG: Waterway Visual Grounding based on Text-Guided Vision and mmWave Radar
by: Guan, Runwei, et al.
Published: (2024)
by: Guan, Runwei, et al.
Published: (2024)
AVE Speech: A Comprehensive Multi-Modal Dataset for Speech Recognition Integrating Audio, Visual, and Electromyographic Signals
by: Zhou, Dongliang, et al.
Published: (2025)
by: Zhou, Dongliang, et al.
Published: (2025)
Landmark-Guided Cross-Speaker Lip Reading with Mutual Information Regularization
by: Wu, Linzhi, et al.
Published: (2024)
by: Wu, Linzhi, et al.
Published: (2024)
RoboTron-Mani: All-in-One Multimodal Large Model for Robotic Manipulation
by: Yan, Feng, et al.
Published: (2024)
by: Yan, Feng, et al.
Published: (2024)
XEmbodied: A Foundation Model with Enhanced Geometric and Physical Cues for Large-Scale Embodied Environments
by: Qian, Kangan, et al.
Published: (2026)
by: Qian, Kangan, et al.
Published: (2026)
PanoDP: Learning Collision-Free Navigation with Panoramic Depth and Differentiable Physics
by: Zhong, Hao, et al.
Published: (2026)
by: Zhong, Hao, et al.
Published: (2026)
CineWild: Balancing Art and Robotics for Ethical Wildlife Documentary Filmmaking
by: Pueyo, Pablo, et al.
Published: (2025)
by: Pueyo, Pablo, et al.
Published: (2025)
Bringing Robots Home: The Rise of AI Robots in Consumer Electronics
by: Dong, Haiwei, et al.
Published: (2024)
by: Dong, Haiwei, et al.
Published: (2024)
A Multimedia Framework for Continuum Robots: Systematic, Computational, and Control Perspectives
by: Hsieh, Po-Yu, et al.
Published: (2024)
by: Hsieh, Po-Yu, et al.
Published: (2024)
COPA: Efficient Vision-Language Pre-training Through Collaborative Object- and Patch-Text Alignment
by: Jiang, Chaoya, et al.
Published: (2023)
by: Jiang, Chaoya, et al.
Published: (2023)
Learning Primitive Embodied World Models: Towards Scalable Robotic Learning
by: Sun, Qiao, et al.
Published: (2025)
by: Sun, Qiao, et al.
Published: (2025)
BLM$_1$: A Boundless Large Model for Cross-Space, Cross-Task, and Cross-Embodiment Learning
by: Tan, Wentao, et al.
Published: (2025)
by: Tan, Wentao, et al.
Published: (2025)
Chain-of-Modality: Learning Manipulation Programs from Multimodal Human Videos with Vision-Language-Models
by: Wang, Chen, et al.
Published: (2025)
by: Wang, Chen, et al.
Published: (2025)
RoboTron-Drive: All-in-One Large Multimodal Model for Autonomous Driving
by: Huang, Zhijian, et al.
Published: (2024)
by: Huang, Zhijian, et al.
Published: (2024)
MIDI-LLM: Adapting Large Language Models for Text-to-MIDI Music Generation
by: Wu, Shih-Lun, et al.
Published: (2025)
by: Wu, Shih-Lun, et al.
Published: (2025)
MotiBo: The Impact of Interactive Digital Storytelling Robots on Student Motivation through Self-Determination Theory
by: Fung, Ka Yan, et al.
Published: (2026)
by: Fung, Ka Yan, et al.
Published: (2026)
InstruGen: Automatic Instruction Generation for Vision-and-Language Navigation Via Large Multimodal Models
by: Yan, Yu, et al.
Published: (2024)
by: Yan, Yu, et al.
Published: (2024)
Flight Patterns for Swarms of Drones
by: Zhu, Shuqin, et al.
Published: (2024)
by: Zhu, Shuqin, et al.
Published: (2024)
WildFusion: Multimodal Implicit 3D Reconstructions in the Wild
by: Liu, Yanbaihui, et al.
Published: (2024)
by: Liu, Yanbaihui, et al.
Published: (2024)
Hybrid Feedback-Guided Optimal Learning for Wireless Interactive Panoramic Scene Delivery
by: Wu, Xiaoyi, et al.
Published: (2026)
by: Wu, Xiaoyi, et al.
Published: (2026)
Listen, Look, Drive: Coupling Audio Instructions for User-aware VLA-based Autonomous Driving
by: Guo, Ziang, et al.
Published: (2026)
by: Guo, Ziang, et al.
Published: (2026)
E-VLA: Event-Augmented Vision-Language-Action Model for Dark and Blurred Scenes
by: Zhai, Jiajun, et al.
Published: (2026)
by: Zhai, Jiajun, et al.
Published: (2026)
UniScene: Multi-Camera Unified Pre-training via 3D Scene Reconstruction for Autonomous Driving
by: Min, Chen, et al.
Published: (2023)
by: Min, Chen, et al.
Published: (2023)
NavigScene: Bridging Local Perception and Global Navigation for Beyond-Visual-Range Autonomous Driving
by: Peng, Qucheng, et al.
Published: (2025)
by: Peng, Qucheng, et al.
Published: (2025)
DreamFoley: Scalable VLMs for High-Fidelity Video-to-Audio Generation
by: Li, Fu, et al.
Published: (2025)
by: Li, Fu, et al.
Published: (2025)
SceneDreamer360: Text-Driven 3D-Consistent Scene Generation with Panoramic Gaussian Splatting
by: Li, Wenrui, et al.
Published: (2024)
by: Li, Wenrui, et al.
Published: (2024)
SurgSora: Object-Aware Diffusion Model for Controllable Surgical Video Generation
by: Chen, Tong, et al.
Published: (2024)
by: Chen, Tong, et al.
Published: (2024)
Physics-Aware Novel-View Acoustic Synthesis with Vision-Language Priors and 3D Acoustic Environment Modeling
by: Fan, Congyi, et al.
Published: (2026)
by: Fan, Congyi, et al.
Published: (2026)
Video-Guided Text-to-Music Generation Using Public Domain Movie Collections
by: Kim, Haven, et al.
Published: (2025)
by: Kim, Haven, et al.
Published: (2025)
Evaluating Magic Leap 2 Tool Tracking for AR Sensor Guidance in Industrial Inspections
by: Masuhr, Christian, et al.
Published: (2025)
by: Masuhr, Christian, et al.
Published: (2025)
TimeSpot: Benchmarking Geo-Temporal Understanding in Vision-Language Models in Real-World Settings
by: Wasi, Azmine Toushik, et al.
Published: (2026)
by: Wasi, Azmine Toushik, et al.
Published: (2026)
Exploring Event-based Human Pose Estimation with 3D Event Representations
by: Yin, Xiaoting, et al.
Published: (2023)
by: Yin, Xiaoting, et al.
Published: (2023)
PanoNav: Mapless Zero-Shot Object Navigation with Panoramic Scene Parsing and Dynamic Memory
by: Jin, Qunchao, et al.
Published: (2025)
by: Jin, Qunchao, et al.
Published: (2025)
OnomatoGen: Onomatopoeia Generation with the Alpha-Channel in Manga
by: Taniguchi, Takara, et al.
Published: (2025)
by: Taniguchi, Takara, et al.
Published: (2025)
Similar Items
-
Safe-VLN: Collision Avoidance for Vision-and-Language Navigation of Autonomous Robots Operating in Continuous Environments
by: Yue, Lu, et al.
Published: (2023) -
LipGen: Viseme-Guided Lip Video Generation for Enhancing Visual Speech Recognition
by: Hao, Bowen, et al.
Published: (2025) -
ST-Booster: An Iterative SpatioTemporal Perception Booster for Vision-and-Language Navigation in Continuous Environments
by: Yue, Lu, et al.
Published: (2025) -
SIG-Chat: Spatial Intent-Guided Conversational Gesture Generation Involving How, When and Where
by: Huang, Yiheng, et al.
Published: (2025) -
DA-PTQ: Drift-Aware Post-Training Quantization for Efficient Vision-Language-Action Models
by: Xu, Siyuan, et al.
Published: (2026)