Bootstrapping Language-Guided Navigation Learning with Self-Refining Data Flywheel
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Wang, Zun, Li, Jialu, Hong, Yicong, Li, Songze, Li, Kunchang, Yu, Shoubin, Wang, Yi, Qiao, Yu, Wang, Yali, Bansal, Mohit, Wang, Limin |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Learning Goal-Oriented Vision-and-Language Navigation with Self-Improving Demonstrations at Scale
von: Li, Songze, et al.
Veröffentlicht: (2025)
von: Li, Songze, et al.
Veröffentlicht: (2025)
SAME: Learning Generic Language-Guided Visual Navigation with State-Adaptive Mixture of Experts
von: Zhou, Gengze, et al.
Veröffentlicht: (2024)
von: Zhou, Gengze, et al.
Veröffentlicht: (2024)
VideoMamba: State Space Model for Efficient Video Understanding
von: Li, Kunchang, et al.
Veröffentlicht: (2024)
von: Li, Kunchang, et al.
Veröffentlicht: (2024)
Unmasked Teacher: Towards Training-Efficient Video Foundation Models
von: Li, Kunchang, et al.
Veröffentlicht: (2023)
von: Li, Kunchang, et al.
Veröffentlicht: (2023)
Vision-and-Language Navigation Today and Tomorrow: A Survey in the Era of Foundation Models
von: Zhang, Yue, et al.
Veröffentlicht: (2024)
von: Zhang, Yue, et al.
Veröffentlicht: (2024)
DreamRunner: Fine-Grained Compositional Story-to-Video Generation with Retrieval-Augmented Motion Adaptation
von: Wang, Zun, et al.
Veröffentlicht: (2024)
von: Wang, Zun, et al.
Veröffentlicht: (2024)
CREMA: Generalizable and Efficient Video-Language Reasoning via Multimodal Modular Fusion
von: Yu, Shoubin, et al.
Veröffentlicht: (2024)
von: Yu, Shoubin, et al.
Veröffentlicht: (2024)
Harvest Video Foundation Models via Efficient Post-Pretraining
von: Li, Yizhuo, et al.
Veröffentlicht: (2023)
von: Li, Yizhuo, et al.
Veröffentlicht: (2023)
MVBench: A Comprehensive Multi-modal Video Understanding Benchmark
von: Li, Kunchang, et al.
Veröffentlicht: (2023)
von: Li, Kunchang, et al.
Veröffentlicht: (2023)
TransAgent: Transfer Vision-Language Foundation Models with Heterogeneous Agent Collaboration
von: Guo, Yiwei, et al.
Veröffentlicht: (2024)
von: Guo, Yiwei, et al.
Veröffentlicht: (2024)
Training-free Guidance in Text-to-Video Generation via Multimodal Planning and Structured Noise Initialization
von: Li, Jialu, et al.
Veröffentlicht: (2025)
von: Li, Jialu, et al.
Veröffentlicht: (2025)
VLN-Video: Utilizing Driving Videos for Outdoor Vision-and-Language Navigation
von: Li, Jialu, et al.
Veröffentlicht: (2024)
von: Li, Jialu, et al.
Veröffentlicht: (2024)
When and How Much to Imagine: Adaptive Test-Time Scaling with World Models for Visual Spatial Reasoning
von: Yu, Shoubin, et al.
Veröffentlicht: (2026)
von: Yu, Shoubin, et al.
Veröffentlicht: (2026)
RACCooN: A Versatile Instructional Video Editing Framework with Auto-Generated Narratives
von: Yoon, Jaehong, et al.
Veröffentlicht: (2024)
von: Yoon, Jaehong, et al.
Veröffentlicht: (2024)
NavGPT-2: Unleashing Navigational Reasoning Capability for Large Vision-Language Models
von: Zhou, Gengze, et al.
Veröffentlicht: (2024)
von: Zhou, Gengze, et al.
Veröffentlicht: (2024)
TimeSuite: Improving MLLMs for Long Video Understanding via Grounded Tuning
von: Zeng, Xiangyu, et al.
Veröffentlicht: (2024)
von: Zeng, Xiangyu, et al.
Veröffentlicht: (2024)
Task Preference Optimization: Improving Multimodal Large Language Models with Vision Task Alignment
von: Yan, Ziang, et al.
Veröffentlicht: (2024)
von: Yan, Ziang, et al.
Veröffentlicht: (2024)
MEXA: Towards General Multimodal Reasoning with Dynamic Multi-Expert Aggregation
von: Yu, Shoubin, et al.
Veröffentlicht: (2025)
von: Yu, Shoubin, et al.
Veröffentlicht: (2025)
Error-Driven Scene Editing for 3D Grounding in Large Language Models
von: Zhang, Yue, et al.
Veröffentlicht: (2025)
von: Zhang, Yue, et al.
Veröffentlicht: (2025)
UniFormer: Unifying Convolution and Self-attention for Visual Recognition
von: Li, Kunchang, et al.
Veröffentlicht: (2022)
von: Li, Kunchang, et al.
Veröffentlicht: (2022)
InternVideo2: Scaling Foundation Models for Multimodal Video Understanding
von: Wang, Yi, et al.
Veröffentlicht: (2024)
von: Wang, Yi, et al.
Veröffentlicht: (2024)
VEGGIE: Instructional Editing and Reasoning Video Concepts with Grounded Generation
von: Yu, Shoubin, et al.
Veröffentlicht: (2025)
von: Yu, Shoubin, et al.
Veröffentlicht: (2025)
Percept, Chat, and then Adapt: Multimodal Knowledge Transfer of Foundation Models for Open-World Video Recognition
von: Chen, Boyu, et al.
Veröffentlicht: (2024)
von: Chen, Boyu, et al.
Veröffentlicht: (2024)
Super Encoding Network: Recursive Association of Multi-Modal Encoders for Video Understanding
von: Chen, Boyu, et al.
Veröffentlicht: (2025)
von: Chen, Boyu, et al.
Veröffentlicht: (2025)
EPiC: Efficient Video Camera Control Learning with Precise Anchor-Video Guidance
von: Wang, Zun, et al.
Veröffentlicht: (2025)
von: Wang, Zun, et al.
Veröffentlicht: (2025)
MUSES: 3D-Controllable Image Generation via Multi-Modal Agent Collaboration
von: Ding, Yanbo, et al.
Veröffentlicht: (2024)
von: Ding, Yanbo, et al.
Veröffentlicht: (2024)
VideoEval: Comprehensive Benchmark Suite for Low-Cost Evaluation of Video Foundation Model
von: Li, Xinhao, et al.
Veröffentlicht: (2024)
von: Li, Xinhao, et al.
Veröffentlicht: (2024)
VisionCoach: Reinforcing Grounded Video Reasoning via Visual-Perception Prompting
von: Lee, Daeun, et al.
Veröffentlicht: (2026)
von: Lee, Daeun, et al.
Veröffentlicht: (2026)
VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling
von: Li, Xinhao, et al.
Veröffentlicht: (2024)
von: Li, Xinhao, et al.
Veröffentlicht: (2024)
Make Your Training Flexible: Towards Deployment-Efficient Video Models
von: Wang, Chenting, et al.
Veröffentlicht: (2025)
von: Wang, Chenting, et al.
Veröffentlicht: (2025)
Video-RTS: Rethinking Reinforcement Learning and Test-Time Scaling for Efficient and Enhanced Video Reasoning
von: Wang, Ziyang, et al.
Veröffentlicht: (2025)
von: Wang, Ziyang, et al.
Veröffentlicht: (2025)
CorrectNav: Self-Correction Flywheel Empowers Vision-Language-Action Navigation Model
von: Yu, Zhuoyuan, et al.
Veröffentlicht: (2025)
von: Yu, Zhuoyuan, et al.
Veröffentlicht: (2025)
Vlogger: Make Your Dream A Vlog
von: Zhuang, Shaobin, et al.
Veröffentlicht: (2024)
von: Zhuang, Shaobin, et al.
Veröffentlicht: (2024)
VideoChat: Chat-Centric Video Understanding
von: Li, KunChang, et al.
Veröffentlicht: (2023)
von: Li, KunChang, et al.
Veröffentlicht: (2023)
SELMA: Learning and Merging Skill-Specific Text-to-Image Experts with Auto-Generated Data
von: Li, Jialu, et al.
Veröffentlicht: (2024)
von: Li, Jialu, et al.
Veröffentlicht: (2024)
Navigating the Data Trading Crossroads: An Interdisciplinary Survey
von: Yu, Yi, et al.
Veröffentlicht: (2024)
von: Yu, Yi, et al.
Veröffentlicht: (2024)
InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation
von: Wang, Yi, et al.
Veröffentlicht: (2023)
von: Wang, Yi, et al.
Veröffentlicht: (2023)
VideoChat-A1: Thinking with Long Videos by Chain-of-Shot Reasoning
von: Wang, Zikang, et al.
Veröffentlicht: (2025)
von: Wang, Zikang, et al.
Veröffentlicht: (2025)
VideoTree: Adaptive Tree-based Video Representation for LLM Reasoning on Long Videos
von: Wang, Ziyang, et al.
Veröffentlicht: (2024)
von: Wang, Ziyang, et al.
Veröffentlicht: (2024)
SAFREE: Training-Free and Adaptive Guard for Safe Text-to-Image And Video Generation
von: Yoon, Jaehong, et al.
Veröffentlicht: (2024)
von: Yoon, Jaehong, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Learning Goal-Oriented Vision-and-Language Navigation with Self-Improving Demonstrations at Scale
von: Li, Songze, et al.
Veröffentlicht: (2025) -
SAME: Learning Generic Language-Guided Visual Navigation with State-Adaptive Mixture of Experts
von: Zhou, Gengze, et al.
Veröffentlicht: (2024) -
VideoMamba: State Space Model for Efficient Video Understanding
von: Li, Kunchang, et al.
Veröffentlicht: (2024) -
Unmasked Teacher: Towards Training-Efficient Video Foundation Models
von: Li, Kunchang, et al.
Veröffentlicht: (2023) -
Vision-and-Language Navigation Today and Tomorrow: A Survey in the Era of Foundation Models
von: Zhang, Yue, et al.
Veröffentlicht: (2024)