RAP: Retrieval-Augmented Planner for Adaptive Procedure Planning in Instructional Videos
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Zare, Ali, Niu, Yulei, Ayyubi, Hammad, Chang, Shih-fu |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
SCHEMA: State CHangEs MAtter for Procedure Planning in Instructional Videos
von: Niu, Yulei, et al.
Veröffentlicht: (2024)
von: Niu, Yulei, et al.
Veröffentlicht: (2024)
SDA-PLANNER: State-Dependency Aware Adaptive Planner for Embodied Task Planning
von: Shen, Zichao, et al.
Veröffentlicht: (2025)
von: Shen, Zichao, et al.
Veröffentlicht: (2025)
PuzzleGPT: Emulating Human Puzzle-Solving Ability for Time and Location Prediction
von: Ayyubi, Hammad, et al.
Veröffentlicht: (2025)
von: Ayyubi, Hammad, et al.
Veröffentlicht: (2025)
RoomTour3D: Geometry-Aware Video-Instruction Tuning for Embodied Navigation
von: Han, Mingfei, et al.
Veröffentlicht: (2024)
von: Han, Mingfei, et al.
Veröffentlicht: (2024)
RAP: 3D Rasterization Augmented End-to-End Planning
von: Feng, Lan, et al.
Veröffentlicht: (2025)
von: Feng, Lan, et al.
Veröffentlicht: (2025)
RAG-3DSG: Enhancing 3D Scene Graphs with Re-Shot Guided Retrieval-Augmented Generation
von: Chang, Yue, et al.
Veröffentlicht: (2026)
von: Chang, Yue, et al.
Veröffentlicht: (2026)
Large Trajectory Models are Scalable Motion Predictors and Planners
von: Sun, Qiao, et al.
Veröffentlicht: (2023)
von: Sun, Qiao, et al.
Veröffentlicht: (2023)
Detecting Multimodal Situations with Insufficient Context and Abstaining from Baseless Predictions
von: Liu, Junzhang, et al.
Veröffentlicht: (2024)
von: Liu, Junzhang, et al.
Veröffentlicht: (2024)
What Matters for Scalable and Robust Learning in End-to-End Driving Planners?
von: Holtz, David, et al.
Veröffentlicht: (2026)
von: Holtz, David, et al.
Veröffentlicht: (2026)
RAAP: Retrieval-Augmented Affordance Prediction with Cross-Image Action Alignment
von: Zhuang, Qiyuan, et al.
Veröffentlicht: (2026)
von: Zhuang, Qiyuan, et al.
Veröffentlicht: (2026)
Video Summarization: Towards Entity-Aware Captions
von: Ayyubi, Hammad A., et al.
Veröffentlicht: (2023)
von: Ayyubi, Hammad A., et al.
Veröffentlicht: (2023)
OmniEVA: Embodied Versatile Planner via Task-Adaptive 3D-Grounded and Embodiment-aware Reasoning
von: Liu, Yuecheng, et al.
Veröffentlicht: (2025)
von: Liu, Yuecheng, et al.
Veröffentlicht: (2025)
This&That: Language-Gesture Controlled Video Generation for Robot Planning
von: Wang, Boyang, et al.
Veröffentlicht: (2024)
von: Wang, Boyang, et al.
Veröffentlicht: (2024)
NovaPlan: Zero-Shot Long-Horizon Manipulation via Closed-Loop Video Language Planning
von: Fu, Jiahui, et al.
Veröffentlicht: (2026)
von: Fu, Jiahui, et al.
Veröffentlicht: (2026)
ENTER: Event Based Interpretable Reasoning for VideoQA
von: Ayyubi, Hammad, et al.
Veröffentlicht: (2025)
von: Ayyubi, Hammad, et al.
Veröffentlicht: (2025)
Towards Unified World Models for Visual Navigation via Memory-Augmented Planning and Foresight
von: Dong, Yifei, et al.
Veröffentlicht: (2025)
von: Dong, Yifei, et al.
Veröffentlicht: (2025)
WIDIn: Wording Image for Domain-Invariant Representation in Single-Source Domain Generalization
von: Ma, Jiawei, et al.
Veröffentlicht: (2024)
von: Ma, Jiawei, et al.
Veröffentlicht: (2024)
MapGPT: Map-Guided Prompting with Adaptive Path Planning for Vision-and-Language Navigation
von: Chen, Jiaqi, et al.
Veröffentlicht: (2024)
von: Chen, Jiaqi, et al.
Veröffentlicht: (2024)
RoboVIP: Multi-View Video Generation with Visual Identity Prompting Augments Robot Manipulation
von: Wang, Boyang, et al.
Veröffentlicht: (2026)
von: Wang, Boyang, et al.
Veröffentlicht: (2026)
Efficient Robotic Policy Learning via Latent Space Backward Planning
von: Liu, Dongxiu, et al.
Veröffentlicht: (2025)
von: Liu, Dongxiu, et al.
Veröffentlicht: (2025)
Robotic Visual Instruction
von: Li, Yanbang, et al.
Veröffentlicht: (2025)
von: Li, Yanbang, et al.
Veröffentlicht: (2025)
RDD: Retrieval-Based Demonstration Decomposer for Planner Alignment in Long-Horizon Tasks
von: Yan, Mingxuan, et al.
Veröffentlicht: (2025)
von: Yan, Mingxuan, et al.
Veröffentlicht: (2025)
Neural MP: A Generalist Neural Motion Planner
von: Dalal, Murtaza, et al.
Veröffentlicht: (2024)
von: Dalal, Murtaza, et al.
Veröffentlicht: (2024)
VLA-Pro: Cross-Task Procedural Memory Transfer for Vision-Language-Action Models
von: Si, Shengyu, et al.
Veröffentlicht: (2026)
von: Si, Shengyu, et al.
Veröffentlicht: (2026)
Unseen from Seen: Rewriting Observation-Instruction Using Foundation Models for Augmenting Vision-Language Navigation
von: Wei, Ziming, et al.
Veröffentlicht: (2025)
von: Wei, Ziming, et al.
Veröffentlicht: (2025)
MotIF: Motion Instruction Fine-tuning
von: Hwang, Minyoung, et al.
Veröffentlicht: (2024)
von: Hwang, Minyoung, et al.
Veröffentlicht: (2024)
Vega: Learning to Drive with Natural Language Instructions
von: Zuo, Sicheng, et al.
Veröffentlicht: (2026)
von: Zuo, Sicheng, et al.
Veröffentlicht: (2026)
LIAM: Multimodal Transformer for Language Instructions, Images, Actions and Semantic Maps
von: Wang, Yihao, et al.
Veröffentlicht: (2025)
von: Wang, Yihao, et al.
Veröffentlicht: (2025)
VideoVLA: Video Generators Can Be Generalizable Robot Manipulators
von: Shen, Yichao, et al.
Veröffentlicht: (2025)
von: Shen, Yichao, et al.
Veröffentlicht: (2025)
IKEA Manuals at Work: 4D Grounding of Assembly Instructions on Internet Videos
von: Liu, Yunong, et al.
Veröffentlicht: (2024)
von: Liu, Yunong, et al.
Veröffentlicht: (2024)
SLAG: Scalable Language-Augmented Gaussian Splatting
von: Szilagyi, Laszlo, et al.
Veröffentlicht: (2025)
von: Szilagyi, Laszlo, et al.
Veröffentlicht: (2025)
Grounding 3D Object Affordance with Language Instructions, Visual Observations and Interactions
von: Zhu, He, et al.
Veröffentlicht: (2025)
von: Zhu, He, et al.
Veröffentlicht: (2025)
RAP: Retrieval-Augmented Personalization for Multimodal Large Language Models
von: Hao, Haoran, et al.
Veröffentlicht: (2024)
von: Hao, Haoran, et al.
Veröffentlicht: (2024)
VideoArtGS: Building Digital Twins of Articulated Objects from Monocular Video
von: Liu, Yu, et al.
Veröffentlicht: (2025)
von: Liu, Yu, et al.
Veröffentlicht: (2025)
Planning with the Views via Scene Self-Exploration
von: Wang, Kangrui, et al.
Veröffentlicht: (2026)
von: Wang, Kangrui, et al.
Veröffentlicht: (2026)
RAPID: Robust and Agile Planner Using Inverse Reinforcement Learning for Vision-Based Drone Navigation
von: Kim, Minwoo, et al.
Veröffentlicht: (2025)
von: Kim, Minwoo, et al.
Veröffentlicht: (2025)
Less is More: Label-Guided Summarization of Procedural and Instructional Videos
von: Rajpal, Shreya, et al.
Veröffentlicht: (2026)
von: Rajpal, Shreya, et al.
Veröffentlicht: (2026)
Sparse Imagination for Efficient Visual World Model Planning
von: Chun, Junha, et al.
Veröffentlicht: (2025)
von: Chun, Junha, et al.
Veröffentlicht: (2025)
Robix: A Unified Model for Robot Interaction, Reasoning and Planning
von: Fang, Huang, et al.
Veröffentlicht: (2025)
von: Fang, Huang, et al.
Veröffentlicht: (2025)
Gondola: Grounded Vision Language Planning for Generalizable Robotic Manipulation
von: Chen, Shizhe, et al.
Veröffentlicht: (2025)
von: Chen, Shizhe, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
SCHEMA: State CHangEs MAtter for Procedure Planning in Instructional Videos
von: Niu, Yulei, et al.
Veröffentlicht: (2024) -
SDA-PLANNER: State-Dependency Aware Adaptive Planner for Embodied Task Planning
von: Shen, Zichao, et al.
Veröffentlicht: (2025) -
PuzzleGPT: Emulating Human Puzzle-Solving Ability for Time and Location Prediction
von: Ayyubi, Hammad, et al.
Veröffentlicht: (2025) -
RoomTour3D: Geometry-Aware Video-Instruction Tuning for Embodied Navigation
von: Han, Mingfei, et al.
Veröffentlicht: (2024) -
RAP: 3D Rasterization Augmented End-to-End Planning
von: Feng, Lan, et al.
Veröffentlicht: (2025)