Can Language Beat Numerical Regression? Language-Based Multimodal Trajectory Prediction
Fuente:
arXiv
Salvato in:
| Autori principali: | Bae, Inhwan, Lee, Junoh, Jeon, Hae-Gon |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Continuous Locomotive Crowd Behavior Generation
di: Bae, Inhwan, et al.
Pubblicazione: (2025)
di: Bae, Inhwan, et al.
Pubblicazione: (2025)
SingularTrajectory: Universal Trajectory Predictor Using Diffusion Model
di: Bae, Inhwan, et al.
Pubblicazione: (2024)
di: Bae, Inhwan, et al.
Pubblicazione: (2024)
Depth Prompting for Sensor-Agnostic Depth Estimation
di: Park, Jin-Hwi, et al.
Pubblicazione: (2024)
di: Park, Jin-Hwi, et al.
Pubblicazione: (2024)
Fully Explicit Dynamic Gaussian Splatting
di: Lee, Junoh, et al.
Pubblicazione: (2024)
di: Lee, Junoh, et al.
Pubblicazione: (2024)
Relaxed Rigidity with Ray-based Grouping for Dynamic Gaussian Splatting
di: Lee, Junoh, et al.
Pubblicazione: (2026)
di: Lee, Junoh, et al.
Pubblicazione: (2026)
ComPose: When to Trust Hands for Object Pose Tracking
di: Shin, Jisu, et al.
Pubblicazione: (2026)
di: Shin, Jisu, et al.
Pubblicazione: (2026)
Kinetic Typography Diffusion Model
di: Park, Seonmi, et al.
Pubblicazione: (2024)
di: Park, Seonmi, et al.
Pubblicazione: (2024)
Signs of Language: Embodied Sign Language Fingerspelling Acquisition from Demonstrations for Human-Robot Interaction
di: Tavella, Federico, et al.
Pubblicazione: (2022)
di: Tavella, Federico, et al.
Pubblicazione: (2022)
SpatialVLM: Endowing Vision-Language Models with Spatial Reasoning Capabilities
di: Chen, Boyuan, et al.
Pubblicazione: (2024)
di: Chen, Boyuan, et al.
Pubblicazione: (2024)
GenSim: Generating Robotic Simulation Tasks via Large Language Models
di: Wang, Lirui, et al.
Pubblicazione: (2023)
di: Wang, Lirui, et al.
Pubblicazione: (2023)
Diffusion-Based Environment-Aware Trajectory Prediction
di: Westny, Theodor, et al.
Pubblicazione: (2024)
di: Westny, Theodor, et al.
Pubblicazione: (2024)
Uncertainty-Aware Diffusion Model for Multimodal Highway Trajectory Prediction via DDIM Sampling
di: Neumeier, Marion, et al.
Pubblicazione: (2026)
di: Neumeier, Marion, et al.
Pubblicazione: (2026)
Logic-RAG: Augmenting Large Multimodal Models with Visual-Spatial Knowledge for Road Scene Understanding
di: Kabir, Imran, et al.
Pubblicazione: (2025)
di: Kabir, Imran, et al.
Pubblicazione: (2025)
ProIn: Learning to Predict Trajectory Based on Progressive Interactions for Autonomous Driving
di: Dong, Yinke, et al.
Pubblicazione: (2024)
di: Dong, Yinke, et al.
Pubblicazione: (2024)
Latent Action Pretraining from Videos
di: Ye, Seonghyeon, et al.
Pubblicazione: (2024)
di: Ye, Seonghyeon, et al.
Pubblicazione: (2024)
Video-Language Critic: Transferable Reward Functions for Language-Conditioned Robotics
di: Alakuijala, Minttu, et al.
Pubblicazione: (2024)
di: Alakuijala, Minttu, et al.
Pubblicazione: (2024)
LanguageMPC: Large Language Models as Decision Makers for Autonomous Driving
di: Sha, Hao, et al.
Pubblicazione: (2023)
di: Sha, Hao, et al.
Pubblicazione: (2023)
Machine Learning-Based Vehicle Intention Trajectory Recognition and Prediction for Autonomous Driving
di: Yu, Hanyi, et al.
Pubblicazione: (2024)
di: Yu, Hanyi, et al.
Pubblicazione: (2024)
LangGap: Diagnosing and Closing the Language Gap in Vision-Language-Action Models
di: Hou, Yuchen, et al.
Pubblicazione: (2026)
di: Hou, Yuchen, et al.
Pubblicazione: (2026)
Holistic Evaluation of Multimodal LLMs on Spatial Intelligence
di: Cai, Zhongang, et al.
Pubblicazione: (2025)
di: Cai, Zhongang, et al.
Pubblicazione: (2025)
LLaRA: Supercharging Robot Learning Data for Vision-Language Policy
di: Li, Xiang, et al.
Pubblicazione: (2024)
di: Li, Xiang, et al.
Pubblicazione: (2024)
MAPS: Preserving Vision-Language Representations via Module-Wise Proximity Scheduling for Better Vision-Language-Action Generalization
di: Huang, Chengyue, et al.
Pubblicazione: (2025)
di: Huang, Chengyue, et al.
Pubblicazione: (2025)
Textual Steering Vectors Can Improve Visual Understanding in Multimodal Large Language Models
di: Gan, Woody Haosheng, et al.
Pubblicazione: (2025)
di: Gan, Woody Haosheng, et al.
Pubblicazione: (2025)
LOC: A General Language-Guided Framework for Open-Set 3D Occupancy Prediction
di: Gao, Yuhang, et al.
Pubblicazione: (2025)
di: Gao, Yuhang, et al.
Pubblicazione: (2025)
Producing and Leveraging Online Map Uncertainty in Trajectory Prediction
di: Gu, Xunjiang, et al.
Pubblicazione: (2024)
di: Gu, Xunjiang, et al.
Pubblicazione: (2024)
Knowledge-aware Graph Transformer for Pedestrian Trajectory Prediction
di: Liu, Yu, et al.
Pubblicazione: (2024)
di: Liu, Yu, et al.
Pubblicazione: (2024)
TRAVEL: Training-Free Retrieval and Alignment for Vision-and-Language Navigation
di: Rajabi, Navid, et al.
Pubblicazione: (2025)
di: Rajabi, Navid, et al.
Pubblicazione: (2025)
Pedestrian Trajectory Prediction with Missing Data: Datasets, Imputation, and Benchmarking
di: Chib, Pranav Singh, et al.
Pubblicazione: (2024)
di: Chib, Pranav Singh, et al.
Pubblicazione: (2024)
Vision-based Multi-future Trajectory Prediction: A Survey
di: Huang, Renhao, et al.
Pubblicazione: (2023)
di: Huang, Renhao, et al.
Pubblicazione: (2023)
Teaching Embodied Reinforcement Learning Agents: Informativeness and Diversity of Language Use
di: Xi, Jiajun, et al.
Pubblicazione: (2024)
di: Xi, Jiajun, et al.
Pubblicazione: (2024)
Meanings and Measurements: Multi-Agent Probabilistic Grounding for Vision-Language Navigation
di: Padhan, Swagat, et al.
Pubblicazione: (2026)
di: Padhan, Swagat, et al.
Pubblicazione: (2026)
Distilled Feature Fields Enable Few-Shot Language-Guided Manipulation
di: Shen, William, et al.
Pubblicazione: (2023)
di: Shen, William, et al.
Pubblicazione: (2023)
PIVOT: Iterative Visual Prompting Elicits Actionable Knowledge for VLMs
di: Nasiriany, Soroush, et al.
Pubblicazione: (2024)
di: Nasiriany, Soroush, et al.
Pubblicazione: (2024)
TrajLearn: Trajectory Prediction Learning using Deep Generative Models
di: Nadiri, Amirhossein, et al.
Pubblicazione: (2024)
di: Nadiri, Amirhossein, et al.
Pubblicazione: (2024)
AMEND: A Mixture of Experts Framework for Long-tailed Trajectory Prediction
di: Mercurius, Ray Coden, et al.
Pubblicazione: (2024)
di: Mercurius, Ray Coden, et al.
Pubblicazione: (2024)
Vision-Language Model Fine-Tuning via Simple Parameter-Efficient Modification
di: Li, Ming, et al.
Pubblicazione: (2024)
di: Li, Ming, et al.
Pubblicazione: (2024)
PhysBench: Benchmarking and Enhancing Vision-Language Models for Physical World Understanding
di: Chow, Wei, et al.
Pubblicazione: (2025)
di: Chow, Wei, et al.
Pubblicazione: (2025)
Language and Planning in Robotic Navigation: A Multilingual Evaluation of State-of-the-Art Models
di: Mansour, Malak, et al.
Pubblicazione: (2025)
di: Mansour, Malak, et al.
Pubblicazione: (2025)
ChatVLA: Unified Multimodal Understanding and Robot Control with Vision-Language-Action Model
di: Zhou, Zhongyi, et al.
Pubblicazione: (2025)
di: Zhou, Zhongyi, et al.
Pubblicazione: (2025)
Moto: Latent Motion Token as the Bridging Language for Learning Robot Manipulation from Videos
di: Chen, Yi, et al.
Pubblicazione: (2024)
di: Chen, Yi, et al.
Pubblicazione: (2024)
Documenti analoghi
-
Continuous Locomotive Crowd Behavior Generation
di: Bae, Inhwan, et al.
Pubblicazione: (2025) -
SingularTrajectory: Universal Trajectory Predictor Using Diffusion Model
di: Bae, Inhwan, et al.
Pubblicazione: (2024) -
Depth Prompting for Sensor-Agnostic Depth Estimation
di: Park, Jin-Hwi, et al.
Pubblicazione: (2024) -
Fully Explicit Dynamic Gaussian Splatting
di: Lee, Junoh, et al.
Pubblicazione: (2024) -
Relaxed Rigidity with Ray-based Grouping for Dynamic Gaussian Splatting
di: Lee, Junoh, et al.
Pubblicazione: (2026)