Enhanced Motion Forecasting with Plug-and-Play Multimodal Large Language Models
Fuente:
arXiv
Saved in:
| Main Authors: | Luo, Katie, Ji, Jingwei, He, Tong, Xu, Runsheng, Xie, Yichen, Anguelov, Dragomir, Tan, Mingxing |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
S4-Driver: Scalable Self-Supervised Driving Multimodal Large Language Modelwith Spatio-Temporal Visual Representation
by: Xie, Yichen, et al.
Published: (2025)
by: Xie, Yichen, et al.
Published: (2025)
PVTransformer: Point-to-Voxel Transformer for Scalable 3D Object Detection
by: Leng, Zhaoqi, et al.
Published: (2024)
by: Leng, Zhaoqi, et al.
Published: (2024)
EMMA: End-to-End Multimodal Model for Autonomous Driving
by: Hwang, Jyh-Jing, et al.
Published: (2024)
by: Hwang, Jyh-Jing, et al.
Published: (2024)
Plug-and-Play Grounding of Reasoning in Multimodal Large Language Models
by: Chen, Jiaxing, et al.
Published: (2024)
by: Chen, Jiaxing, et al.
Published: (2024)
MotionCraft: Crafting Whole-Body Motion with Plug-and-Play Multimodal Controls
by: Bian, Yuxuan, et al.
Published: (2024)
by: Bian, Yuxuan, et al.
Published: (2024)
WOMD-LiDAR: Raw Sensor Dataset Benchmark for Motion Forecasting
by: Chen, Kan, et al.
Published: (2023)
by: Chen, Kan, et al.
Published: (2023)
SceneDiffuser++: City-Scale Traffic Simulation via a Generative World Model
by: Tan, Shuhan, et al.
Published: (2025)
by: Tan, Shuhan, et al.
Published: (2025)
Plug-and-Play Tri-Branch Invertible Block for Image Rescaling
by: Bao, Jingwei, et al.
Published: (2024)
by: Bao, Jingwei, et al.
Published: (2024)
WOD-E2E: Waymo Open Dataset for End-to-End Driving in Challenging Long-tail Scenarios
by: Xu, Runsheng, et al.
Published: (2025)
by: Xu, Runsheng, et al.
Published: (2025)
GS-Net: Generalizable Plug-and-Play 3D Gaussian Splatting Module
by: Zhang, Yichen, et al.
Published: (2024)
by: Zhang, Yichen, et al.
Published: (2024)
Video Compression Commander: Plug-and-Play Inference Acceleration for Video Large Language Models
by: Liu, Xuyang, et al.
Published: (2025)
by: Liu, Xuyang, et al.
Published: (2025)
Interactive Text-to-Image Retrieval with Large Language Models: A Plug-and-Play Approach
by: Lee, Saehyung, et al.
Published: (2024)
by: Lee, Saehyung, et al.
Published: (2024)
MoST: Multi-modality Scene Tokenization for Motion Prediction
by: Mu, Norman, et al.
Published: (2024)
by: Mu, Norman, et al.
Published: (2024)
Plug, Play, and Fortify: A Low-Cost Module for Robust Multimodal Image Understanding Models
by: Lu, Siqi, et al.
Published: (2026)
by: Lu, Siqi, et al.
Published: (2026)
Scene Reconstruction as Mapping Priors for 3D Detection
by: Fu, Yang, et al.
Published: (2026)
by: Fu, Yang, et al.
Published: (2026)
Segmentation as A Plug-and-Play Capability for Frozen Multimodal LLMs
by: Liu, Jiazhen, et al.
Published: (2025)
by: Liu, Jiazhen, et al.
Published: (2025)
Plug-and-Play Logit Fusion for Heterogeneous Pathology Foundation Models
by: Huang, Gexin, et al.
Published: (2026)
by: Huang, Gexin, et al.
Published: (2026)
Plug-and-Play 1.x-Bit KV Cache Quantization for Video Large Language Models
by: Tao, Keda, et al.
Published: (2025)
by: Tao, Keda, et al.
Published: (2025)
A Plug-and-Play Framework for Volumetric Light-Sheet Image Reconstruction
by: Gong, Yi, et al.
Published: (2025)
by: Gong, Yi, et al.
Published: (2025)
A Plug-and-Play Multi-Criteria Guidance for Diverse In-Betweening Human Motion Generation
by: Yu, Hua, et al.
Published: (2025)
by: Yu, Hua, et al.
Published: (2025)
Global Compression Commander: Plug-and-Play Inference Acceleration for High-Resolution Large Vision-Language Models
by: Liu, Xuyang, et al.
Published: (2025)
by: Liu, Xuyang, et al.
Published: (2025)
A Plug-and-Play Physical Motion Restoration Approach for In-the-Wild High-Difficulty Motions
by: Zhang, Youliang, et al.
Published: (2024)
by: Zhang, Youliang, et al.
Published: (2024)
RAG-Adapter: A Plug-and-Play RAG-enhanced Framework for Long Video Understanding
by: Tan, Xichen, et al.
Published: (2025)
by: Tan, Xichen, et al.
Published: (2025)
LET-3D-AP: Longitudinal Error Tolerant 3D Average Precision for Camera-Only 3D Detection
by: Hung, Wei-Chih, et al.
Published: (2022)
by: Hung, Wei-Chih, et al.
Published: (2022)
Plug-and-Play Diffusion Distillation
by: Hsiao, Yi-Ting, et al.
Published: (2024)
by: Hsiao, Yi-Ting, et al.
Published: (2024)
Protecting NeRFs' Copyright via Plug-And-Play Watermarking Base Model
by: Song, Qi, et al.
Published: (2024)
by: Song, Qi, et al.
Published: (2024)
SR$^{2}$-Net: A General Plug-and-Play Model for Spectral Refinement in Hyperspectral Image Super-Resolution
by: He, Ji-Xuan, et al.
Published: (2026)
by: He, Ji-Xuan, et al.
Published: (2026)
MQADet: A Plug-and-Play Paradigm for Enhancing Open-Vocabulary Object Detection via Multimodal Question Answering
by: Li, Caixiong, et al.
Published: (2025)
by: Li, Caixiong, et al.
Published: (2025)
SceneCrafter: Controllable Multi-View Driving Scene Editing
by: Zhu, Zehao, et al.
Published: (2025)
by: Zhu, Zehao, et al.
Published: (2025)
RoleMotion: A Large-Scale Dataset towards Robust Scene-Specific Role-Playing Motion Synthesis with Fine-grained Descriptions
by: Peng, Junran, et al.
Published: (2025)
by: Peng, Junran, et al.
Published: (2025)
Integrating Reweighted Least Squares with Plug-and-Play Diffusion Priors for Noisy Image Restoration
by: Li, Ji, et al.
Published: (2025)
by: Li, Ji, et al.
Published: (2025)
Enhancing Logits Distillation with Plug\&Play Kendall's $τ$ Ranking Loss
by: Guan, Yuchen, et al.
Published: (2024)
by: Guan, Yuchen, et al.
Published: (2024)
Prune2Drive: A Plug-and-Play Framework for Accelerating Vision-Language Models in Autonomous Driving
by: Xiong, Minhao, et al.
Published: (2025)
by: Xiong, Minhao, et al.
Published: (2025)
SearchLVLMs: A Plug-and-Play Framework for Augmenting Large Vision-Language Models by Searching Up-to-Date Internet Knowledge
by: Li, Chuanhao, et al.
Published: (2024)
by: Li, Chuanhao, et al.
Published: (2024)
Seeing Clearly, Reasoning Confidently: Plug-and-Play Remedies for Vision Language Model Blindness
by: Hu, Xin, et al.
Published: (2026)
by: Hu, Xin, et al.
Published: (2026)
Next Token Is Enough: Realistic Image Quality and Aesthetic Scoring with Multimodal Large Language Model
by: Li, Mingxing, et al.
Published: (2025)
by: Li, Mingxing, et al.
Published: (2025)
PADS: Plug-and-Play 3D Human Pose Analysis via Diffusion Generative Modeling
by: Ji, Haorui, et al.
Published: (2024)
by: Ji, Haorui, et al.
Published: (2024)
Safety of Multimodal Large Language Models on Images and Texts
by: Liu, Xin, et al.
Published: (2024)
by: Liu, Xin, et al.
Published: (2024)
FasterVAR: Plug-and-Play Acceleration for Visual Autoregressive Models
by: Li, Senmao, et al.
Published: (2025)
by: Li, Senmao, et al.
Published: (2025)
Drive&Gen: Co-Evaluating End-to-End Driving and Video Generation Models
by: Wang, Jiahao, et al.
Published: (2025)
by: Wang, Jiahao, et al.
Published: (2025)
Similar Items
-
S4-Driver: Scalable Self-Supervised Driving Multimodal Large Language Modelwith Spatio-Temporal Visual Representation
by: Xie, Yichen, et al.
Published: (2025) -
PVTransformer: Point-to-Voxel Transformer for Scalable 3D Object Detection
by: Leng, Zhaoqi, et al.
Published: (2024) -
EMMA: End-to-End Multimodal Model for Autonomous Driving
by: Hwang, Jyh-Jing, et al.
Published: (2024) -
Plug-and-Play Grounding of Reasoning in Multimodal Large Language Models
by: Chen, Jiaxing, et al.
Published: (2024) -
MotionCraft: Crafting Whole-Body Motion with Plug-and-Play Multimodal Controls
by: Bian, Yuxuan, et al.
Published: (2024)