From Seeing to Predicting: A Vision-Language Framework for Trajectory Forecasting and Controlled Video Generation
Fuente:
arXiv
Saved in:
| Main Authors: | Yang, Fan, Chen, Zhiyang, Zhu, Yousong, Li, Xin, Wang, Jinqiao |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
TraceVision: Trajectory-Aware Vision-Language Model for Human-Like Spatial Understanding
by: Yang, Fan, et al.
Published: (2026)
by: Yang, Fan, et al.
Published: (2026)
Griffon: Spelling out All Object Locations at Any Granularity with Large Language Models
by: Zhan, Yufei, et al.
Published: (2023)
by: Zhan, Yufei, et al.
Published: (2023)
Griffon-G: Bridging Vision-Language and Vision-Centric Tasks via Large Multimodal Models
by: Zhan, Yufei, et al.
Published: (2024)
by: Zhan, Yufei, et al.
Published: (2024)
FOCUS: Unified Vision-Language Modeling for Interactive Editing Driven by Referential Segmentation
by: Yang, Fan, et al.
Published: (2025)
by: Yang, Fan, et al.
Published: (2025)
Vision-R1: Evolving Human-Free Alignment in Large Vision-Language Models via Vision-Guided Reinforcement Learning
by: Zhan, Yufei, et al.
Published: (2025)
by: Zhan, Yufei, et al.
Published: (2025)
Efficient Masked Autoencoders with Self-Consistency
by: Li, Zhaowen, et al.
Published: (2023)
by: Li, Zhaowen, et al.
Published: (2023)
GeM-VG: Towards Generalized Multi-image Visual Grounding with Multimodal Large Language Models
by: Zheng, Shurong, et al.
Published: (2026)
by: Zheng, Shurong, et al.
Published: (2026)
Griffon v2: Advancing Multimodal Perception with High-Resolution Scaling and Visual-Language Co-Referring
by: Zhan, Yufei, et al.
Published: (2024)
by: Zhan, Yufei, et al.
Published: (2024)
VFaith: Do Large Multimodal Models Really Reason on Seen Images Rather than Previous Memories?
by: Yu, Jiachen, et al.
Published: (2025)
by: Yu, Jiachen, et al.
Published: (2025)
Understand, Think, and Answer: Advancing Visual Reasoning with Large Multimodal Models
by: Zhan, Yufei, et al.
Published: (2025)
by: Zhan, Yufei, et al.
Published: (2025)
LINK: Adaptive Modality Interaction for Audio-Visual Video Parsing
by: Wang, Langyu, et al.
Published: (2024)
by: Wang, Langyu, et al.
Published: (2024)
Mojito: Motion Trajectory and Intensity Control for Video Generation
by: He, Xuehai, et al.
Published: (2024)
by: He, Xuehai, et al.
Published: (2024)
Optimization of Prompt Learning via Multi-Knowledge Representation for Vision-Language Models
by: Zhang, Enming, et al.
Published: (2024)
by: Zhang, Enming, et al.
Published: (2024)
From Plausibility to Verifiability: Risk-Controlled Generative OCR with Vision-Language Models
by: Gong, Weile, et al.
Published: (2026)
by: Gong, Weile, et al.
Published: (2026)
PhysCtrl: Generative Physics for Controllable and Physics-Grounded Video Generation
by: Wang, Chen, et al.
Published: (2025)
by: Wang, Chen, et al.
Published: (2025)
TLControl: Trajectory and Language Control for Human Motion Synthesis
by: Wan, Weilin, et al.
Published: (2023)
by: Wan, Weilin, et al.
Published: (2023)
TGT: Text-Grounded Trajectories for Locally Controlled Video Generation
by: Zhang, Guofeng, et al.
Published: (2025)
by: Zhang, Guofeng, et al.
Published: (2025)
Decoupled Similarity for Task-Aware Token Pruning in Large Vision-Language Models
by: Ma, Kexin, et al.
Published: (2026)
by: Ma, Kexin, et al.
Published: (2026)
Reading $\neq$ Seeing: Diagnosing and Closing the Typography Gap in Vision-Language Models
by: Zhou, Heng, et al.
Published: (2026)
by: Zhou, Heng, et al.
Published: (2026)
FlexTraj: Image-to-Video Generation with Flexible Point Trajectory Control
by: Zhang, Zhiyuan, et al.
Published: (2025)
by: Zhang, Zhiyuan, et al.
Published: (2025)
Motion Prompting: Controlling Video Generation with Motion Trajectories
by: Geng, Daniel, et al.
Published: (2024)
by: Geng, Daniel, et al.
Published: (2024)
MROVSeg: Breaking the Resolution Curse of Vision-Language Models in Open-Vocabulary Image Segmentation
by: Zhu, Yuanbing, et al.
Published: (2024)
by: Zhu, Yuanbing, et al.
Published: (2024)
Focus on What Really Matters in Low-Altitude Governance: A Management-Centric Multi-Modal Benchmark with Implicitly Coordinated Vision-Language Reasoning Framework
by: Chang, Hao, et al.
Published: (2026)
by: Chang, Hao, et al.
Published: (2026)
A Benchmark for Crime Surveillance Video Analysis with Large Models
by: Chen, Haoran, et al.
Published: (2025)
by: Chen, Haoran, et al.
Published: (2025)
MUG: Pseudo Labeling Augmented Audio-Visual Mamba Network for Audio-Visual Video Parsing
by: Wang, Langyu, et al.
Published: (2025)
by: Wang, Langyu, et al.
Published: (2025)
WorldMAP: Bootstrapping Vision-Language Navigation Trajectory Prediction with Generative World Models
by: Chen, Hongjin, et al.
Published: (2026)
by: Chen, Hongjin, et al.
Published: (2026)
Seeing Clearly, Reasoning Confidently: Plug-and-Play Remedies for Vision Language Model Blindness
by: Hu, Xin, et al.
Published: (2026)
by: Hu, Xin, et al.
Published: (2026)
HPNet: Dynamic Trajectory Forecasting with Historical Prediction Attention
by: Tang, Xiaolong, et al.
Published: (2024)
by: Tang, Xiaolong, et al.
Published: (2024)
Optimizing Diffusion Models for Joint Trajectory Prediction and Controllable Generation
by: Wang, Yixiao, et al.
Published: (2024)
by: Wang, Yixiao, et al.
Published: (2024)
ATI: Any Trajectory Instruction for Controllable Video Generation
by: Wang, Angtian, et al.
Published: (2025)
by: Wang, Angtian, et al.
Published: (2025)
Learning Video Generation for Robotic Manipulation with Collaborative Trajectory Control
by: Fu, Xiao, et al.
Published: (2025)
by: Fu, Xiao, et al.
Published: (2025)
Learning Cooperative Trajectory Representations for Motion Forecasting
by: Ruan, Hongzhi, et al.
Published: (2023)
by: Ruan, Hongzhi, et al.
Published: (2023)
DCARL: A Divide-and-Conquer Framework for Autoregressive Long-Trajectory Video Generation
by: Ouyang, Junyi, et al.
Published: (2026)
by: Ouyang, Junyi, et al.
Published: (2026)
Quality-Aware Language-Conditioned Local Auto-Regressive Anomaly Synthesis and Detection
by: Qian, Long, et al.
Published: (2025)
by: Qian, Long, et al.
Published: (2025)
Tora: Trajectory-oriented Diffusion Transformer for Video Generation
by: Zhang, Zhenghao, et al.
Published: (2024)
by: Zhang, Zhenghao, et al.
Published: (2024)
Seeing Realism from Simulation: Efficient Video Transfer for Vision-Language-Action Data Augmentation
by: Hui, Chenyu, et al.
Published: (2026)
by: Hui, Chenyu, et al.
Published: (2026)
Synthetic Data is an Elegant GIFT for Continual Vision-Language Models
by: Wu, Bin, et al.
Published: (2025)
by: Wu, Bin, et al.
Published: (2025)
GThinker: Towards General Multimodal Reasoning via Cue-Guided Rethinking
by: Zhan, Yufei, et al.
Published: (2025)
by: Zhan, Yufei, et al.
Published: (2025)
Motion-Zero: Zero-Shot Moving Object Control Framework for Diffusion-Based Video Generation
by: Chen, Changgu, et al.
Published: (2024)
by: Chen, Changgu, et al.
Published: (2024)
Seeing Across Views: Benchmarking Spatial Reasoning of Vision-Language Models in Robotic Scenes
by: Feng, Zhiyuan, et al.
Published: (2025)
by: Feng, Zhiyuan, et al.
Published: (2025)
Similar Items
-
TraceVision: Trajectory-Aware Vision-Language Model for Human-Like Spatial Understanding
by: Yang, Fan, et al.
Published: (2026) -
Griffon: Spelling out All Object Locations at Any Granularity with Large Language Models
by: Zhan, Yufei, et al.
Published: (2023) -
Griffon-G: Bridging Vision-Language and Vision-Centric Tasks via Large Multimodal Models
by: Zhan, Yufei, et al.
Published: (2024) -
FOCUS: Unified Vision-Language Modeling for Interactive Editing Driven by Referential Segmentation
by: Yang, Fan, et al.
Published: (2025) -
Vision-R1: Evolving Human-Free Alignment in Large Vision-Language Models via Vision-Guided Reinforcement Learning
by: Zhan, Yufei, et al.
Published: (2025)