AutoTraces: Autoregressive Trajectory Forecasting via Multimodal Large Language Models
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Wang, Teng, Lu, Yanting, Wang, Ruize |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Unified Multimodal Models as Auto-Encoders
von: Yan, Zhiyuan, et al.
Veröffentlicht: (2025)
von: Yan, Zhiyuan, et al.
Veröffentlicht: (2025)
VARGPT: Unified Understanding and Generation in a Visual Autoregressive Multimodal Large Language Model
von: Zhuang, Xianwei, et al.
Veröffentlicht: (2025)
von: Zhuang, Xianwei, et al.
Veröffentlicht: (2025)
InternSVG: Towards Unified SVG Tasks with Multimodal Large Language Models
von: Wang, Haomin, et al.
Veröffentlicht: (2025)
von: Wang, Haomin, et al.
Veröffentlicht: (2025)
Beyond Words: Advancing Long-Text Image Generation via Multimodal Autoregressive Models
von: Wang, Alex Jinpeng, et al.
Veröffentlicht: (2025)
von: Wang, Alex Jinpeng, et al.
Veröffentlicht: (2025)
Debiasing Multimodal Large Language Models via Penalization of Language Priors
von: Zhang, YiFan, et al.
Veröffentlicht: (2024)
von: Zhang, YiFan, et al.
Veröffentlicht: (2024)
DCARL: A Divide-and-Conquer Framework for Autoregressive Long-Trajectory Video Generation
von: Ouyang, Junyi, et al.
Veröffentlicht: (2026)
von: Ouyang, Junyi, et al.
Veröffentlicht: (2026)
Masked Autoregressive Model for Weather Forecasting
von: Kim, Doyi, et al.
Veröffentlicht: (2024)
von: Kim, Doyi, et al.
Veröffentlicht: (2024)
Enhanced Motion Forecasting with Plug-and-Play Multimodal Large Language Models
von: Luo, Katie, et al.
Veröffentlicht: (2025)
von: Luo, Katie, et al.
Veröffentlicht: (2025)
Learning Trajectory-Aware Multimodal Large Language Models for Video Reasoning Segmentation
von: Luo, Jingnan, et al.
Veröffentlicht: (2026)
von: Luo, Jingnan, et al.
Veröffentlicht: (2026)
Reflective Instruction Tuning: Mitigating Hallucinations in Large Vision-Language Models
von: Zhang, Jinrui, et al.
Veröffentlicht: (2024)
von: Zhang, Jinrui, et al.
Veröffentlicht: (2024)
TraceVision: Trajectory-Aware Vision-Language Model for Human-Like Spatial Understanding
von: Yang, Fan, et al.
Veröffentlicht: (2026)
von: Yang, Fan, et al.
Veröffentlicht: (2026)
Hierarchical Light Transformer Ensembles for Multimodal Trajectory Forecasting
von: Lafage, Adrien, et al.
Veröffentlicht: (2024)
von: Lafage, Adrien, et al.
Veröffentlicht: (2024)
BiomechGPT: Towards a Biomechanically Fluent Multimodal Foundation Model for Clinically Relevant Motion Tasks
von: Yang, Ruize, et al.
Veröffentlicht: (2025)
von: Yang, Ruize, et al.
Veröffentlicht: (2025)
Auto DragGAN: Editing the Generative Image Manifold in an Autoregressive Manner
von: Cai, Pengxiang, et al.
Veröffentlicht: (2024)
von: Cai, Pengxiang, et al.
Veröffentlicht: (2024)
MMaDA-Parallel: Multimodal Large Diffusion Language Models for Thinking-Aware Editing and Generation
von: Tian, Ye, et al.
Veröffentlicht: (2025)
von: Tian, Ye, et al.
Veröffentlicht: (2025)
OVT-B: A New Large-Scale Benchmark for Open-Vocabulary Multi-Object Tracking
von: Liang, Haiji, et al.
Veröffentlicht: (2024)
von: Liang, Haiji, et al.
Veröffentlicht: (2024)
Trajectory Prediction Meets Large Language Models: A Survey
von: Xu, Yi, et al.
Veröffentlicht: (2025)
von: Xu, Yi, et al.
Veröffentlicht: (2025)
Dynamic Pyramid Network for Efficient Multimodal Large Language Model
von: Ai, Hao, et al.
Veröffentlicht: (2025)
von: Ai, Hao, et al.
Veröffentlicht: (2025)
Vision-Centric Activation and Coordination for Multimodal Large Language Models
von: Wang, Yunnan, et al.
Veröffentlicht: (2025)
von: Wang, Yunnan, et al.
Veröffentlicht: (2025)
Conditional Panoramic Image Generation via Masked Autoregressive Modeling
von: Wang, Chaoyang, et al.
Veröffentlicht: (2025)
von: Wang, Chaoyang, et al.
Veröffentlicht: (2025)
Learning Multimodal Volumetric Features for Large-Scale Neuron Tracing
von: Chen, Qihua, et al.
Veröffentlicht: (2024)
von: Chen, Qihua, et al.
Veröffentlicht: (2024)
Grounding Everything in Tokens for Multimodal Large Language Models
von: Ren, Xiangxuan, et al.
Veröffentlicht: (2025)
von: Ren, Xiangxuan, et al.
Veröffentlicht: (2025)
Enhancing the Reasoning Ability of Multimodal Large Language Models via Mixed Preference Optimization
von: Wang, Weiyun, et al.
Veröffentlicht: (2024)
von: Wang, Weiyun, et al.
Veröffentlicht: (2024)
EfficientLLaVA:Generalizable Auto-Pruning for Large Vision-language Models
von: Liang, Yinan, et al.
Veröffentlicht: (2025)
von: Liang, Yinan, et al.
Veröffentlicht: (2025)
Period-LLM: Extending the Periodic Capability of Multimodal Large Language Model
von: Zhang, Yuting, et al.
Veröffentlicht: (2025)
von: Zhang, Yuting, et al.
Veröffentlicht: (2025)
DUALVISION: RGB-Infrared Multimodal Large Language Models for Robust Visual Reasoning
von: Majeedi, Abrar, et al.
Veröffentlicht: (2026)
von: Majeedi, Abrar, et al.
Veröffentlicht: (2026)
BoxTuning: Directly Injecting the Object Box for Multimodal Model Fine-Tuning
von: Qian, Zekun, et al.
Veröffentlicht: (2026)
von: Qian, Zekun, et al.
Veröffentlicht: (2026)
LLMGA: Multimodal Large Language Model based Generation Assistant
von: Xia, Bin, et al.
Veröffentlicht: (2023)
von: Xia, Bin, et al.
Veröffentlicht: (2023)
Guiding Instruction-based Image Editing via Multimodal Large Language Models
von: Fu, Tsu-Jui, et al.
Veröffentlicht: (2023)
von: Fu, Tsu-Jui, et al.
Veröffentlicht: (2023)
Hierarchical Auto-Organizing System for Open-Ended Multi-Agent Navigation
von: Zhao, Zhonghan, et al.
Veröffentlicht: (2024)
von: Zhao, Zhonghan, et al.
Veröffentlicht: (2024)
AutoRefiner: Improving Autoregressive Video Diffusion Models via Reflective Refinement Over the Stochastic Sampling Path
von: Yu, Zhengyang, et al.
Veröffentlicht: (2025)
von: Yu, Zhengyang, et al.
Veröffentlicht: (2025)
HATS: Hardness-Aware Trajectory Synthesis for GUI Agents
von: Shao, Rui, et al.
Veröffentlicht: (2026)
von: Shao, Rui, et al.
Veröffentlicht: (2026)
Assessment of Multimodal Large Language Models in Alignment with Human Values
von: Shi, Zhelun, et al.
Veröffentlicht: (2024)
von: Shi, Zhelun, et al.
Veröffentlicht: (2024)
Survey of Adversarial Robustness in Multimodal Large Language Models
von: Jiang, Chengze, et al.
Veröffentlicht: (2025)
von: Jiang, Chengze, et al.
Veröffentlicht: (2025)
Seg-VAR: Image Segmentation with Visual Autoregressive Modeling
von: Zheng, Rongkun, et al.
Veröffentlicht: (2025)
von: Zheng, Rongkun, et al.
Veröffentlicht: (2025)
DragEntity: Trajectory Guided Video Generation using Entity and Positional Relationships
von: Wan, Zhang, et al.
Veröffentlicht: (2024)
von: Wan, Zhang, et al.
Veröffentlicht: (2024)
VL-Trojan: Multimodal Instruction Backdoor Attacks against Autoregressive Visual Language Models
von: Liang, Jiawei, et al.
Veröffentlicht: (2024)
von: Liang, Jiawei, et al.
Veröffentlicht: (2024)
Trajectory Mamba: Efficient Attention-Mamba Forecasting Model Based on Selective SSM
von: Huang, Yizhou, et al.
Veröffentlicht: (2025)
von: Huang, Yizhou, et al.
Veröffentlicht: (2025)
Verb Mirage: Unveiling and Assessing Verb Concept Hallucinations in Multimodal Large Language Models
von: Wang, Zehao, et al.
Veröffentlicht: (2024)
von: Wang, Zehao, et al.
Veröffentlicht: (2024)
Semantic Alignment for Multimodal Large Language Models
von: Wu, Tao, et al.
Veröffentlicht: (2024)
von: Wu, Tao, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Unified Multimodal Models as Auto-Encoders
von: Yan, Zhiyuan, et al.
Veröffentlicht: (2025) -
VARGPT: Unified Understanding and Generation in a Visual Autoregressive Multimodal Large Language Model
von: Zhuang, Xianwei, et al.
Veröffentlicht: (2025) -
InternSVG: Towards Unified SVG Tasks with Multimodal Large Language Models
von: Wang, Haomin, et al.
Veröffentlicht: (2025) -
Beyond Words: Advancing Long-Text Image Generation via Multimodal Autoregressive Models
von: Wang, Alex Jinpeng, et al.
Veröffentlicht: (2025) -
Debiasing Multimodal Large Language Models via Penalization of Language Priors
von: Zhang, YiFan, et al.
Veröffentlicht: (2024)