Mode-as-Sequence: Translating Multimodal Motion Prediction into Unified Sequential Mode Modeling
Fuente:
arXiv
Saved in:
| Main Authors: | Zhou, Zikang, Hu, Haibo, Chen, Xinhong, Zhang, Yifan, Guan, Nan, Li, Yung-Hui, Xue, Chun Jason, Wang, Jianping |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
ModeSeq: Taming Sparse Multimodal Motion Prediction with Sequential Mode Modeling
by: Zhou, Zikang, et al.
Published: (2024)
by: Zhou, Zikang, et al.
Published: (2024)
RALAD: Bridging the Real-to-Sim Domain Gap in Autonomous Driving with Retrieval-Augmented Learning
by: Zuo, Jiacheng, et al.
Published: (2025)
by: Zuo, Jiacheng, et al.
Published: (2025)
BehaviorGPT: Smart Agent Simulation for Autonomous Driving with Next-Patch Prediction
by: Zhou, Zikang, et al.
Published: (2024)
by: Zhou, Zikang, et al.
Published: (2024)
GM-Skip: Metric-Guided Transformer Block Skipping for Efficient Vision-Language Models
by: Huang, Lianming, et al.
Published: (2025)
by: Huang, Lianming, et al.
Published: (2025)
On-Demand Multi-Task Sparsity for Efficient Large-Model Deployment on Edge Devices
by: Huang, Lianming, et al.
Published: (2025)
by: Huang, Lianming, et al.
Published: (2025)
Global Regulation and Excitation via Attention Tuning for Stereo Matching
by: Li, Jiahao, et al.
Published: (2025)
by: Li, Jiahao, et al.
Published: (2025)
DeeAD: Dynamic Early Exit of Vision-Language Action for Efficient Autonomous Driving
by: HU, Haibo, et al.
Published: (2025)
by: HU, Haibo, et al.
Published: (2025)
AD-EE: Early Exiting for Fast and Reliable Vision-Language Models in Autonomous Driving
by: Huang, Lianming, et al.
Published: (2025)
by: Huang, Lianming, et al.
Published: (2025)
Geometry Reinforced Efficient Attention Tuning Equipped with Normals for Robust Stereo Matching
by: Li, Jiahao, et al.
Published: (2026)
by: Li, Jiahao, et al.
Published: (2026)
Real-Time Motion Detection Using Dynamic Mode Decomposition
by: Mignacca, Marco, et al.
Published: (2024)
by: Mignacca, Marco, et al.
Published: (2024)
MMCM: Multimodality-aware Metric using Clustering-based Modes for Probabilistic Human Motion Prediction
by: Tokoro, Kyotaro, et al.
Published: (2025)
by: Tokoro, Kyotaro, et al.
Published: (2025)
IHC Matters: Incorporating IHC analysis to H&E Whole Slide Image Analysis for Improved Cancer Grading via Two-stage Multimodal Bilinear Pooling Fusion
by: Wang, Jun, et al.
Published: (2024)
by: Wang, Jun, et al.
Published: (2024)
SHAP-CAT: A interpretable multi-modal framework enhancing WSI classification via virtual staining and shapley-value-based multimodal fusion
by: Wang, Jun, et al.
Published: (2024)
by: Wang, Jun, et al.
Published: (2024)
Deep Non-rigid Structure-from-Motion: A Sequence-to-Sequence Translation Perspective
by: Deng, Hui, et al.
Published: (2022)
by: Deng, Hui, et al.
Published: (2022)
Motion Modes: What Could Happen Next?
by: Pandey, Karran, et al.
Published: (2024)
by: Pandey, Karran, et al.
Published: (2024)
Advances in Multiple Instance Learning for Whole Slide Image Analysis: Techniques, Challenges, and Future Directions
by: Wang, Jun, et al.
Published: (2024)
by: Wang, Jun, et al.
Published: (2024)
VRSA: Jailbreaking Multimodal Large Language Models through Visual Reasoning Sequential Attack
by: Zhao, Shiji, et al.
Published: (2025)
by: Zhao, Shiji, et al.
Published: (2025)
WISE: A Framework for Gigapixel Whole-Slide-Image Lossless Compression
by: Mao, Yu, et al.
Published: (2025)
by: Mao, Yu, et al.
Published: (2025)
Sequential Gaussian Avatars with Hierarchical Motion Context
by: Xu, Wangze, et al.
Published: (2024)
by: Xu, Wangze, et al.
Published: (2024)
NextFlow: Unified Sequential Modeling Activates Multimodal Understanding and Generation
by: Zhang, Huichao, et al.
Published: (2026)
by: Zhang, Huichao, et al.
Published: (2026)
Timely Fusion of Surround Radar/Lidar for Object Detection in Autonomous Driving Systems
by: Xie, Wenjing, et al.
Published: (2023)
by: Xie, Wenjing, et al.
Published: (2023)
ModeT: Learning Deformable Image Registration via Motion Decomposition Transformer
by: Wang, Haiqiao, et al.
Published: (2023)
by: Wang, Haiqiao, et al.
Published: (2023)
BAHOP: Similarity-based Basin Hopping for A fast hyper-parameter search in WSI classification
by: Wang, Jun, et al.
Published: (2024)
by: Wang, Jun, et al.
Published: (2024)
Query-Kontext: An Unified Multimodal Model for Image Generation and Editing
by: Song, Yuxin, et al.
Published: (2025)
by: Song, Yuxin, et al.
Published: (2025)
RSHallu: Dual-Mode Hallucination Evaluation for Remote-Sensing Multimodal Large Language Models with Domain-Tailored Mitigation
by: Zhou, Zihui, et al.
Published: (2026)
by: Zhou, Zihui, et al.
Published: (2026)
Bundle Adjustment in the Eager Mode
by: Zhan, Zitong, et al.
Published: (2024)
by: Zhan, Zitong, et al.
Published: (2024)
MotionVerse: A Unified Multimodal Framework for Motion Comprehension, Generation and Editing
by: Hou, Ruibing, et al.
Published: (2025)
by: Hou, Ruibing, et al.
Published: (2025)
EnLVAM: Enhanced Left Ventricle Linear Measurements Utilizing Anatomical Motion Mode
by: Singh, Durgesh K., et al.
Published: (2025)
by: Singh, Durgesh K., et al.
Published: (2025)
MotionSight: Boosting Fine-Grained Motion Understanding in Multimodal LLMs
by: Du, Yipeng, et al.
Published: (2025)
by: Du, Yipeng, et al.
Published: (2025)
Multimodal Sense-Informed Prediction of 3D Human Motions
by: Lou, Zhenyu, et al.
Published: (2024)
by: Lou, Zhenyu, et al.
Published: (2024)
AdaptMMBench: Benchmarking Adaptive Multimodal Reasoning for Mode Selection and Reasoning Process
by: Zhang, Xintong, et al.
Published: (2026)
by: Zhang, Xintong, et al.
Published: (2026)
Flow to the Mode: Mode-Seeking Diffusion Autoencoders for State-of-the-Art Image Tokenization
by: Sargent, Kyle, et al.
Published: (2025)
by: Sargent, Kyle, et al.
Published: (2025)
Deep Non-rigid Structure-from-Motion Revisited: Canonicalization and Sequence Modeling
by: Deng, Hui, et al.
Published: (2024)
by: Deng, Hui, et al.
Published: (2024)
MotionLab: Unified Human Motion Generation and Editing via the Motion-Condition-Motion Paradigm
by: Guo, Ziyan, et al.
Published: (2025)
by: Guo, Ziyan, et al.
Published: (2025)
LUMA: Low-Dimension Unified Motion Alignment with Dual-Path Anchoring for Text-to-Motion Diffusion Model
by: Jia, Haozhe, et al.
Published: (2025)
by: Jia, Haozhe, et al.
Published: (2025)
AnyGPT: Unified Multimodal LLM with Discrete Sequence Modeling
by: Zhan, Jun, et al.
Published: (2024)
by: Zhan, Jun, et al.
Published: (2024)
Unified Multimodal Models as Auto-Encoders
by: Yan, Zhiyuan, et al.
Published: (2025)
by: Yan, Zhiyuan, et al.
Published: (2025)
BA-SAM: Scalable Bias-Mode Attention Mask for Segment Anything Model
by: Song, Yiran, et al.
Published: (2024)
by: Song, Yiran, et al.
Published: (2024)
ModeTv2: GPU-accelerated Motion Decomposition Transformer for Pairwise Optimization in Medical Image Registration
by: Wang, Haiqiao, et al.
Published: (2024)
by: Wang, Haiqiao, et al.
Published: (2024)
Collaborative Multi-Mode Pruning for Vision-Language Models
by: Wu, Zimeng, et al.
Published: (2026)
by: Wu, Zimeng, et al.
Published: (2026)
Similar Items
-
ModeSeq: Taming Sparse Multimodal Motion Prediction with Sequential Mode Modeling
by: Zhou, Zikang, et al.
Published: (2024) -
RALAD: Bridging the Real-to-Sim Domain Gap in Autonomous Driving with Retrieval-Augmented Learning
by: Zuo, Jiacheng, et al.
Published: (2025) -
BehaviorGPT: Smart Agent Simulation for Autonomous Driving with Next-Patch Prediction
by: Zhou, Zikang, et al.
Published: (2024) -
GM-Skip: Metric-Guided Transformer Block Skipping for Efficient Vision-Language Models
by: Huang, Lianming, et al.
Published: (2025) -
On-Demand Multi-Task Sparsity for Efficient Large-Model Deployment on Edge Devices
by: Huang, Lianming, et al.
Published: (2025)