VideoOrion: Tokenizing Object Dynamics in Videos
Fuente:
arXiv
Saved in:
| Main Authors: | Feng, Yicheng, Li, Yijiang, Zhang, Wanpeng, Luo, Hao, Yue, Zihao, Zheng, Sipeng, Lu, Zongqing |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Unified Multimodal Understanding via Byte-Pair Visual Encoding
by: Zhang, Wanpeng, et al.
Published: (2025)
by: Zhang, Wanpeng, et al.
Published: (2025)
Being-H0.7: A Latent World-Action Model from Egocentric Videos
by: Luo, Hao, et al.
Published: (2026)
by: Luo, Hao, et al.
Published: (2026)
Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos
by: Luo, Hao, et al.
Published: (2025)
by: Luo, Hao, et al.
Published: (2025)
From Pixels to Tokens: Byte-Pair Encoding on Quantized Visual Modalities
by: Zhang, Wanpeng, et al.
Published: (2024)
by: Zhang, Wanpeng, et al.
Published: (2024)
Pre-trained Visual Dynamics Representations for Efficient Policy Learning
by: Luo, Hao, et al.
Published: (2024)
by: Luo, Hao, et al.
Published: (2024)
Scaling Large Motion Models with Million-Level Human Motions
by: Wang, Ye, et al.
Published: (2024)
by: Wang, Ye, et al.
Published: (2024)
Reinforcement Learning Friendly Vision-Language Model for Minecraft
by: Jiang, Haobin, et al.
Published: (2023)
by: Jiang, Haobin, et al.
Published: (2023)
Being-M0.5: A Real-Time Controllable Vision-Language-Motion Model
by: Cao, Bin, et al.
Published: (2025)
by: Cao, Bin, et al.
Published: (2025)
Robust Motion Generation using Part-level Reliable Data from Videos
by: Li, Boyuan, et al.
Published: (2025)
by: Li, Boyuan, et al.
Published: (2025)
HyperTokens: Controlling Token Dynamics for Continual Video-Language Understanding
by: Nguyen, Toan, et al.
Published: (2026)
by: Nguyen, Toan, et al.
Published: (2026)
Dancing with Still Images: Video Distillation via Static-Dynamic Disentanglement
by: Wang, Ziyu, et al.
Published: (2023)
by: Wang, Ziyu, et al.
Published: (2023)
CMTM: Cross-Modal Token Modulation for Unsupervised Video Object Segmentation
by: Jeon, Inseok, et al.
Published: (2026)
by: Jeon, Inseok, et al.
Published: (2026)
UniCode: Learning a Unified Codebook for Multimodal Large Language Models
by: Zheng, Sipeng, et al.
Published: (2024)
by: Zheng, Sipeng, et al.
Published: (2024)
Beyond Training: Dynamic Token Merging for Zero-Shot Video Understanding
by: Zhang, Yiming, et al.
Published: (2024)
by: Zhang, Yiming, et al.
Published: (2024)
CETCAM: Camera-Controllable Video Generation via Consistent and Extensible Tokenization
by: Zhao, Zelin, et al.
Published: (2025)
by: Zhao, Zelin, et al.
Published: (2025)
Saliency-Motion Guided Trunk-Collateral Network for Unsupervised Video Object Segmentation
by: Zheng, Xiangyu, et al.
Published: (2025)
by: Zheng, Xiangyu, et al.
Published: (2025)
VideoFlexTok: Flexible-Length Coarse-to-Fine Video Tokenization
by: Atanov, Andrei, et al.
Published: (2026)
by: Atanov, Andrei, et al.
Published: (2026)
DyCoke: Dynamic Compression of Tokens for Fast Video Large Language Models
by: Tao, Keda, et al.
Published: (2024)
by: Tao, Keda, et al.
Published: (2024)
Reasoning-Enhanced Object-Centric Learning for Videos
by: Li, Jian, et al.
Published: (2024)
by: Li, Jian, et al.
Published: (2024)
Particle-Grid Neural Dynamics for Learning Deformable Object Models from RGB-D Videos
by: Zhang, Kaifeng, et al.
Published: (2025)
by: Zhang, Kaifeng, et al.
Published: (2025)
Lecture Video Visual Objects (LVVO) Dataset: A Benchmark for Visual Object Detection in Educational Videos
by: Biswas, Dipayan, et al.
Published: (2025)
by: Biswas, Dipayan, et al.
Published: (2025)
MaskVD: Region Masking for Efficient Video Object Detection
by: Sarkar, Sreetama, et al.
Published: (2024)
by: Sarkar, Sreetama, et al.
Published: (2024)
DDLP: Unsupervised Object-Centric Video Prediction with Deep Dynamic Latent Particles
by: Daniel, Tal, et al.
Published: (2023)
by: Daniel, Tal, et al.
Published: (2023)
VideoRefer Suite: Advancing Spatial-Temporal Object Understanding with Video LLM
by: Yuan, Yuqian, et al.
Published: (2024)
by: Yuan, Yuqian, et al.
Published: (2024)
Unified Spatio-Temporal Token Scoring for Efficient Video VLMs
by: Zhang, Jianrui, et al.
Published: (2026)
by: Zhang, Jianrui, et al.
Published: (2026)
VistaDPO: Video Hierarchical Spatial-Temporal Direct Preference Optimization for Large Video Models
by: Huang, Haojian, et al.
Published: (2025)
by: Huang, Haojian, et al.
Published: (2025)
iVideoGPT: Interactive VideoGPTs are Scalable World Models
by: Wu, Jialong, et al.
Published: (2024)
by: Wu, Jialong, et al.
Published: (2024)
Graph2Video: Leveraging Video Models to Model Dynamic Graph Evolution
by: Liu, Hua, et al.
Published: (2026)
by: Liu, Hua, et al.
Published: (2026)
Object-Centric Diffusion for Efficient Video Editing
by: Kahatapitiya, Kumara, et al.
Published: (2024)
by: Kahatapitiya, Kumara, et al.
Published: (2024)
OpenT2M: No-frill Motion Generation with Open-source,Large-scale, High-quality Data
by: Cao, Bin, et al.
Published: (2026)
by: Cao, Bin, et al.
Published: (2026)
Taking Notes Brings Focus? Towards Multi-Turn Multimodal Dialogue Learning
by: Liu, Jiazheng, et al.
Published: (2025)
by: Liu, Jiazheng, et al.
Published: (2025)
Uni-AdaFocus: Spatial-temporal Dynamic Computation for Video Recognition
by: Wang, Yulin, et al.
Published: (2024)
by: Wang, Yulin, et al.
Published: (2024)
H$_{2}$OT: Hierarchical Hourglass Tokenizer for Efficient Video Pose Transformers
by: Li, Wenhao, et al.
Published: (2025)
by: Li, Wenhao, et al.
Published: (2025)
Recurrent Attention-based Token Selection for Efficient Streaming Video-LLMs
by: Dorovatas, Vaggelis, et al.
Published: (2025)
by: Dorovatas, Vaggelis, et al.
Published: (2025)
EVEREST: Efficient Masked Video Autoencoder by Removing Redundant Spatiotemporal Tokens
by: Hwang, Sunil, et al.
Published: (2022)
by: Hwang, Sunil, et al.
Published: (2022)
Don't Look Twice: Faster Video Transformers with Run-Length Tokenization
by: Choudhury, Rohan, et al.
Published: (2024)
by: Choudhury, Rohan, et al.
Published: (2024)
Vidu: a Highly Consistent, Dynamic and Skilled Text-to-Video Generator with Diffusion Models
by: Bao, Fan, et al.
Published: (2024)
by: Bao, Fan, et al.
Published: (2024)
Physics-Driven Spatiotemporal Modeling for AI-Generated Video Detection
by: Zhang, Shuhai, et al.
Published: (2025)
by: Zhang, Shuhai, et al.
Published: (2025)
Image and Video Tokenization with Binary Spherical Quantization
by: Zhao, Yue, et al.
Published: (2024)
by: Zhao, Yue, et al.
Published: (2024)
One Trajectory, One Token: Grounded Video Tokenization via Panoptic Sub-object Trajectory
by: Zheng, Chenhao, et al.
Published: (2025)
by: Zheng, Chenhao, et al.
Published: (2025)
Similar Items
-
Unified Multimodal Understanding via Byte-Pair Visual Encoding
by: Zhang, Wanpeng, et al.
Published: (2025) -
Being-H0.7: A Latent World-Action Model from Egocentric Videos
by: Luo, Hao, et al.
Published: (2026) -
Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos
by: Luo, Hao, et al.
Published: (2025) -
From Pixels to Tokens: Byte-Pair Encoding on Quantized Visual Modalities
by: Zhang, Wanpeng, et al.
Published: (2024) -
Pre-trained Visual Dynamics Representations for Efficient Policy Learning
by: Luo, Hao, et al.
Published: (2024)