VTok: A Unified Video Tokenizer with Decoupled Spatial-Temporal Latents
Fuente:
arXiv
Salvato in:
| Autori principali: | Wang, Feng, Shi, Yichun, Yang, Ceyuan, Guo, Qiushan, Sun, Jingxiang, Yuille, Alan, Wang, Peng |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Video Generation with Predictive Latents
di: Zhao, Yian, et al.
Pubblicazione: (2026)
di: Zhao, Yian, et al.
Pubblicazione: (2026)
ViMix-14M: A Curated Multi-Source Video-Text Dataset with Long-Form, High-Quality Captions and Crawl-Free Access
di: Yang, Timing, et al.
Pubblicazione: (2025)
di: Yang, Timing, et al.
Pubblicazione: (2025)
M-VAR: Decoupled Scale-wise Autoregressive Modeling for High-Quality Image Generation
di: Ren, Sucheng, et al.
Pubblicazione: (2024)
di: Ren, Sucheng, et al.
Pubblicazione: (2024)
SweetTok: Semantic-Aware Spatial-Temporal Tokenizer for Compact Video Discretization
di: Tan, Zhentao, et al.
Pubblicazione: (2024)
di: Tan, Zhentao, et al.
Pubblicazione: (2024)
Video-LaVIT: Unified Video-Language Pre-training with Decoupled Visual-Motional Tokenization
di: Jin, Yang, et al.
Pubblicazione: (2024)
di: Jin, Yang, et al.
Pubblicazione: (2024)
Composable Visual Tokenizers with Generator-Free Diagnostics of Learnability
di: Zhao, Bingchen, et al.
Pubblicazione: (2026)
di: Zhao, Bingchen, et al.
Pubblicazione: (2026)
Identity-Preserving Text-to-Video Generation Guided by Simple yet Effective Spatial-Temporal Decoupled Representations
di: Wang, Yuji, et al.
Pubblicazione: (2025)
di: Wang, Yuji, et al.
Pubblicazione: (2025)
DeRA: Decoupled Representation Alignment for Video Tokenization
di: Guo, Pengbo, et al.
Pubblicazione: (2025)
di: Guo, Pengbo, et al.
Pubblicazione: (2025)
Autoregressive Video Autoencoder with Decoupled Temporal and Spatial Context
di: Shen, Cuifeng, et al.
Pubblicazione: (2025)
di: Shen, Cuifeng, et al.
Pubblicazione: (2025)
SCLIP: Rethinking Self-Attention for Dense Vision-Language Inference
di: Wang, Feng, et al.
Pubblicazione: (2023)
di: Wang, Feng, et al.
Pubblicazione: (2023)
Spatial-Temporal Decoupled Reference Conditioning for Identity-Preserving Text-to-Video Generation
di: Chen, Yuheng, et al.
Pubblicazione: (2026)
di: Chen, Yuheng, et al.
Pubblicazione: (2026)
SpatialRGPT: Grounded Spatial Reasoning in Vision Language Models
di: Cheng, An-Chieh, et al.
Pubblicazione: (2024)
di: Cheng, An-Chieh, et al.
Pubblicazione: (2024)
Thinking with Spatial Code for Physical-World Video Reasoning
di: Chen, Jieneng, et al.
Pubblicazione: (2026)
di: Chen, Jieneng, et al.
Pubblicazione: (2026)
RNN as Linear Transformer: A Closer Investigation into Representational Potentials of Visual Mamba Models
di: Yang, Timing, et al.
Pubblicazione: (2025)
di: Yang, Timing, et al.
Pubblicazione: (2025)
Learning Adaptive and Temporally Causal Video Tokenization in a 1D Latent Space
di: Li, Yan, et al.
Pubblicazione: (2025)
di: Li, Yan, et al.
Pubblicazione: (2025)
MMCORE: MultiModal COnnection with Representation Aligned Latent Embeddings
di: Li, Zijie, et al.
Pubblicazione: (2026)
di: Li, Zijie, et al.
Pubblicazione: (2026)
Scaling Laws in Patchification: An Image Is Worth 50,176 Tokens And More
di: Wang, Feng, et al.
Pubblicazione: (2025)
di: Wang, Feng, et al.
Pubblicazione: (2025)
Follow-Your-Motion: Video Motion Transfer via Efficient Spatial-Temporal Decoupled Finetuning
di: Ma, Yue, et al.
Pubblicazione: (2025)
di: Ma, Yue, et al.
Pubblicazione: (2025)
Captain Cinema: Towards Short Movie Generation
di: Xiao, Junfei, et al.
Pubblicazione: (2025)
di: Xiao, Junfei, et al.
Pubblicazione: (2025)
VoGE: A Differentiable Volume Renderer using Gaussian Ellipsoids for Analysis-by-Synthesis
di: Wang, Angtian, et al.
Pubblicazione: (2022)
di: Wang, Angtian, et al.
Pubblicazione: (2022)
SeedEdit: Align Image Re-Generation to Image Editing
di: Shi, Yichun, et al.
Pubblicazione: (2024)
di: Shi, Yichun, et al.
Pubblicazione: (2024)
Temporal2Seq: A Unified Framework for Temporal Video Understanding Tasks
di: Yang, Min, et al.
Pubblicazione: (2024)
di: Yang, Min, et al.
Pubblicazione: (2024)
Mixture of Contexts for Long Video Generation
di: Cai, Shengqu, et al.
Pubblicazione: (2025)
di: Cai, Shengqu, et al.
Pubblicazione: (2025)
Unified Static and Dynamic Network: Efficient Temporal Filtering for Video Grounding
di: Hu, Jingjing, et al.
Pubblicazione: (2024)
di: Hu, Jingjing, et al.
Pubblicazione: (2024)
Unified Medical Image Tokenizer for Autoregressive Synthesis and Understanding
di: Ma, Chenglong, et al.
Pubblicazione: (2025)
di: Ma, Chenglong, et al.
Pubblicazione: (2025)
Computer Vision and Its Relationship to Cognitive Science: A perspective from Bayes Decision Theory
di: Yuille, Alan, et al.
Pubblicazione: (2026)
di: Yuille, Alan, et al.
Pubblicazione: (2026)
VideoLoom: A Video Large Language Model for Joint Spatial-Temporal Understanding
di: Shi, Jiapeng, et al.
Pubblicazione: (2026)
di: Shi, Jiapeng, et al.
Pubblicazione: (2026)
CausalSpatial: A Benchmark for Object-Centric Causal Spatial Reasoning
di: Ma, Wenxin, et al.
Pubblicazione: (2026)
di: Ma, Wenxin, et al.
Pubblicazione: (2026)
Progressive Growing of Video Tokenizers for Temporally Compact Latent Spaces
di: Mahapatra, Aniruddha, et al.
Pubblicazione: (2025)
di: Mahapatra, Aniruddha, et al.
Pubblicazione: (2025)
Spatial Steerability of GANs via Self-Supervision from Discriminator
di: Wang, Jianyuan, et al.
Pubblicazione: (2023)
di: Wang, Jianyuan, et al.
Pubblicazione: (2023)
Long Context Tuning for Video Generation
di: Guo, Yuwei, et al.
Pubblicazione: (2025)
di: Guo, Yuwei, et al.
Pubblicazione: (2025)
SpatialReasoner: Towards Explicit and Generalizable 3D Spatial Reasoning
di: Ma, Wufei, et al.
Pubblicazione: (2025)
di: Ma, Wufei, et al.
Pubblicazione: (2025)
Memorize When Needed: Decoupled Memory Control for Spatially Consistent Long-Horizon Video Generation
di: Guo, Yanjun, et al.
Pubblicazione: (2026)
di: Guo, Yanjun, et al.
Pubblicazione: (2026)
TSdetector: Temporal-Spatial Self-correction Collaborative Learning for Colonoscopy Video Detection
di: Wang, Kaini, et al.
Pubblicazione: (2024)
di: Wang, Kaini, et al.
Pubblicazione: (2024)
Spatial-Temporal-Spectral Unified Modeling for Remote Sensing Dense Prediction
di: Zhao, Sijie, et al.
Pubblicazione: (2025)
di: Zhao, Sijie, et al.
Pubblicazione: (2025)
MetaNeRV: Meta Neural Representations for Videos with Spatial-Temporal Guidance
di: Guo, Jialong, et al.
Pubblicazione: (2025)
di: Guo, Jialong, et al.
Pubblicazione: (2025)
CamFreeDiff: Camera-free Image to Panorama Generation with Diffusion Model
di: Yuan, Xiaoding, et al.
Pubblicazione: (2024)
di: Yuan, Xiaoding, et al.
Pubblicazione: (2024)
HECTOR: Hybrid Editable Compositional Object References for Video Generation
di: Zhang, Guofeng, et al.
Pubblicazione: (2026)
di: Zhang, Guofeng, et al.
Pubblicazione: (2026)
VideoAuteur: Towards Long Narrative Video Generation
di: Xiao, Junfei, et al.
Pubblicazione: (2025)
di: Xiao, Junfei, et al.
Pubblicazione: (2025)
EvoTok: A Unified Image Tokenizer via Residual Latent Evolution for Visual Understanding and Generation
di: Li, Yan, et al.
Pubblicazione: (2026)
di: Li, Yan, et al.
Pubblicazione: (2026)
Documenti analoghi
-
Video Generation with Predictive Latents
di: Zhao, Yian, et al.
Pubblicazione: (2026) -
ViMix-14M: A Curated Multi-Source Video-Text Dataset with Long-Form, High-Quality Captions and Crawl-Free Access
di: Yang, Timing, et al.
Pubblicazione: (2025) -
M-VAR: Decoupled Scale-wise Autoregressive Modeling for High-Quality Image Generation
di: Ren, Sucheng, et al.
Pubblicazione: (2024) -
SweetTok: Semantic-Aware Spatial-Temporal Tokenizer for Compact Video Discretization
di: Tan, Zhentao, et al.
Pubblicazione: (2024) -
Video-LaVIT: Unified Video-Language Pre-training with Decoupled Visual-Motional Tokenization
di: Jin, Yang, et al.
Pubblicazione: (2024)