OmniTokenizer: A Joint Image-Video Tokenizer for Visual Generation
Fuente:
arXiv
Salvato in:
| Autori principali: | Wang, Junke, Jiang, Yi, Yuan, Zehuan, Peng, Binyue, Wu, Zuxuan, Jiang, Yu-Gang |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
OmniVid: A Generative Framework for Universal Video Understanding
di: Wang, Junke, et al.
Pubblicazione: (2024)
di: Wang, Junke, et al.
Pubblicazione: (2024)
Rethinking Discrete Tokens: Treating Them as Conditions for Continuous Autoregressive Image Synthesis
di: Zheng, Peng, et al.
Pubblicazione: (2025)
di: Zheng, Peng, et al.
Pubblicazione: (2025)
OmniTracker: Unifying Object Tracking by Tracking-with-Detection
di: Wang, Junke, et al.
Pubblicazione: (2023)
di: Wang, Junke, et al.
Pubblicazione: (2023)
UniTok: A Unified Tokenizer for Visual Generation and Understanding
di: Ma, Chuofan, et al.
Pubblicazione: (2025)
di: Ma, Chuofan, et al.
Pubblicazione: (2025)
CaTok: Taming Mean Flows for One-Dimensional Causal Image Tokenization
di: Chen, Yitong, et al.
Pubblicazione: (2026)
di: Chen, Yitong, et al.
Pubblicazione: (2026)
DeRA: Decoupled Representation Alignment for Video Tokenization
di: Guo, Pengbo, et al.
Pubblicazione: (2025)
di: Guo, Pengbo, et al.
Pubblicazione: (2025)
VideoLoom: A Video Large Language Model for Joint Spatial-Temporal Understanding
di: Shi, Jiapeng, et al.
Pubblicazione: (2026)
di: Shi, Jiapeng, et al.
Pubblicazione: (2026)
TokenFlow: Unified Image Tokenizer for Multimodal Understanding and Generation
di: Qu, Liao, et al.
Pubblicazione: (2024)
di: Qu, Liao, et al.
Pubblicazione: (2024)
SimpleAR: Pushing the Frontier of Autoregressive Visual Generation through Pretraining, SFT, and RL
di: Wang, Junke, et al.
Pubblicazione: (2025)
di: Wang, Junke, et al.
Pubblicazione: (2025)
Daily-Omni: Towards Audio-Visual Reasoning with Temporal Alignment across Modalities
di: Zhou, Ziwei, et al.
Pubblicazione: (2025)
di: Zhou, Ziwei, et al.
Pubblicazione: (2025)
DeepStack: Deeply Stacking Visual Tokens is Surprisingly Simple and Effective for LMMs
di: Meng, Lingchen, et al.
Pubblicazione: (2024)
di: Meng, Lingchen, et al.
Pubblicazione: (2024)
Groma: Localized Visual Tokenization for Grounding Multimodal Large Language Models
di: Ma, Chuofan, et al.
Pubblicazione: (2024)
di: Ma, Chuofan, et al.
Pubblicazione: (2024)
Visual Autoregressive Modeling: Scalable Image Generation via Next-Scale Prediction
di: Tian, Keyu, et al.
Pubblicazione: (2024)
di: Tian, Keyu, et al.
Pubblicazione: (2024)
A Glimpse to Compress: Dynamic Visual Token Pruning for Large Vision-Language Models
di: Zeng, Quan-Sheng, et al.
Pubblicazione: (2025)
di: Zeng, Quan-Sheng, et al.
Pubblicazione: (2025)
OmniLottie: Generating Vector Animations via Parameterized Lottie Tokens
di: Yang, Yiying, et al.
Pubblicazione: (2026)
di: Yang, Yiying, et al.
Pubblicazione: (2026)
Omni-RGPT: Unifying Image and Video Region-level Understanding via Token Marks
di: Heo, Miran, et al.
Pubblicazione: (2025)
di: Heo, Miran, et al.
Pubblicazione: (2025)
GenRec: Unifying Video Generation and Recognition with Diffusion Models
di: Weng, Zejia, et al.
Pubblicazione: (2024)
di: Weng, Zejia, et al.
Pubblicazione: (2024)
Learning Accurate Segmentation Purely from Self-Supervision
di: You, Zuyao, et al.
Pubblicazione: (2026)
di: You, Zuyao, et al.
Pubblicazione: (2026)
Baton: Explicit Semantic Blueprints for Joint Video-Audio Generation
di: Tu, Shuyuan, et al.
Pubblicazione: (2026)
di: Tu, Shuyuan, et al.
Pubblicazione: (2026)
Cubic Discrete Diffusion: Discrete Visual Generation on High-Dimensional Representation Tokens
di: Wang, Yuqing, et al.
Pubblicazione: (2026)
di: Wang, Yuqing, et al.
Pubblicazione: (2026)
DCDM: Divide-and-Conquer Diffusion Models for Consistency-Preserving Video Generation
di: Zhao, Haoyu, et al.
Pubblicazione: (2026)
di: Zhao, Haoyu, et al.
Pubblicazione: (2026)
Learning to Rank Patches for Unbiased Image Redundancy Reduction
di: Luo, Yang, et al.
Pubblicazione: (2024)
di: Luo, Yang, et al.
Pubblicazione: (2024)
InfinityStar: Unified Spacetime AutoRegressive Modeling for Visual Generation
di: Liu, Jinlai, et al.
Pubblicazione: (2025)
di: Liu, Jinlai, et al.
Pubblicazione: (2025)
CreatiLayout: Siamese Multimodal Diffusion Transformer for Creative Layout-to-Image Generation
di: Zhang, Hui, et al.
Pubblicazione: (2024)
di: Zhang, Hui, et al.
Pubblicazione: (2024)
Autoregressive Model Beats Diffusion: Llama for Scalable Image Generation
di: Sun, Peize, et al.
Pubblicazione: (2024)
di: Sun, Peize, et al.
Pubblicazione: (2024)
TempoMaster: Efficient Long Video Generation via Next-Frame-Rate Prediction
di: Ma, Yukuo, et al.
Pubblicazione: (2025)
di: Ma, Yukuo, et al.
Pubblicazione: (2025)
Vision Foundation Models as Effective Visual Tokenizers for Autoregressive Image Generation
di: Zheng, Anlin, et al.
Pubblicazione: (2025)
di: Zheng, Anlin, et al.
Pubblicazione: (2025)
Importance-Based Token Merging for Efficient Image and Video Generation
di: Wu, Haoyu, et al.
Pubblicazione: (2024)
di: Wu, Haoyu, et al.
Pubblicazione: (2024)
NativeTok: Native Visual Tokenization for Improved Image Generation
di: Wu, Bin, et al.
Pubblicazione: (2026)
di: Wu, Bin, et al.
Pubblicazione: (2026)
Generative Refinement Networks for Visual Synthesis
di: Han, Jian, et al.
Pubblicazione: (2026)
di: Han, Jian, et al.
Pubblicazione: (2026)
UniToken: Harmonizing Multimodal Understanding and Generation through Unified Visual Encoding
di: Jiao, Yang, et al.
Pubblicazione: (2025)
di: Jiao, Yang, et al.
Pubblicazione: (2025)
Factorized Visual Tokenization and Generation
di: Bai, Zechen, et al.
Pubblicazione: (2024)
di: Bai, Zechen, et al.
Pubblicazione: (2024)
FluxMem: Adaptive Hierarchical Memory for Streaming Video Understanding
di: Xie, Yiweng, et al.
Pubblicazione: (2026)
di: Xie, Yiweng, et al.
Pubblicazione: (2026)
Vision Foundation Models as Generalist Tokenizers for Image Generation
di: Zheng, Anlin, et al.
Pubblicazione: (2026)
di: Zheng, Anlin, et al.
Pubblicazione: (2026)
CoMP: Continual Multimodal Pre-training for Vision Foundation Models
di: Chen, Yitong, et al.
Pubblicazione: (2025)
di: Chen, Yitong, et al.
Pubblicazione: (2025)
AID: Adapting Image2Video Diffusion Models for Instruction-guided Video Prediction
di: Xing, Zhen, et al.
Pubblicazione: (2024)
di: Xing, Zhen, et al.
Pubblicazione: (2024)
Pix2Cap-COCO: Advancing Visual Comprehension via Pixel-Level Captioning
di: You, Zuyao, et al.
Pubblicazione: (2025)
di: You, Zuyao, et al.
Pubblicazione: (2025)
Video-KTR: Reinforcing Video Reasoning via Key Token Attribution
di: Wang, Ziyue, et al.
Pubblicazione: (2026)
di: Wang, Ziyue, et al.
Pubblicazione: (2026)
TokensGen: Harnessing Condensed Tokens for Long Video Generation
di: Ouyang, Wenqi, et al.
Pubblicazione: (2025)
di: Ouyang, Wenqi, et al.
Pubblicazione: (2025)
Attention Itself Could Retrieve.RetrieveVGGT: Training-Free Long Context Streaming 3D Reconstruction via Query-Key Similarity Retrieval
di: Zou, Zichen, et al.
Pubblicazione: (2026)
di: Zou, Zichen, et al.
Pubblicazione: (2026)
Documenti analoghi
-
OmniVid: A Generative Framework for Universal Video Understanding
di: Wang, Junke, et al.
Pubblicazione: (2024) -
Rethinking Discrete Tokens: Treating Them as Conditions for Continuous Autoregressive Image Synthesis
di: Zheng, Peng, et al.
Pubblicazione: (2025) -
OmniTracker: Unifying Object Tracking by Tracking-with-Detection
di: Wang, Junke, et al.
Pubblicazione: (2023) -
UniTok: A Unified Tokenizer for Visual Generation and Understanding
di: Ma, Chuofan, et al.
Pubblicazione: (2025) -
CaTok: Taming Mean Flows for One-Dimensional Causal Image Tokenization
di: Chen, Yitong, et al.
Pubblicazione: (2026)