WinTok: A Win-Win Hybrid Tokenizer via Decomposing Visual Understanding and Generation with Transferable Tokens
Fuente:
arXiv
Saved in:
| Main Authors: | Guo, Yiwei, Zhuang, Shaobin, Huang, Zhipeng, Fu, Canmiao, Li, Chen, Lyu, Jing, Wang, Yali |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
WeTok: Powerful Discrete Tokenization for High-Fidelity Visual Reconstruction
by: Zhuang, Shaobin, et al.
Published: (2025)
by: Zhuang, Shaobin, et al.
Published: (2025)
Random Wins All: Rethinking Grouping Strategies for Vision Tokens
by: Fan, Qihang, et al.
Published: (2026)
by: Fan, Qihang, et al.
Published: (2026)
WinT3R: Window-Based Streaming Reconstruction with Camera Token Pool
by: Li, Zizun, et al.
Published: (2025)
by: Li, Zizun, et al.
Published: (2025)
Win-Win: Training High-Resolution Vision Transformers from Two Windows
by: Leroy, Vincent, et al.
Published: (2023)
by: Leroy, Vincent, et al.
Published: (2023)
UniFlow: A Unified Pixel Flow Tokenizer for Visual Understanding and Generation
by: Yue, Zhengrong, et al.
Published: (2025)
by: Yue, Zhengrong, et al.
Published: (2025)
TransAgent: Transfer Vision-Language Foundation Models with Heterogeneous Agent Collaboration
by: Guo, Yiwei, et al.
Published: (2024)
by: Guo, Yiwei, et al.
Published: (2024)
Revealing the Two Sides of Data Augmentation: An Asymmetric Distillation-based Win-Win Solution for Open-Set Recognition
by: Jia, Yunbing, et al.
Published: (2024)
by: Jia, Yunbing, et al.
Published: (2024)
Video-GPT via Next Clip Diffusion
by: Zhuang, Shaobin, et al.
Published: (2025)
by: Zhuang, Shaobin, et al.
Published: (2025)
WeGen: A Unified Model for Interactive Multimodal Generation as We Chat
by: Huang, Zhipeng, et al.
Published: (2025)
by: Huang, Zhipeng, et al.
Published: (2025)
UniTok: A Unified Tokenizer for Visual Generation and Understanding
by: Ma, Chuofan, et al.
Published: (2025)
by: Ma, Chuofan, et al.
Published: (2025)
NativeTok: Native Visual Tokenization for Improved Image Generation
by: Wu, Bin, et al.
Published: (2026)
by: Wu, Bin, et al.
Published: (2026)
DINO-Tok: Adapting DINO for Visual Tokenizers
by: Jia, Mingkai, et al.
Published: (2025)
by: Jia, Mingkai, et al.
Published: (2025)
MergeTok: Unified Continuous and Discrete Visual Tokenization via Token Merging
by: Zhang, Luyuan, et al.
Published: (2026)
by: Zhang, Luyuan, et al.
Published: (2026)
Get In Video: Add Anything You Want to the Video
by: Zhuang, Shaobin, et al.
Published: (2025)
by: Zhuang, Shaobin, et al.
Published: (2025)
Evaluation of Winning Solutions of 2025 Low Power Computer Vision Challenge
by: Ye, Zihao, et al.
Published: (2026)
by: Ye, Zihao, et al.
Published: (2026)
The Championship-Winning Solution for the 5th CLVISION Challenge 2024
by: Pan, Sishun, et al.
Published: (2024)
by: Pan, Sishun, et al.
Published: (2024)
EvoTok: A Unified Image Tokenizer via Residual Latent Evolution for Visual Understanding and Generation
by: Li, Yan, et al.
Published: (2026)
by: Li, Yan, et al.
Published: (2026)
Text-guided Visual Prompt DINO for Generic Segmentation
by: Guan, Yuchen, et al.
Published: (2025)
by: Guan, Yuchen, et al.
Published: (2025)
Layer Pruning with Consensus: A Triple-Win Solution
by: Mugnaini, Leandro Giusti, et al.
Published: (2024)
by: Mugnaini, Leandro Giusti, et al.
Published: (2024)
Predicting Winning Captions for Weekly New Yorker Comics
by: Cao, Stanley, et al.
Published: (2024)
by: Cao, Stanley, et al.
Published: (2024)
WinSyn: A High Resolution Testbed for Synthetic Data
by: Kelly, Tom, et al.
Published: (2023)
by: Kelly, Tom, et al.
Published: (2023)
CAIT: Triple-Win Compression towards High Accuracy, Fast Inference, and Favorable Transferability For ViTs
by: Wang, Ao, et al.
Published: (2023)
by: Wang, Ao, et al.
Published: (2023)
PhaseWin Search Framework Enable Efficient Object-Level Interpretation
by: Gu, Zihan, et al.
Published: (2025)
by: Gu, Zihan, et al.
Published: (2025)
TokBench: Evaluating Your Visual Tokenizer before Visual Generation
by: Wu, Junfeng, et al.
Published: (2025)
by: Wu, Junfeng, et al.
Published: (2025)
SliceFine: The Universal Winning-Slice Hypothesis for Pretrained Networks
by: Kowsher, Md, et al.
Published: (2025)
by: Kowsher, Md, et al.
Published: (2025)
GloTok: Global Perspective Tokenizer for Image Reconstruction and Generation
by: Zhao, Xuan, et al.
Published: (2025)
by: Zhao, Xuan, et al.
Published: (2025)
GigaTok: Scaling Visual Tokenizers to 3 Billion Parameters for Autoregressive Image Generation
by: Xiong, Tianwei, et al.
Published: (2025)
by: Xiong, Tianwei, et al.
Published: (2025)
BitDance: Scaling Autoregressive Generative Models with Binary Tokens
by: Ai, Yuang, et al.
Published: (2026)
by: Ai, Yuang, et al.
Published: (2026)
TrajTok: Learning Trajectory Tokens enables better Video Understanding
by: Zheng, Chenhao, et al.
Published: (2026)
by: Zheng, Chenhao, et al.
Published: (2026)
MVP: Winning Solution to SMP Challenge 2025 Video Track
by: Ye, Liliang, et al.
Published: (2025)
by: Ye, Liliang, et al.
Published: (2025)
UniCompress: Token Compression for Unified Vision-Language Understanding and Generation
by: Wang, Ziyao, et al.
Published: (2026)
by: Wang, Ziyao, et al.
Published: (2026)
UniWeTok: An Unified Binary Tokenizer with Codebook Size $\mathit{2^{128}}$ for Unified Multimodal Large Language Model
by: Zhuang, Shaobin, et al.
Published: (2026)
by: Zhuang, Shaobin, et al.
Published: (2026)
DaWin: Training-free Dynamic Weight Interpolation for Robust Adaptation
by: Oh, Changdae, et al.
Published: (2024)
by: Oh, Changdae, et al.
Published: (2024)
DynTok: Dynamic Compression of Visual Tokens for Efficient and Effective Video Understanding
by: Zhang, Hongzhi, et al.
Published: (2025)
by: Zhang, Hongzhi, et al.
Published: (2025)
pMoE: Prompting Diverse Experts Together Wins More in Visual Adaptation
by: Mo, Shentong, et al.
Published: (2026)
by: Mo, Shentong, et al.
Published: (2026)
Structured Semantic 3D Reconstruction (S23DR) Challenge 2025 -- Winning solution
by: Skvrna, Jan, et al.
Published: (2025)
by: Skvrna, Jan, et al.
Published: (2025)
PyraTok: Language-Aligned Pyramidal Tokenizer for Video Understanding and Generation
by: Susladkar, Onkar, et al.
Published: (2026)
by: Susladkar, Onkar, et al.
Published: (2026)
Continual Learning: Forget-free Winning Subnetworks for Video Representations
by: Kang, Haeyong, et al.
Published: (2023)
by: Kang, Haeyong, et al.
Published: (2023)
HieraTok: Multi-Scale Visual Tokenizer Improves Image Reconstruction and Generation
by: Chen, Cong, et al.
Published: (2025)
by: Chen, Cong, et al.
Published: (2025)
RefTok: Reference-Based Tokenization for Video Generation
by: Fan, Xiang, et al.
Published: (2025)
by: Fan, Xiang, et al.
Published: (2025)
Similar Items
-
WeTok: Powerful Discrete Tokenization for High-Fidelity Visual Reconstruction
by: Zhuang, Shaobin, et al.
Published: (2025) -
Random Wins All: Rethinking Grouping Strategies for Vision Tokens
by: Fan, Qihang, et al.
Published: (2026) -
WinT3R: Window-Based Streaming Reconstruction with Camera Token Pool
by: Li, Zizun, et al.
Published: (2025) -
Win-Win: Training High-Resolution Vision Transformers from Two Windows
by: Leroy, Vincent, et al.
Published: (2023) -
UniFlow: A Unified Pixel Flow Tokenizer for Visual Understanding and Generation
by: Yue, Zhengrong, et al.
Published: (2025)