TokBench: Evaluating Your Visual Tokenizer before Visual Generation
Fuente:
arXiv
Saved in:
| Main Authors: | Wu, Junfeng, Luo, Dongliang, Zhao, Weizhi, Xie, Zhihao, Wang, Yuanhao, Li, Junyi, Xie, Xudong, Liu, Yuliang, Bai, Xiang |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
UniTok: A Unified Tokenizer for Visual Generation and Understanding
by: Ma, Chuofan, et al.
Published: (2025)
by: Ma, Chuofan, et al.
Published: (2025)
PartGLEE: A Foundation Model for Recognizing and Parsing Any Objects
by: Li, Junyi, et al.
Published: (2024)
by: Li, Junyi, et al.
Published: (2024)
SemiETS: Integrating Spatial and Content Consistencies for Semi-Supervised End-to-end Text Spotting
by: Luo, Dongliang, et al.
Published: (2025)
by: Luo, Dongliang, et al.
Published: (2025)
NativeTok: Native Visual Tokenization for Improved Image Generation
by: Wu, Bin, et al.
Published: (2026)
by: Wu, Bin, et al.
Published: (2026)
MergeTok: Unified Continuous and Discrete Visual Tokenization via Token Merging
by: Zhang, Luyuan, et al.
Published: (2026)
by: Zhang, Luyuan, et al.
Published: (2026)
HAIChart: Human and AI Paired Visualization System
by: Xie, Yupeng, et al.
Published: (2024)
by: Xie, Yupeng, et al.
Published: (2024)
RadixGraph: A Fast, Space-Optimized Data Structure for Dynamic Graph Storage (Extended Version)
by: Xie, Haoxuan, et al.
Published: (2026)
by: Xie, Haoxuan, et al.
Published: (2026)
[Extended Version] ArceKV: Towards Workload-driven LSM-compactions for Key-Value Store Under Dynamic Workloads
by: Liu, Junfeng, et al.
Published: (2025)
by: Liu, Junfeng, et al.
Published: (2025)
DINO-Tok: Adapting DINO for Visual Tokenizers
by: Jia, Mingkai, et al.
Published: (2025)
by: Jia, Mingkai, et al.
Published: (2025)
Towards a Flexible Scale-out Framework for Efficient Visual Data Query Processing
by: Verma, Rohit, et al.
Published: (2024)
by: Verma, Rohit, et al.
Published: (2024)
MSTAR: Box-free Multi-query Scene Text Retrieval with Attention Recycling
by: Yin, Liang, et al.
Published: (2025)
by: Yin, Liang, et al.
Published: (2025)
TokLIP: Marry Visual Tokens to CLIP for Multimodal Comprehension and Generation
by: Lin, Haokun, et al.
Published: (2025)
by: Lin, Haokun, et al.
Published: (2025)
HieraTok: Multi-Scale Visual Tokenizer Improves Image Reconstruction and Generation
by: Chen, Cong, et al.
Published: (2025)
by: Chen, Cong, et al.
Published: (2025)
GigaTok: Scaling Visual Tokenizers to 3 Billion Parameters for Autoregressive Image Generation
by: Xiong, Tianwei, et al.
Published: (2025)
by: Xiong, Tianwei, et al.
Published: (2025)
A Formalism and Library for Database Visualization
by: Wu, Eugene, et al.
Published: (2025)
by: Wu, Eugene, et al.
Published: (2025)
Exploring Agentic Visual Analytics: A Co-Evolutionary Framework of Roles and Workflows
by: Luo, Tianqi, et al.
Published: (2026)
by: Luo, Tianqi, et al.
Published: (2026)
EvoTok: A Unified Image Tokenizer via Residual Latent Evolution for Visual Understanding and Generation
by: Li, Yan, et al.
Published: (2026)
by: Li, Yan, et al.
Published: (2026)
FineState-Bench: Benchmarking State-Conditioned Grounding for Fine-grained GUI State Setting
by: Ji, Fengxian, et al.
Published: (2026)
by: Ji, Fengxian, et al.
Published: (2026)
WinTok: A Win-Win Hybrid Tokenizer via Decomposing Visual Understanding and Generation with Transferable Tokens
by: Guo, Yiwei, et al.
Published: (2026)
by: Guo, Yiwei, et al.
Published: (2026)
RefTok: Reference-Based Tokenization for Video Generation
by: Fan, Xiang, et al.
Published: (2025)
by: Fan, Xiang, et al.
Published: (2025)
WeTok: Powerful Discrete Tokenization for High-Fidelity Visual Reconstruction
by: Zhuang, Shaobin, et al.
Published: (2025)
by: Zhuang, Shaobin, et al.
Published: (2025)
AlignTok: Aligning Visual Foundation Encoders to Tokenizers for Diffusion Models
by: Chen, Bowei, et al.
Published: (2025)
by: Chen, Bowei, et al.
Published: (2025)
VectorMaton: Efficient Vector Search with Pattern Constraints via an Enhanced Suffix Automaton
by: Xie, Haoxuan, et al.
Published: (2026)
by: Xie, Haoxuan, et al.
Published: (2026)
ResTok: Learning Hierarchical Residuals in 1D Visual Tokenizers for Autoregressive Image Generation
by: Zhang, Xu, et al.
Published: (2026)
by: Zhang, Xu, et al.
Published: (2026)
From Tokens to Materials: Leveraging Language Models for Scientific Discovery
by: Wan, Yuwei, et al.
Published: (2024)
by: Wan, Yuwei, et al.
Published: (2024)
ResBench: A Comprehensive Framework for Evaluating Database Resilience
by: Hu, Puyun, et al.
Published: (2025)
by: Hu, Puyun, et al.
Published: (2025)
VALLR-Pin: Uncertainty-Factorized Visual Speech Recognition for Mandarin with Pinyin Guidance
by: Sun, Chang, et al.
Published: (2025)
by: Sun, Chang, et al.
Published: (2025)
Progressive Evolution from Single-Point to Polygon for Scene Text
by: Deng, Linger, et al.
Published: (2023)
by: Deng, Linger, et al.
Published: (2023)
From Words to Structured Visuals: A Benchmark and Framework for Text-to-Diagram Generation and Editing
by: Wei, Jingxuan, et al.
Published: (2024)
by: Wei, Jingxuan, et al.
Published: (2024)
Factorized Visual Tokenization and Generation
by: Bai, Zechen, et al.
Published: (2024)
by: Bai, Zechen, et al.
Published: (2024)
RSL-SQL: Robust Schema Linking in Text-to-SQL Generation
by: Cao, Zhenbiao, et al.
Published: (2024)
by: Cao, Zhenbiao, et al.
Published: (2024)
Saliency-Bench: A Comprehensive Benchmark for Evaluating Visual Explanations
by: Zhang, Yifei, et al.
Published: (2023)
by: Zhang, Yifei, et al.
Published: (2023)
Liquid: Language Models are Scalable and Unified Multi-modal Generators
by: Wu, Junfeng, et al.
Published: (2024)
by: Wu, Junfeng, et al.
Published: (2024)
DynTok: Dynamic Compression of Visual Tokens for Efficient and Effective Video Understanding
by: Zhang, Hongzhi, et al.
Published: (2025)
by: Zhang, Hongzhi, et al.
Published: (2025)
Database Theory + X: Database Visualization
by: Wu, Eugene
Published: (2024)
by: Wu, Eugene
Published: (2024)
OS-W2S: An Automatic Labeling Engine for Language-Guided Open-Set Aerial Object Detection
by: Wei, Guoting, et al.
Published: (2025)
by: Wei, Guoting, et al.
Published: (2025)
Aster: Enhancing LSM-structures for Scalable Graph Database
by: Mo, Dingheng, et al.
Published: (2025)
by: Mo, Dingheng, et al.
Published: (2025)
LipGen: Viseme-Guided Lip Video Generation for Enhancing Visual Speech Recognition
by: Hao, Bowen, et al.
Published: (2025)
by: Hao, Bowen, et al.
Published: (2025)
VTBench: Evaluating Visual Tokenizers for Autoregressive Image Generation
by: Lin, Huawei, et al.
Published: (2025)
by: Lin, Huawei, et al.
Published: (2025)
Towards Cost-effective LLMs Routing with Batch Prompting
by: Xu, Haotian, et al.
Published: (2026)
by: Xu, Haotian, et al.
Published: (2026)
Similar Items
-
UniTok: A Unified Tokenizer for Visual Generation and Understanding
by: Ma, Chuofan, et al.
Published: (2025) -
PartGLEE: A Foundation Model for Recognizing and Parsing Any Objects
by: Li, Junyi, et al.
Published: (2024) -
SemiETS: Integrating Spatial and Content Consistencies for Semi-Supervised End-to-end Text Spotting
by: Luo, Dongliang, et al.
Published: (2025) -
NativeTok: Native Visual Tokenization for Improved Image Generation
by: Wu, Bin, et al.
Published: (2026) -
MergeTok: Unified Continuous and Discrete Visual Tokenization via Token Merging
by: Zhang, Luyuan, et al.
Published: (2026)