Semantic One-Dimensional Tokenizer for Image Reconstruction and Generation
Fuente:
arXiv
Saved in:
| Main Authors: | Qu, Yunpeng, Zhang, Kaidong, Ding, Yukang, Chen, Ying, Wang, Jian |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Mind the Generative Details: Direct Localized Detail Preference Optimization for Video Diffusion Models
by: Huang, Zitong, et al.
Published: (2026)
by: Huang, Zitong, et al.
Published: (2026)
GloTok: Global Perspective Tokenizer for Image Reconstruction and Generation
by: Zhao, Xuan, et al.
Published: (2025)
by: Zhao, Xuan, et al.
Published: (2025)
Democratizing Text-to-Image Masked Generative Models with Compact Text-Aware One-Dimensional Tokens
by: Kim, Dongwon, et al.
Published: (2025)
by: Kim, Dongwon, et al.
Published: (2025)
CaTok: Taming Mean Flows for One-Dimensional Causal Image Tokenization
by: Chen, Yitong, et al.
Published: (2026)
by: Chen, Yitong, et al.
Published: (2026)
An Image is Worth 32 Tokens for Reconstruction and Generation
by: Yu, Qihang, et al.
Published: (2024)
by: Yu, Qihang, et al.
Published: (2024)
Layton: Latent Consistency Tokenizer for 1024-pixel Image Reconstruction and Generation by 256 Tokens
by: Xie, Qingsong, et al.
Published: (2025)
by: Xie, Qingsong, et al.
Published: (2025)
LaRE$^2$: Latent Reconstruction Error Based Method for Diffusion-Generated Image Detection
by: Luo, Yunpeng, et al.
Published: (2024)
by: Luo, Yunpeng, et al.
Published: (2024)
Semantic-Aware Prefix Learning for Token-Efficient Image Generation
by: Li, Qingfeng, et al.
Published: (2026)
by: Li, Qingfeng, et al.
Published: (2026)
Improving Flexible Image Tokenizers for Autoregressive Image Generation
by: Fu, Zixuan, et al.
Published: (2026)
by: Fu, Zixuan, et al.
Published: (2026)
A Survey on 3D Human Avatar Modeling -- From Reconstruction to Generation
by: Wang, Ruihe, et al.
Published: (2024)
by: Wang, Ruihe, et al.
Published: (2024)
FreeMorph: Tuning-Free Generalized Image Morphing with Diffusion Model
by: Cao, Yukang, et al.
Published: (2025)
by: Cao, Yukang, et al.
Published: (2025)
IGGT: Instance-Grounded Geometry Transformer for Semantic 3D Reconstruction
by: Li, Hao, et al.
Published: (2025)
by: Li, Hao, et al.
Published: (2025)
Memory-Guided Point Cloud Completion for Dental Reconstruction
by: Sun, Jianan, et al.
Published: (2025)
by: Sun, Jianan, et al.
Published: (2025)
Beyond Pixel Simulation: Pathology Image Generation via Diagnostic Semantic Tokens and Prototype Control
by: Han, Minghao, et al.
Published: (2025)
by: Han, Minghao, et al.
Published: (2025)
ExpoCM: Exposure-Aware One-Step Generative Single-Image HDR Reconstruction
by: Liu, Aoyu, et al.
Published: (2026)
by: Liu, Aoyu, et al.
Published: (2026)
TokenCompose: Text-to-Image Diffusion with Token-level Supervision
by: Wang, Zirui, et al.
Published: (2023)
by: Wang, Zirui, et al.
Published: (2023)
Consistent123: One Image to Highly Consistent 3D Asset Using Case-Aware Diffusion Priors
by: Lin, Yukang, et al.
Published: (2023)
by: Lin, Yukang, et al.
Published: (2023)
SSD: Spatial-Semantic Head Decoupling for Efficient Autoregressive Image Generation
by: Jian, Siyong, et al.
Published: (2025)
by: Jian, Siyong, et al.
Published: (2025)
TokenFlow: Unified Image Tokenizer for Multimodal Understanding and Generation
by: Qu, Liao, et al.
Published: (2024)
by: Qu, Liao, et al.
Published: (2024)
WeTok: Powerful Discrete Tokenization for High-Fidelity Visual Reconstruction
by: Zhuang, Shaobin, et al.
Published: (2025)
by: Zhuang, Shaobin, et al.
Published: (2025)
ATP-LLaVA: Adaptive Token Pruning for Large Vision Language Models
by: Ye, Xubing, et al.
Published: (2024)
by: Ye, Xubing, et al.
Published: (2024)
Semantic-Aware Reconstruction Error for Detecting AI-Generated Images
by: Kang, Ju Yeon, et al.
Published: (2025)
by: Kang, Ju Yeon, et al.
Published: (2025)
Switchable Token-Specific Codebook Quantization For Face Image Compression
by: Wang, Yongbo, et al.
Published: (2025)
by: Wang, Yongbo, et al.
Published: (2025)
One Token, Two Fates: A Unified Framework via Vision Token Manipulation Against MLLMs Hallucination
by: Fa, Zhan, et al.
Published: (2026)
by: Fa, Zhan, et al.
Published: (2026)
A Frame is Worth One Token: Efficient Generative World Modeling with Delta Tokens
by: Kerssies, Tommie, et al.
Published: (2026)
by: Kerssies, Tommie, et al.
Published: (2026)
MUC: Mixture of Uncalibrated Cameras for Robust 3D Human Body Reconstruction
by: Zhu, Yitao, et al.
Published: (2024)
by: Zhu, Yitao, et al.
Published: (2024)
HieraTok: Multi-Scale Visual Tokenizer Improves Image Reconstruction and Generation
by: Chen, Cong, et al.
Published: (2025)
by: Chen, Cong, et al.
Published: (2025)
Generative Video Compression with One-Dimensional Latent Representation
by: Zheng, Zihan, et al.
Published: (2026)
by: Zheng, Zihan, et al.
Published: (2026)
GSAlign: Geometric and Semantic Alignment Network for Aerial-Ground Person Re-Identification
by: Li, Qiao, et al.
Published: (2025)
by: Li, Qiao, et al.
Published: (2025)
Image Understanding Makes for A Good Tokenizer for Image Generation
by: Wang, Luting, et al.
Published: (2024)
by: Wang, Luting, et al.
Published: (2024)
Learning Dual Transformers for All-In-One Image Restoration from a Frequency Perspective
by: Chu, Jie, et al.
Published: (2024)
by: Chu, Jie, et al.
Published: (2024)
Towards Interactive Image Inpainting via Sketch Refinement
by: Liu, Chang, et al.
Published: (2023)
by: Liu, Chang, et al.
Published: (2023)
Both Semantics and Reconstruction Matter: Making Representation Encoders Ready for Text-to-Image Generation and Editing
by: Zhang, Shilong, et al.
Published: (2025)
by: Zhang, Shilong, et al.
Published: (2025)
Semantic Iterative Reconstruction: One-Shot Universal Anomaly Detection
by: Zhu, Ning
Published: (2026)
by: Zhu, Ning
Published: (2026)
End-to-End Autoregressive Image Generation with 1D Semantic Tokenizer
by: Chu, Wenda, et al.
Published: (2026)
by: Chu, Wenda, et al.
Published: (2026)
A Survey of fMRI to Image Reconstruction
by: Guo, Weiyu, et al.
Published: (2025)
by: Guo, Weiyu, et al.
Published: (2025)
TROPHIES: Temporal Reconstruction of Places, Humans, and Cameras from Multi-view Videos
by: Liu, Jinpeng, et al.
Published: (2026)
by: Liu, Jinpeng, et al.
Published: (2026)
LaCon: Late-Constraint Diffusion for Steerable Guided Image Synthesis
by: Liu, Chang, et al.
Published: (2023)
by: Liu, Chang, et al.
Published: (2023)
ImageFolder: Autoregressive Image Generation with Folded Tokens
by: Li, Xiang, et al.
Published: (2024)
by: Li, Xiang, et al.
Published: (2024)
Generalizable Whole Slide Image Classification with Fine-Grained Visual-Semantic Interaction
by: Li, Hao, et al.
Published: (2024)
by: Li, Hao, et al.
Published: (2024)
Similar Items
-
Mind the Generative Details: Direct Localized Detail Preference Optimization for Video Diffusion Models
by: Huang, Zitong, et al.
Published: (2026) -
GloTok: Global Perspective Tokenizer for Image Reconstruction and Generation
by: Zhao, Xuan, et al.
Published: (2025) -
Democratizing Text-to-Image Masked Generative Models with Compact Text-Aware One-Dimensional Tokens
by: Kim, Dongwon, et al.
Published: (2025) -
CaTok: Taming Mean Flows for One-Dimensional Causal Image Tokenization
by: Chen, Yitong, et al.
Published: (2026) -
An Image is Worth 32 Tokens for Reconstruction and Generation
by: Yu, Qihang, et al.
Published: (2024)