When Worse is Better: Navigating the compression-generation tradeoff in visual tokenization
Fuente:
arXiv
Saved in:
| Main Authors: | Ramanujan, Vivek, Tirumala, Kushal, Aghajanyan, Armen, Zettlemoyer, Luke, Farhadi, Ali |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
VideoNet: A Large-Scale Dataset for Domain-Specific Action Recognition
by: Yadav, Tanush, et al.
Published: (2026)
by: Yadav, Tanush, et al.
Published: (2026)
For Better or For Worse? Learning Minimum Variance Features With Label Augmentation
by: Chidambaram, Muthu, et al.
Published: (2024)
by: Chidambaram, Muthu, et al.
Published: (2024)
From an Image to a Scene: Learning to Imagine the World from a Million 360 Videos
by: Wallingford, Matthew, et al.
Published: (2024)
by: Wallingford, Matthew, et al.
Published: (2024)
Matryoshka Representation Learning
by: Kusupati, Aditya, et al.
Published: (2022)
by: Kusupati, Aditya, et al.
Published: (2022)
CAT: Content-Adaptive Image Tokenization
by: Shen, Junhong, et al.
Published: (2025)
by: Shen, Junhong, et al.
Published: (2025)
Contrastive Flow Matching
by: Stoica, George, et al.
Published: (2025)
by: Stoica, George, et al.
Published: (2025)
Few-shot multi-token DreamBooth with LoRa for style-consistent character generation
by: Pascual, Ruben, et al.
Published: (2025)
by: Pascual, Ruben, et al.
Published: (2025)
LookupViT: Compressing visual information to a limited number of tokens
by: Koner, Rajat, et al.
Published: (2024)
by: Koner, Rajat, et al.
Published: (2024)
Comparison of Autoencoders for tokenization of ASL datasets
by: Praun-Petrovic, Vouk, et al.
Published: (2025)
by: Praun-Petrovic, Vouk, et al.
Published: (2025)
Bayesian computation with generative diffusion models by Multilevel Monte Carlo
by: Haji-Ali, Abdul-Lateef, et al.
Published: (2024)
by: Haji-Ali, Abdul-Lateef, et al.
Published: (2024)
Reconstruction Alignment Improves Unified Multimodal Models
by: Xie, Ji, et al.
Published: (2025)
by: Xie, Ji, et al.
Published: (2025)
Using Machine Learning for move sequence visualization and generation in climbing
by: Rimbot, Thomas, et al.
Published: (2025)
by: Rimbot, Thomas, et al.
Published: (2025)
Reducing catastrophic forgetting of incremental learning in the absence of rehearsal memory with task-specific token
by: Choi, Young Jo, et al.
Published: (2024)
by: Choi, Young Jo, et al.
Published: (2024)
SCE-LITE-HQ: Smooth visual counterfactual explanations with generative foundation models
by: Zeid, Ahmed, et al.
Published: (2026)
by: Zeid, Ahmed, et al.
Published: (2026)
Register and [CLS] tokens yield a decoupling of local and global features in large ViTs
by: Lappe, Alexander, et al.
Published: (2025)
by: Lappe, Alexander, et al.
Published: (2025)
Diffusion Forcing: Next-token Prediction Meets Full-Sequence Diffusion
by: Chen, Boyuan, et al.
Published: (2024)
by: Chen, Boyuan, et al.
Published: (2024)
LookWhen? Fast Video Recognition by Learning When, Where, and What to Compute
by: Salamatian, Ali, et al.
Published: (2026)
by: Salamatian, Ali, et al.
Published: (2026)
This Looks Better than That: Better Interpretable Models with ProtoPNeXt
by: Willard, Frank, et al.
Published: (2024)
by: Willard, Frank, et al.
Published: (2024)
Mixture-of-Mamba: Enhancing Multi-Modal State-Space Models with Modality-Aware Sparsity
by: Liang, Weixin, et al.
Published: (2025)
by: Liang, Weixin, et al.
Published: (2025)
Posterior Augmented Flow Matching
by: Stoica, George, et al.
Published: (2026)
by: Stoica, George, et al.
Published: (2026)
Transfusion: Predict the Next Token and Diffuse Images with One Multi-Modal Model
by: Zhou, Chunting, et al.
Published: (2024)
by: Zhou, Chunting, et al.
Published: (2024)
Text Quality-Based Pruning for Efficient Training of Language Models
by: Sharma, Vasu, et al.
Published: (2024)
by: Sharma, Vasu, et al.
Published: (2024)
tenSVD algorithm for compression
by: Gallo, Michele
Published: (2025)
by: Gallo, Michele
Published: (2025)
ViTok-v2: Scaling Native Resolution Auto-Encoders to 5 Billion Parameters
by: Hansen-Estruch, Philippe, et al.
Published: (2026)
by: Hansen-Estruch, Philippe, et al.
Published: (2026)
Sample what you cant compress
by: Birodkar, Vighnesh, et al.
Published: (2024)
by: Birodkar, Vighnesh, et al.
Published: (2024)
CycleGAN with Better Cycles
by: Wang, Tongzhou, et al.
Published: (2024)
by: Wang, Tongzhou, et al.
Published: (2024)
Imitating What Works: Simulation-Filtered Modular Policy Learning from Human Videos
by: Zhai, Albert J., et al.
Published: (2026)
by: Zhai, Albert J., et al.
Published: (2026)
Variational Bayes image restoration with compressive autoencoders
by: Biquard, Maud, et al.
Published: (2023)
by: Biquard, Maud, et al.
Published: (2023)
Improving MoE Compute Efficiency by Composing Weight and Data Sparsity
by: Kilian, Maciej, et al.
Published: (2026)
by: Kilian, Maciej, et al.
Published: (2026)
Towards Better Understanding Attribution Methods
by: Rao, Sukrut, et al.
Published: (2022)
by: Rao, Sukrut, et al.
Published: (2022)
Unpaired Translation of Point Clouds for Modeling Detector Response
by: Li, Mingyang, et al.
Published: (2025)
by: Li, Mingyang, et al.
Published: (2025)
Improving Visual Grounding by Encouraging Consistent Gradient-based Explanations
by: Yang, Ziyan, et al.
Published: (2022)
by: Yang, Ziyan, et al.
Published: (2022)
Robustly overfitting latents for flexible neural image compression
by: Perugachi-Diaz, Yura, et al.
Published: (2024)
by: Perugachi-Diaz, Yura, et al.
Published: (2024)
Are Object-Centric Representations Better At Compositional Generalization?
by: Kapl, Ferdinand, et al.
Published: (2026)
by: Kapl, Ferdinand, et al.
Published: (2026)
LMFusion: Adapting Pretrained Language Models for Multimodal Generation
by: Shi, Weijia, et al.
Published: (2024)
by: Shi, Weijia, et al.
Published: (2024)
Are Bias Mitigation Techniques for Deep Learning Effective?
by: Shrestha, Robik, et al.
Published: (2021)
by: Shrestha, Robik, et al.
Published: (2021)
Robustness to distribution shifts of compressed networks for edge devices
by: Shen, Lulan, et al.
Published: (2024)
by: Shen, Lulan, et al.
Published: (2024)
Efficient training for compact compression models via sequential distillation
by: Rodrigues, Caroline Mazini, et al.
Published: (2026)
by: Rodrigues, Caroline Mazini, et al.
Published: (2026)
Disentangling Mean Embeddings for Better Diagnostics of Image Generators
by: Gruber, Sebastian G., et al.
Published: (2024)
by: Gruber, Sebastian G., et al.
Published: (2024)
Better Together: Evaluating the Complementarity of Earth Embedding Models
by: van der Plas, Thijs L, et al.
Published: (2026)
by: van der Plas, Thijs L, et al.
Published: (2026)
Similar Items
-
VideoNet: A Large-Scale Dataset for Domain-Specific Action Recognition
by: Yadav, Tanush, et al.
Published: (2026) -
For Better or For Worse? Learning Minimum Variance Features With Label Augmentation
by: Chidambaram, Muthu, et al.
Published: (2024) -
From an Image to a Scene: Learning to Imagine the World from a Million 360 Videos
by: Wallingford, Matthew, et al.
Published: (2024) -
Matryoshka Representation Learning
by: Kusupati, Aditya, et al.
Published: (2022) -
CAT: Content-Adaptive Image Tokenization
by: Shen, Junhong, et al.
Published: (2025)