CAT: Content-Adaptive Image Tokenization
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Shen, Junhong, Tirumala, Kushal, Yasunaga, Michihiro, Misra, Ishan, Zettlemoyer, Luke, Yu, Lili, Zhou, Chunting |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Transfusion: Predict the Next Token and Diffuse Images with One Multi-Modal Model
von: Zhou, Chunting, et al.
Veröffentlicht: (2024)
von: Zhou, Chunting, et al.
Veröffentlicht: (2024)
Multimodal RewardBench: Holistic Evaluation of Reward Models for Vision Language Models
von: Yasunaga, Michihiro, et al.
Veröffentlicht: (2025)
von: Yasunaga, Michihiro, et al.
Veröffentlicht: (2025)
When Worse is Better: Navigating the compression-generation tradeoff in visual tokenization
von: Ramanujan, Vivek, et al.
Veröffentlicht: (2024)
von: Ramanujan, Vivek, et al.
Veröffentlicht: (2024)
Computational Tradeoffs in Image Synthesis: Diffusion, Masked-Token, and Next-Token Prediction
von: Kilian, Maciej, et al.
Veröffentlicht: (2024)
von: Kilian, Maciej, et al.
Veröffentlicht: (2024)
Mixture-of-Mamba: Enhancing Multi-Modal State-Space Models with Modality-Aware Sparsity
von: Liang, Weixin, et al.
Veröffentlicht: (2025)
von: Liang, Weixin, et al.
Veröffentlicht: (2025)
LMFusion: Adapting Pretrained Language Models for Multimodal Generation
von: Shi, Weijia, et al.
Veröffentlicht: (2024)
von: Shi, Weijia, et al.
Veröffentlicht: (2024)
Diffusion Autoencoders are Scalable Image Tokenizers
von: Chen, Yinbo, et al.
Veröffentlicht: (2025)
von: Chen, Yinbo, et al.
Veröffentlicht: (2025)
CAT Pruning: Cluster-Aware Token Pruning For Text-to-Image Diffusion Models
von: Cheng, Xinle, et al.
Veröffentlicht: (2025)
von: Cheng, Xinle, et al.
Veröffentlicht: (2025)
Effective pruning of web-scale datasets based on complexity of concept clusters
von: Abbas, Amro, et al.
Veröffentlicht: (2024)
von: Abbas, Amro, et al.
Veröffentlicht: (2024)
SelfEval: Leveraging the discriminative nature of generative models for evaluation
von: Rambhatla, Sai Saketh, et al.
Veröffentlicht: (2023)
von: Rambhatla, Sai Saketh, et al.
Veröffentlicht: (2023)
Image2Struct: Benchmarking Structure Extraction for Vision-Language Models
von: Roberts, Josselin Somerville, et al.
Veröffentlicht: (2024)
von: Roberts, Josselin Somerville, et al.
Veröffentlicht: (2024)
Pisces: An Auto-regressive Foundation Model for Image Understanding and Generation
von: Xu, Zhiyang, et al.
Veröffentlicht: (2025)
von: Xu, Zhiyang, et al.
Veröffentlicht: (2025)
ATD: Improved Transformer with Adaptive Token Dictionary for Image Restoration
von: Zhang, Leheng, et al.
Veröffentlicht: (2026)
von: Zhang, Leheng, et al.
Veröffentlicht: (2026)
Generating Multi-Image Synthetic Data for Text-to-Image Customization
von: Kumari, Nupur, et al.
Veröffentlicht: (2025)
von: Kumari, Nupur, et al.
Veröffentlicht: (2025)
CAT: Class Aware Adaptive Thresholding for Semi-Supervised Domain Generalization
von: Zoha, Sumaiya, et al.
Veröffentlicht: (2024)
von: Zoha, Sumaiya, et al.
Veröffentlicht: (2024)
CAT: Exploiting Inter-Class Dynamics for Domain Adaptive Object Detection
von: Kennerley, Mikhail, et al.
Veröffentlicht: (2024)
von: Kennerley, Mikhail, et al.
Veröffentlicht: (2024)
Multimodal RewardBench 2: Evaluating Omni Reward Models for Interleaved Text and Image
von: Hu, Yushi, et al.
Veröffentlicht: (2025)
von: Hu, Yushi, et al.
Veröffentlicht: (2025)
AdaNAT: Exploring Adaptive Policy for Token-Based Image Generation
von: Ni, Zanlin, et al.
Veröffentlicht: (2024)
von: Ni, Zanlin, et al.
Veröffentlicht: (2024)
Beyond Night Visibility: Adaptive Multi-Scale Fusion of Infrared and Visible Images
von: Pei, Shufan, et al.
Veröffentlicht: (2024)
von: Pei, Shufan, et al.
Veröffentlicht: (2024)
CAT-3DGS: A Context-Adaptive Triplane Approach to Rate-Distortion-Optimized 3DGS Compression
von: Zhan, Yu-Ting, et al.
Veröffentlicht: (2025)
von: Zhan, Yu-Ting, et al.
Veröffentlicht: (2025)
An Image is Worth 32 Tokens for Reconstruction and Generation
von: Yu, Qihang, et al.
Veröffentlicht: (2024)
von: Yu, Qihang, et al.
Veröffentlicht: (2024)
Multi-Expert Adaptive Selection: Task-Balancing for All-in-One Image Restoration
von: Yu, Xiaoyan, et al.
Veröffentlicht: (2024)
von: Yu, Xiaoyan, et al.
Veröffentlicht: (2024)
CE-FAM: Concept-Based Explanation via Fusion of Activation Maps
von: Kuroki, Michihiro, et al.
Veröffentlicht: (2025)
von: Kuroki, Michihiro, et al.
Veröffentlicht: (2025)
BSED: Baseline Shapley-Based Explainable Detector
von: Kuroki, Michihiro, et al.
Veröffentlicht: (2023)
von: Kuroki, Michihiro, et al.
Veröffentlicht: (2023)
CATANet: Efficient Content-Aware Token Aggregation for Lightweight Image Super-Resolution
von: Liu, Xin, et al.
Veröffentlicht: (2025)
von: Liu, Xin, et al.
Veröffentlicht: (2025)
Generating Illustrated Instructions
von: Menon, Sachit, et al.
Veröffentlicht: (2023)
von: Menon, Sachit, et al.
Veröffentlicht: (2023)
AITTI: Learning Adaptive Inclusive Token for Text-to-Image Generation
von: Hou, Xinyu, et al.
Veröffentlicht: (2024)
von: Hou, Xinyu, et al.
Veröffentlicht: (2024)
InstanceDiffusion: Instance-level Control for Image Generation
von: Wang, Xudong, et al.
Veröffentlicht: (2024)
von: Wang, Xudong, et al.
Veröffentlicht: (2024)
Efficient Hybrid CNN-GNN Architecture for Monocular Depth Estimation
von: Narayan, Ishan
Veröffentlicht: (2026)
von: Narayan, Ishan
Veröffentlicht: (2026)
GenEval 2: Addressing Benchmark Drift in Text-to-Image Evaluation
von: Kamath, Amita, et al.
Veröffentlicht: (2025)
von: Kamath, Amita, et al.
Veröffentlicht: (2025)
Fit Pixels, Get Labels: Meta-learned Implicit Networks for Image Segmentation
von: Vyas, Kushal, et al.
Veröffentlicht: (2025)
von: Vyas, Kushal, et al.
Veröffentlicht: (2025)
Negative Token Merging: Image-based Adversarial Feature Guidance
von: Singh, Jaskirat, et al.
Veröffentlicht: (2024)
von: Singh, Jaskirat, et al.
Veröffentlicht: (2024)
ADCD-Net: Robust Document Image Forgery Localization via Adaptive DCT Feature and Hierarchical Content Disentanglement
von: Wong, Kahim, et al.
Veröffentlicht: (2025)
von: Wong, Kahim, et al.
Veröffentlicht: (2025)
CADC: Content Adaptive Diffusion-Based Generative Image Compression
von: Sheng, Xihua, et al.
Veröffentlicht: (2026)
von: Sheng, Xihua, et al.
Veröffentlicht: (2026)
Adaptive Diffusion Terrain Generator for Autonomous Uneven Terrain Navigation
von: Yu, Youwei, et al.
Veröffentlicht: (2024)
von: Yu, Youwei, et al.
Veröffentlicht: (2024)
CAT: Contrastive Adapter Training for Personalized Image Generation
von: Park, Jae Wan, et al.
Veröffentlicht: (2024)
von: Park, Jae Wan, et al.
Veröffentlicht: (2024)
Content-Adaptive Image Retouching Guided by Attribute-Based Text Representation
von: Zhu, Hancheng, et al.
Veröffentlicht: (2025)
von: Zhu, Hancheng, et al.
Veröffentlicht: (2025)
Instant GaussianImage: A Generalizable and Self-Adaptive Image Representation via 2D Gaussian Splatting
von: Zeng, Zhaojie, et al.
Veröffentlicht: (2025)
von: Zeng, Zhaojie, et al.
Veröffentlicht: (2025)
User-in-the-Loop View Sampling with Error Peaking Visualization
von: Yasunaga, Ayaka, et al.
Veröffentlicht: (2025)
von: Yasunaga, Ayaka, et al.
Veröffentlicht: (2025)
OmniTokenizer: A Joint Image-Video Tokenizer for Visual Generation
von: Wang, Junke, et al.
Veröffentlicht: (2024)
von: Wang, Junke, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Transfusion: Predict the Next Token and Diffuse Images with One Multi-Modal Model
von: Zhou, Chunting, et al.
Veröffentlicht: (2024) -
Multimodal RewardBench: Holistic Evaluation of Reward Models for Vision Language Models
von: Yasunaga, Michihiro, et al.
Veröffentlicht: (2025) -
When Worse is Better: Navigating the compression-generation tradeoff in visual tokenization
von: Ramanujan, Vivek, et al.
Veröffentlicht: (2024) -
Computational Tradeoffs in Image Synthesis: Diffusion, Masked-Token, and Next-Token Prediction
von: Kilian, Maciej, et al.
Veröffentlicht: (2024) -
Mixture-of-Mamba: Enhancing Multi-Modal State-Space Models with Modality-Aware Sparsity
von: Liang, Weixin, et al.
Veröffentlicht: (2025)