Scalable Image Tokenization with Index Backpropagation Quantization
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Shi, Fengyuan, Luo, Zhuoyan, Ge, Yixiao, Yang, Yujiu, Shan, Ying, Wang, Limin |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Open-MAGVIT2: An Open-Source Project Toward Democratizing Auto-regressive Visual Generation
von: Luo, Zhuoyan, et al.
Veröffentlicht: (2024)
von: Luo, Zhuoyan, et al.
Veröffentlicht: (2024)
1st Place Solution for 5th LSVOS Challenge: Referring Video Object Segmentation
von: Luo, Zhuoyan, et al.
Veröffentlicht: (2024)
von: Luo, Zhuoyan, et al.
Veröffentlicht: (2024)
VEN-VL: A Visual Ensemble MoE Framework for Effective and Efficient Multi-Modal Understanding
von: Wu, Yinghao, et al.
Veröffentlicht: (2026)
von: Wu, Yinghao, et al.
Veröffentlicht: (2026)
TimeLens: Rethinking Video Temporal Grounding with Multimodal LLMs
von: Zhang, Jun, et al.
Veröffentlicht: (2025)
von: Zhang, Jun, et al.
Veröffentlicht: (2025)
TokLIP: Marry Visual Tokens to CLIP for Multimodal Comprehension and Generation
von: Lin, Haokun, et al.
Veröffentlicht: (2025)
von: Lin, Haokun, et al.
Veröffentlicht: (2025)
Supervised Fine-tuning in turn Improves Visual Foundation Models
von: Jiang, Xiaohu, et al.
Veröffentlicht: (2024)
von: Jiang, Xiaohu, et al.
Veröffentlicht: (2024)
BIVDiff: A Training-Free Framework for General-Purpose Video Synthesis via Bridging Image and Video Diffusion Models
von: Shi, Fengyuan, et al.
Veröffentlicht: (2023)
von: Shi, Fengyuan, et al.
Veröffentlicht: (2023)
Efficient Personalization of Quantized Diffusion Model without Backpropagation
von: Seo, Hoigi, et al.
Veröffentlicht: (2025)
von: Seo, Hoigi, et al.
Veröffentlicht: (2025)
CoHD: A Counting-Aware Hierarchical Decoding Framework for Generalized Referring Expression Segmentation
von: Luo, Zhuoyan, et al.
Veröffentlicht: (2024)
von: Luo, Zhuoyan, et al.
Veröffentlicht: (2024)
EgoPlan-Bench2: A Benchmark for Multimodal Large Language Model Planning in Real-World Scenarios
von: Qiu, Lu, et al.
Veröffentlicht: (2024)
von: Qiu, Lu, et al.
Veröffentlicht: (2024)
Scaling Image Tokenizers with Grouped Spherical Quantization
von: Wang, Jiangtao, et al.
Veröffentlicht: (2024)
von: Wang, Jiangtao, et al.
Veröffentlicht: (2024)
Moto: Latent Motion Token as the Bridging Language for Learning Robot Manipulation from Videos
von: Chen, Yi, et al.
Veröffentlicht: (2024)
von: Chen, Yi, et al.
Veröffentlicht: (2024)
MaskCD: Mitigating LVLM Hallucinations by Image Head Masked Contrastive Decoding
von: Deng, Jingyuan, et al.
Veröffentlicht: (2025)
von: Deng, Jingyuan, et al.
Veröffentlicht: (2025)
Advanced Sign Language Video Generation with Compressed and Quantized Multi-Condition Tokenization
von: Wang, Cong, et al.
Veröffentlicht: (2025)
von: Wang, Cong, et al.
Veröffentlicht: (2025)
VQ-Map: Bird's-Eye-View Map Layout Estimation in Tokenized Discrete Space via Vector Quantization
von: Zhang, Yiwei, et al.
Veröffentlicht: (2024)
von: Zhang, Yiwei, et al.
Veröffentlicht: (2024)
Diffusion Autoencoders are Scalable Image Tokenizers
von: Chen, Yinbo, et al.
Veröffentlicht: (2025)
von: Chen, Yinbo, et al.
Veröffentlicht: (2025)
Mitigating the Reasoning Tax in Vision-Language Fine-Tuning with Input-Adaptive Depth Aggregation
von: Ren, Yiming, et al.
Veröffentlicht: (2026)
von: Ren, Yiming, et al.
Veröffentlicht: (2026)
Divot: Diffusion Powers Video Tokenizer for Comprehension and Generation
von: Ge, Yuying, et al.
Veröffentlicht: (2024)
von: Ge, Yuying, et al.
Veröffentlicht: (2024)
UniRepLKNet: A Universal Perception Large-Kernel ConvNet for Audio, Video, Point Cloud, Time-Series and Image Recognition
von: Ding, Xiaohan, et al.
Veröffentlicht: (2023)
von: Ding, Xiaohan, et al.
Veröffentlicht: (2023)
DiCoDe: Diffusion-Compressed Deep Tokens for Autoregressive Video Generation with Language Models
von: Li, Yizhuo, et al.
Veröffentlicht: (2024)
von: Li, Yizhuo, et al.
Veröffentlicht: (2024)
HaploOmni: Unified Single Transformer for Multimodal Video Understanding and Generation
von: Xiao, Yicheng, et al.
Veröffentlicht: (2025)
von: Xiao, Yicheng, et al.
Veröffentlicht: (2025)
ViT-Lens: Towards Omni-modal Representations
von: Lei, Weixian, et al.
Veröffentlicht: (2023)
von: Lei, Weixian, et al.
Veröffentlicht: (2023)
StyleCrafter: Enhancing Stylized Text-to-Video Generation with Style Adapter
von: Liu, Gongye, et al.
Veröffentlicht: (2023)
von: Liu, Gongye, et al.
Veröffentlicht: (2023)
DiffMoE: Dynamic Token Selection for Scalable Diffusion Transformers
von: Shi, Minglei, et al.
Veröffentlicht: (2025)
von: Shi, Minglei, et al.
Veröffentlicht: (2025)
AdaTP: Attention-Debiased Token Pruning for Video Large Language Models
von: Sun, Fengyuan, et al.
Veröffentlicht: (2025)
von: Sun, Fengyuan, et al.
Veröffentlicht: (2025)
Multimodal Pathway: Improve Transformers with Irrelevant Data from Other Modalities
von: Zhang, Yiyuan, et al.
Veröffentlicht: (2024)
von: Zhang, Yiyuan, et al.
Veröffentlicht: (2024)
GRPO-CARE: Consistency-Aware Reinforcement Learning for Multimodal Reasoning
von: Chen, Yi, et al.
Veröffentlicht: (2025)
von: Chen, Yi, et al.
Veröffentlicht: (2025)
Exploring the Effect of Reinforcement Learning on Video Understanding: Insights from SEED-Bench-R1
von: Chen, Yi, et al.
Veröffentlicht: (2025)
von: Chen, Yi, et al.
Veröffentlicht: (2025)
QAPruner: Quantization-Aware Vision Token Pruning for Multimodal Large Language Models
von: Wang, Xinhao, et al.
Veröffentlicht: (2026)
von: Wang, Xinhao, et al.
Veröffentlicht: (2026)
FlowSteer: Guiding Few-Step Image Synthesis with Authentic Trajectories
von: Ke, Lei, et al.
Veröffentlicht: (2025)
von: Ke, Lei, et al.
Veröffentlicht: (2025)
Frequency Autoregressive Image Generation with Continuous Tokens
von: Yu, Hu, et al.
Veröffentlicht: (2025)
von: Yu, Hu, et al.
Veröffentlicht: (2025)
MergeVQ: A Unified Framework for Visual Generation and Representation with Disentangled Token Merging and Quantization
von: Li, Siyuan, et al.
Veröffentlicht: (2025)
von: Li, Siyuan, et al.
Veröffentlicht: (2025)
TokenFlow: Unified Image Tokenizer for Multimodal Understanding and Generation
von: Qu, Liao, et al.
Veröffentlicht: (2024)
von: Qu, Liao, et al.
Veröffentlicht: (2024)
GaussianToken: An Effective Image Tokenizer with 2D Gaussian Splatting
von: Dong, Jiajun, et al.
Veröffentlicht: (2025)
von: Dong, Jiajun, et al.
Veröffentlicht: (2025)
OccamToken: Efficient VLM Inference with Training-Free and Budget-Adaptive Token Pruning
von: Li, Geng, et al.
Veröffentlicht: (2026)
von: Li, Geng, et al.
Veröffentlicht: (2026)
VideoZoomer: Reinforcement-Learned Temporal Focusing for Long Video Reasoning
von: Ding, Yang, et al.
Veröffentlicht: (2025)
von: Ding, Yang, et al.
Veröffentlicht: (2025)
Hita: Holistic Tokenizer for Autoregressive Image Generation
von: Zheng, Anlin, et al.
Veröffentlicht: (2025)
von: Zheng, Anlin, et al.
Veröffentlicht: (2025)
Image Tokens Matter: Mitigating Hallucination in Discrete Tokenizer-based Large Vision-Language Models via Latent Editing
von: Wang, Weixing, et al.
Veröffentlicht: (2025)
von: Wang, Weixing, et al.
Veröffentlicht: (2025)
Aligning Text-to-Image Diffusion Models with Reward Backpropagation
von: Prabhudesai, Mihir, et al.
Veröffentlicht: (2023)
von: Prabhudesai, Mihir, et al.
Veröffentlicht: (2023)
Token Merging for Training-Free Semantic Binding in Text-to-Image Synthesis
von: Hu, Taihang, et al.
Veröffentlicht: (2024)
von: Hu, Taihang, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Open-MAGVIT2: An Open-Source Project Toward Democratizing Auto-regressive Visual Generation
von: Luo, Zhuoyan, et al.
Veröffentlicht: (2024) -
1st Place Solution for 5th LSVOS Challenge: Referring Video Object Segmentation
von: Luo, Zhuoyan, et al.
Veröffentlicht: (2024) -
VEN-VL: A Visual Ensemble MoE Framework for Effective and Efficient Multi-Modal Understanding
von: Wu, Yinghao, et al.
Veröffentlicht: (2026) -
TimeLens: Rethinking Video Temporal Grounding with Multimodal LLMs
von: Zhang, Jun, et al.
Veröffentlicht: (2025) -
TokLIP: Marry Visual Tokens to CLIP for Multimodal Comprehension and Generation
von: Lin, Haokun, et al.
Veröffentlicht: (2025)