QLIP: Text-Aligned Visual Tokenization Unifies Auto-Regressive Multimodal Understanding and Generation
Fuente:
arXiv
Guardado en:
| Autores principales: | Zhao, Yue, Xue, Fuzhao, Reed, Scott, Fan, Linxi, Zhu, Yuke, Kautz, Jan, Yu, Zhiding, Krähenbühl, Philipp, Huang, De-An |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Image and Video Tokenization with Binary Spherical Quantization
por: Zhao, Yue, et al.
Publicado: (2024)
por: Zhao, Yue, et al.
Publicado: (2024)
FRAG: Frame Selection Augmented Generation for Long Video and Long Document Understanding
por: Huang, De-An, et al.
Publicado: (2025)
por: Huang, De-An, et al.
Publicado: (2025)
AMAGO: Scalable In-Context Reinforcement Learning for Adaptive Agents
por: Grigsby, Jake, et al.
Publicado: (2023)
por: Grigsby, Jake, et al.
Publicado: (2023)
Spherical Leech Quantization for Visual Tokenization and Generation
por: Zhao, Yue, et al.
Publicado: (2025)
por: Zhao, Yue, et al.
Publicado: (2025)
LongVILA: Scaling Long-Context Visual Language Models for Long Videos
por: Chen, Yukang, et al.
Publicado: (2024)
por: Chen, Yukang, et al.
Publicado: (2024)
STORM: Token-Efficient Long Video Understanding for Multimodal LLMs
por: Jiang, Jindong, et al.
Publicado: (2025)
por: Jiang, Jindong, et al.
Publicado: (2025)
Vision as a Dialect: Unifying Visual Understanding and Generation via Text-Aligned Representations
por: Han, Jiaming, et al.
Publicado: (2025)
por: Han, Jiaming, et al.
Publicado: (2025)
Unified Autoregressive Visual Generation and Understanding with Continuous Tokens
por: Fan, Lijie, et al.
Publicado: (2025)
por: Fan, Lijie, et al.
Publicado: (2025)
UM-Text: A Unified Multimodal Model for Image Understanding and Visual Text Editing
por: Ma, Lichen, et al.
Publicado: (2026)
por: Ma, Lichen, et al.
Publicado: (2026)
Sim-to-Real Reinforcement Learning for Vision-Based Dexterous Manipulation on Humanoids
por: Lin, Toru, et al.
Publicado: (2025)
por: Lin, Toru, et al.
Publicado: (2025)
InfinityStar: Unified Spacetime AutoRegressive Modeling for Visual Generation
por: Liu, Jinlai, et al.
Publicado: (2025)
por: Liu, Jinlai, et al.
Publicado: (2025)
STAR: STacked AutoRegressive Scheme for Unified Multimodal Learning
por: Qin, Jie, et al.
Publicado: (2025)
por: Qin, Jie, et al.
Publicado: (2025)
UniToken: Harmonizing Multimodal Understanding and Generation through Unified Visual Encoding
por: Jiao, Yang, et al.
Publicado: (2025)
por: Jiao, Yang, et al.
Publicado: (2025)
Compressed Map Priors for 3D Perception
por: Zhou, Brady, et al.
Publicado: (2025)
por: Zhou, Brady, et al.
Publicado: (2025)
TokenFlow: Unified Image Tokenizer for Multimodal Understanding and Generation
por: Qu, Liao, et al.
Publicado: (2024)
por: Qu, Liao, et al.
Publicado: (2024)
Auto-Encoding Morph-Tokens for Multimodal LLM
por: Pan, Kaihang, et al.
Publicado: (2024)
por: Pan, Kaihang, et al.
Publicado: (2024)
OneCAT: Decoder-Only Auto-Regressive Model for Unified Understanding and Generation
por: Li, Han, et al.
Publicado: (2025)
por: Li, Han, et al.
Publicado: (2025)
Interactive Post-Training for Vision-Language-Action Models
por: Tan, Shuhan, et al.
Publicado: (2025)
por: Tan, Shuhan, et al.
Publicado: (2025)
VAR-CLIP: Text-to-Image Generator with Visual Auto-Regressive Modeling
por: Zhang, Qian, et al.
Publicado: (2024)
por: Zhang, Qian, et al.
Publicado: (2024)
Prismer: A Vision-Language Model with Multi-Task Experts
por: Liu, Shikun, et al.
Publicado: (2023)
por: Liu, Shikun, et al.
Publicado: (2023)
Unified Multimodal Understanding via Byte-Pair Visual Encoding
por: Zhang, Wanpeng, et al.
Publicado: (2025)
por: Zhang, Wanpeng, et al.
Publicado: (2025)
QLIP: A Dynamic Quadtree Vision Prior Enhances MLLM Performance Without Retraining
por: Chickering, Kyle R., et al.
Publicado: (2025)
por: Chickering, Kyle R., et al.
Publicado: (2025)
HAVEN: Hierarchically Aligned Multimodal Benchmark for Unified Video Understanding
por: Shi, Mengqi, et al.
Publicado: (2026)
por: Shi, Mengqi, et al.
Publicado: (2026)
Steering Visual Generation in Unified Multimodal Models with Understanding Supervision
por: Liu, Zeyu, et al.
Publicado: (2026)
por: Liu, Zeyu, et al.
Publicado: (2026)
PhyCritic: Multimodal Critic Models for Physical AI
por: Xiong, Tianyi, et al.
Publicado: (2026)
por: Xiong, Tianyi, et al.
Publicado: (2026)
UniFlow: A Unified Pixel Flow Tokenizer for Visual Understanding and Generation
por: Yue, Zhengrong, et al.
Publicado: (2025)
por: Yue, Zhengrong, et al.
Publicado: (2025)
FlashVLM: Text-Guided Visual Token Selection for Large Multimodal Models
por: Cai, Kaitong, et al.
Publicado: (2025)
por: Cai, Kaitong, et al.
Publicado: (2025)
Wave-Particle (Continuous-Discrete) Dualistic Visual Tokenization for Unified Understanding and Generation
por: Chen, Yizhu, et al.
Publicado: (2025)
por: Chen, Yizhu, et al.
Publicado: (2025)
Auto-Regressive Next-Token Predictors are Universal Learners
por: Malach, Eran
Publicado: (2023)
por: Malach, Eran
Publicado: (2023)
GEOMETRIA ÓSSEA E ATIVIDADE FÍSICA EM CRIANÇAS E ADOLESCENTES: REVISÃO SISTEMÁTICA
por: Tathyane Krahenbühl
Publicado: (2018)
por: Tathyane Krahenbühl
Publicado: (2018)
The use of the additional field player in handball: analysis of the Rio 2016 Olympic Games
por: Tathyane Krahenbühl
Publicado: (2019)
por: Tathyane Krahenbühl
Publicado: (2019)
Fatores que influenciam a massa óssea de crianças e adolescentes saudáveis mensurada pelo ultrassom quantitativo de falanges: revisão sistemática
por: Tathyane Krahenbühl
Publicado: (2014)
por: Tathyane Krahenbühl
Publicado: (2014)
HOVER: Versatile Neural Whole-Body Controller for Humanoid Robots
por: He, Tairan, et al.
Publicado: (2024)
por: He, Tairan, et al.
Publicado: (2024)
UniTok: A Unified Tokenizer for Visual Generation and Understanding
por: Ma, Chuofan, et al.
Publicado: (2025)
por: Ma, Chuofan, et al.
Publicado: (2025)
Stateful Token Reduction for Long-Video Hybrid VLMs
por: Jiang, Jindong, et al.
Publicado: (2026)
por: Jiang, Jindong, et al.
Publicado: (2026)
ARDuP: Active Region Video Diffusion for Universal Policies
por: Huang, Shuaiyi, et al.
Publicado: (2024)
por: Huang, Shuaiyi, et al.
Publicado: (2024)
Prompt-Aware Adapter: Towards Learning Adaptive Visual Tokens for Multimodal Large Language Models
por: Zhang, Yue, et al.
Publicado: (2024)
por: Zhang, Yue, et al.
Publicado: (2024)
DualToken: Towards Unifying Visual Understanding and Generation with Dual Visual Vocabularies
por: Song, Wei, et al.
Publicado: (2025)
por: Song, Wei, et al.
Publicado: (2025)
Domain Adaptation Through Task Distillation
por: Zhou, Brady, et al.
Publicado: (2020)
por: Zhou, Brady, et al.
Publicado: (2020)
Long-term Traffic Simulation with Interleaved Autoregressive Motion and Scenario Generation
por: Yang, Xiuyu, et al.
Publicado: (2025)
por: Yang, Xiuyu, et al.
Publicado: (2025)
Ejemplares similares
-
Image and Video Tokenization with Binary Spherical Quantization
por: Zhao, Yue, et al.
Publicado: (2024) -
FRAG: Frame Selection Augmented Generation for Long Video and Long Document Understanding
por: Huang, De-An, et al.
Publicado: (2025) -
AMAGO: Scalable In-Context Reinforcement Learning for Adaptive Agents
por: Grigsby, Jake, et al.
Publicado: (2023) -
Spherical Leech Quantization for Visual Tokenization and Generation
por: Zhao, Yue, et al.
Publicado: (2025) -
LongVILA: Scaling Long-Context Visual Language Models for Long Videos
por: Chen, Yukang, et al.
Publicado: (2024)