Composable Visual Tokenizers with Generator-Free Diagnostics of Learnability
Fuente:
arXiv
Salvato in:
| Autori principali: | Zhao, Bingchen, Guo, Qiushan, Wang, Ye, Huang, Yixuan, Zhai, Zhonghua, Tian, Yu |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Free-VSC: Free Semantics from Visual Foundation Models for Unsupervised Video Semantic Compression
di: Tian, Yuan, et al.
Pubblicazione: (2024)
di: Tian, Yuan, et al.
Pubblicazione: (2024)
Efficient Visual Transformer by Learnable Token Merging
di: Wang, Yancheng, et al.
Pubblicazione: (2024)
di: Wang, Yancheng, et al.
Pubblicazione: (2024)
VTok: A Unified Video Tokenizer with Decoupled Spatial-Temporal Latents
di: Wang, Feng, et al.
Pubblicazione: (2026)
di: Wang, Feng, et al.
Pubblicazione: (2026)
Revisiting Multi-Task Visual Representation Learning
di: Di, Shangzhe, et al.
Pubblicazione: (2026)
di: Di, Shangzhe, et al.
Pubblicazione: (2026)
CROP: Contextual Region-Oriented Visual Token Pruning
di: Guo, Jiawei, et al.
Pubblicazione: (2025)
di: Guo, Jiawei, et al.
Pubblicazione: (2025)
DanceGRPO: Unleashing GRPO on Visual Generation
di: Xue, Zeyue, et al.
Pubblicazione: (2025)
di: Xue, Zeyue, et al.
Pubblicazione: (2025)
TokenCompose: Text-to-Image Diffusion with Token-level Supervision
di: Wang, Zirui, et al.
Pubblicazione: (2023)
di: Wang, Zirui, et al.
Pubblicazione: (2023)
Generating a Paracosm for Training-Free Zero-Shot Composed Image Retrieval
di: Wang, Tong, et al.
Pubblicazione: (2026)
di: Wang, Tong, et al.
Pubblicazione: (2026)
End-to-End Autoregressive Image Generation with 1D Semantic Tokenizer
di: Chu, Wenda, et al.
Pubblicazione: (2026)
di: Chu, Wenda, et al.
Pubblicazione: (2026)
Generalized Category Discovery under the Long-Tailed Distribution
di: Zhao, Bingchen, et al.
Pubblicazione: (2025)
di: Zhao, Bingchen, et al.
Pubblicazione: (2025)
Video Generation with Predictive Latents
di: Zhao, Yian, et al.
Pubblicazione: (2026)
di: Zhao, Yian, et al.
Pubblicazione: (2026)
TokenGS: Decoupling 3D Gaussian Prediction from Pixels with Learnable Tokens
di: Ren, Jiawei, et al.
Pubblicazione: (2026)
di: Ren, Jiawei, et al.
Pubblicazione: (2026)
VisionSelector: End-to-End Learnable Visual Token Compression for Efficient Multimodal LLMs
di: Zhu, Jiaying, et al.
Pubblicazione: (2025)
di: Zhu, Jiaying, et al.
Pubblicazione: (2025)
LGU-SLAM: Learnable Gaussian Uncertainty Matching with Deformable Correlation Sampling for Deep Visual SLAM
di: Huang, Yucheng, et al.
Pubblicazione: (2024)
di: Huang, Yucheng, et al.
Pubblicazione: (2024)
WinTok: A Win-Win Hybrid Tokenizer via Decomposing Visual Understanding and Generation with Transferable Tokens
di: Guo, Yiwei, et al.
Pubblicazione: (2026)
di: Guo, Yiwei, et al.
Pubblicazione: (2026)
From Mapping to Composing: A Two-Stage Framework for Zero-shot Composed Image Retrieval
di: Wang, Yabing, et al.
Pubblicazione: (2025)
di: Wang, Yabing, et al.
Pubblicazione: (2025)
OmniTokenizer: A Joint Image-Video Tokenizer for Visual Generation
di: Wang, Junke, et al.
Pubblicazione: (2024)
di: Wang, Junke, et al.
Pubblicazione: (2024)
FreeCompose: Generic Zero-Shot Image Composition with Diffusion Prior
di: Chen, Zhekai, et al.
Pubblicazione: (2024)
di: Chen, Zhekai, et al.
Pubblicazione: (2024)
RAPHAEL: Text-to-Image Generation via Large Mixture of Diffusion Paths
di: Xue, Zeyue, et al.
Pubblicazione: (2023)
di: Xue, Zeyue, et al.
Pubblicazione: (2023)
TARA: Token-Aware LoRA for Composable Personalization in Diffusion Models
di: Peng, Yuqi, et al.
Pubblicazione: (2025)
di: Peng, Yuqi, et al.
Pubblicazione: (2025)
Controllable Image Generation with Composed Parallel Token Prediction
di: Stirling, Jamie, et al.
Pubblicazione: (2024)
di: Stirling, Jamie, et al.
Pubblicazione: (2024)
Dynamic-VLM: Simple Dynamic Visual Token Compression for VideoLLM
di: Wang, Han, et al.
Pubblicazione: (2024)
di: Wang, Han, et al.
Pubblicazione: (2024)
Universal Video Temporal Grounding with Generative Multi-modal Large Language Models
di: Li, Zeqian, et al.
Pubblicazione: (2025)
di: Li, Zeqian, et al.
Pubblicazione: (2025)
Factorized Visual Tokenization and Generation
di: Bai, Zechen, et al.
Pubblicazione: (2024)
di: Bai, Zechen, et al.
Pubblicazione: (2024)
ClawMachine: Learning to Fetch Visual Tokens for Referential Comprehension
di: Ma, Tianren, et al.
Pubblicazione: (2024)
di: Ma, Tianren, et al.
Pubblicazione: (2024)
TokenMotion: Motion-Guided Vision Transformer for Video Camouflaged Object Detection Via Learnable Token Selection
di: Yu, Zifan, et al.
Pubblicazione: (2023)
di: Yu, Zifan, et al.
Pubblicazione: (2023)
WeTok: Powerful Discrete Tokenization for High-Fidelity Visual Reconstruction
di: Zhuang, Shaobin, et al.
Pubblicazione: (2025)
di: Zhuang, Shaobin, et al.
Pubblicazione: (2025)
Bridge the Gap Between Visual and Linguistic Comprehension for Generalized Zero-shot Semantic Segmentation
di: Guo, Xiaoqing, et al.
Pubblicazione: (2025)
di: Guo, Xiaoqing, et al.
Pubblicazione: (2025)
Sparse-Up: Learnable Sparse Upsampling for 3D Generation with High-Fidelity Textures
di: Xiao, Lu, et al.
Pubblicazione: (2025)
di: Xiao, Lu, et al.
Pubblicazione: (2025)
MMCORE: MultiModal COnnection with Representation Aligned Latent Embeddings
di: Li, Zijie, et al.
Pubblicazione: (2026)
di: Li, Zijie, et al.
Pubblicazione: (2026)
FashionComposer: Compositional Fashion Image Generation
di: Ji, Sihui, et al.
Pubblicazione: (2024)
di: Ji, Sihui, et al.
Pubblicazione: (2024)
QLIP: Text-Aligned Visual Tokenization Unifies Auto-Regressive Multimodal Understanding and Generation
di: Zhao, Yue, et al.
Pubblicazione: (2025)
di: Zhao, Yue, et al.
Pubblicazione: (2025)
Feature Aligning Few shot Learning Method Using Local Descriptors Weighted Rules
di: Yan, Bingchen
Pubblicazione: (2024)
di: Yan, Bingchen
Pubblicazione: (2024)
Beyond Visual Memory: Mechanistic Diagnostics of Latent Visual Reasoning
di: Guo, Garvin, et al.
Pubblicazione: (2026)
di: Guo, Garvin, et al.
Pubblicazione: (2026)
NativeTok: Native Visual Tokenization for Improved Image Generation
di: Wu, Bin, et al.
Pubblicazione: (2026)
di: Wu, Bin, et al.
Pubblicazione: (2026)
Interpretable Text-Guided Image Clustering via Iterative Search
di: Zhao, Bingchen, et al.
Pubblicazione: (2025)
di: Zhao, Bingchen, et al.
Pubblicazione: (2025)
TMCIR: Token Merge Benefits Composed Image Retrieval
di: Wang, Chaoyang, et al.
Pubblicazione: (2025)
di: Wang, Chaoyang, et al.
Pubblicazione: (2025)
LDCA: Local Descriptors with Contextual Augmentation for Few-Shot Learning
di: Wang, Maofa, et al.
Pubblicazione: (2024)
di: Wang, Maofa, et al.
Pubblicazione: (2024)
Contextuality Helps Representation Learning for Generalized Category Discovery
di: Luo, Tingzhang, et al.
Pubblicazione: (2024)
di: Luo, Tingzhang, et al.
Pubblicazione: (2024)
On the Learnability of Out-of-distribution Detection
di: Fang, Zhen, et al.
Pubblicazione: (2024)
di: Fang, Zhen, et al.
Pubblicazione: (2024)
Documenti analoghi
-
Free-VSC: Free Semantics from Visual Foundation Models for Unsupervised Video Semantic Compression
di: Tian, Yuan, et al.
Pubblicazione: (2024) -
Efficient Visual Transformer by Learnable Token Merging
di: Wang, Yancheng, et al.
Pubblicazione: (2024) -
VTok: A Unified Video Tokenizer with Decoupled Spatial-Temporal Latents
di: Wang, Feng, et al.
Pubblicazione: (2026) -
Revisiting Multi-Task Visual Representation Learning
di: Di, Shangzhe, et al.
Pubblicazione: (2026) -
CROP: Contextual Region-Oriented Visual Token Pruning
di: Guo, Jiawei, et al.
Pubblicazione: (2025)