Descriminative-Generative Custom Tokens for Vision-Language Models
Fuente:
arXiv
Saved in:
| Main Authors: | Perera, Pramuditha, Trager, Matthew, Zancato, Luca, Achille, Alessandro, Soatto, Stefano |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Linear Spaces of Meanings: Compositional Structures in Vision-Language Models
by: Trager, Matthew, et al.
Published: (2023)
by: Trager, Matthew, et al.
Published: (2023)
Multi-Modal Hallucination Control by Visual Information Grounding
by: Favero, Alessandro, et al.
Published: (2024)
by: Favero, Alessandro, et al.
Published: (2024)
Compositional Structures in Neural Embedding and Interaction Decompositions
by: Trager, Matthew, et al.
Published: (2024)
by: Trager, Matthew, et al.
Published: (2024)
NeRF-Insert: 3D Local Editing with Multimodal Control Signals
by: Sabat, Benet Oriol, et al.
Published: (2024)
by: Sabat, Benet Oriol, et al.
Published: (2024)
CPR: Retrieval Augmented Generation for Copyright Protection
by: Golatkar, Aditya, et al.
Published: (2024)
by: Golatkar, Aditya, et al.
Published: (2024)
Interpretable Measures of Conceptual Similarity by Complexity-Constrained Descriptive Auto-Encoding
by: Achille, Alessandro, et al.
Published: (2024)
by: Achille, Alessandro, et al.
Published: (2024)
Training Data Protection with Compositional Diffusion Models
by: Golatkar, Aditya, et al.
Published: (2023)
by: Golatkar, Aditya, et al.
Published: (2023)
Detecting Korean Food Using Image using Hierarchical Model
by: Lam, Hoang Khanh, et al.
Published: (2024)
by: Lam, Hoang Khanh, et al.
Published: (2024)
PICASO: Permutation-Invariant Context Composition with State Space Models
by: Liu, Tian Yu, et al.
Published: (2025)
by: Liu, Tian Yu, et al.
Published: (2025)
Non-autoregressive Sequence-to-Sequence Vision-Language Models
by: Shi, Kunyu, et al.
Published: (2024)
by: Shi, Kunyu, et al.
Published: (2024)
Beyond Sole Strength: Customized Ensembles for Generalized Vision-Language Models
by: Lu, Zhihe, et al.
Published: (2023)
by: Lu, Zhihe, et al.
Published: (2023)
THRONE: An Object-based Hallucination Benchmark for the Free-form Generations of Large Vision-Language Models
by: Kaul, Prannay, et al.
Published: (2024)
by: Kaul, Prannay, et al.
Published: (2024)
Enhancing Vision-Language Pre-training with Rich Supervisions
by: Gao, Yuan, et al.
Published: (2024)
by: Gao, Yuan, et al.
Published: (2024)
Diffeomorphic Template Registration for Atmospheric Turbulence Mitigation
by: Lao, Dong, et al.
Published: (2024)
by: Lao, Dong, et al.
Published: (2024)
Diffusion Soup: Model Merging for Text-to-Image Diffusion Models
by: Biggs, Benjamin, et al.
Published: (2024)
by: Biggs, Benjamin, et al.
Published: (2024)
Attention Debiasing for Token Pruning in Vision Language Models
by: Zhao, Kai, et al.
Published: (2025)
by: Zhao, Kai, et al.
Published: (2025)
Enhancing Vision-Language Model with Unmasked Token Alignment
by: Liu, Jihao, et al.
Published: (2024)
by: Liu, Jihao, et al.
Published: (2024)
HALP: Detecting Hallucinations in Vision-Language Models without Generating a Single Token
by: Kogilathota, Sai Akhil, et al.
Published: (2026)
by: Kogilathota, Sai Akhil, et al.
Published: (2026)
B'MOJO: Hybrid State Space Realizations of Foundation Models with Eidetic and Fading Memory
by: Zancato, Luca, et al.
Published: (2024)
by: Zancato, Luca, et al.
Published: (2024)
Vision Foundation Models as Generalist Tokenizers for Image Generation
by: Zheng, Anlin, et al.
Published: (2026)
by: Zheng, Anlin, et al.
Published: (2026)
Variation-aware Vision Token Dropping for Faster Large Vision-Language Models
by: Chen, Junjie, et al.
Published: (2025)
by: Chen, Junjie, et al.
Published: (2025)
Dynamic Token Reduction during Generation for Vision Language Models
by: Liang, Xiaoyu, et al.
Published: (2025)
by: Liang, Xiaoyu, et al.
Published: (2025)
Unified Pix Token And Word Token Generative Language Model
by: Leung, Haun, et al.
Published: (2026)
by: Leung, Haun, et al.
Published: (2026)
UniCompress: Token Compression for Unified Vision-Language Understanding and Generation
by: Wang, Ziyao, et al.
Published: (2026)
by: Wang, Ziyao, et al.
Published: (2026)
Do Vision Language Models Need to Process Image Tokens?
by: Ghosh, Sambit, et al.
Published: (2026)
by: Ghosh, Sambit, et al.
Published: (2026)
Towards Joint Quantization and Token Pruning of Vision-Language Models
by: Li, Xinqing, et al.
Published: (2026)
by: Li, Xinqing, et al.
Published: (2026)
Language-Guided Token Compression with Reinforcement Learning in Large Vision-Language Models
by: Cao, Sihan, et al.
Published: (2026)
by: Cao, Sihan, et al.
Published: (2026)
Musketeer: Joint Training for Multi-task Vision Language Model with Task Explanation Prompts
by: Zhang, Zhaoyang, et al.
Published: (2023)
by: Zhang, Zhaoyang, et al.
Published: (2023)
Dynamic Token Reweighting for Robust Vision-Language Models
by: Jiang, Tanqiu, et al.
Published: (2025)
by: Jiang, Tanqiu, et al.
Published: (2025)
Vision-centric Token Compression in Large Language Model
by: Xing, Ling, et al.
Published: (2025)
by: Xing, Ling, et al.
Published: (2025)
Scenes as Tokens: Multi-Scale Normal Distributions Transform Tokenizer for General 3D Vision-Language Understanding
by: Tang, Yutao, et al.
Published: (2025)
by: Tang, Yutao, et al.
Published: (2025)
Leveraging Visual Signals for Robust Token-Level Uncertainty in Vision-Language Generation
by: Hoche, Joseph, et al.
Published: (2026)
by: Hoche, Joseph, et al.
Published: (2026)
HalLoc: Token-level Localization of Hallucinations for Vision Language Models
by: Park, Eunkyu, et al.
Published: (2025)
by: Park, Eunkyu, et al.
Published: (2025)
Skip-Vision: Efficient and Scalable Acceleration of Vision-Language Models via Adaptive Token Skipping
by: Zeng, Weili, et al.
Published: (2025)
by: Zeng, Weili, et al.
Published: (2025)
HybridToken-VLM: Hybrid Token Compression for Vision-Language Models
by: Zhang, Jusheng, et al.
Published: (2025)
by: Zhang, Jusheng, et al.
Published: (2025)
TokenCom: Vision-Language Model for Multimodal and Multitask Token Communications
by: Jiang, Feibo, et al.
Published: (2026)
by: Jiang, Feibo, et al.
Published: (2026)
Object-Centric Vision Token Pruning for Vision Language Models
by: Li, Guangyuan, et al.
Published: (2025)
by: Li, Guangyuan, et al.
Published: (2025)
CustomKD: Customizing Large Vision Foundation for Edge Model Improvement via Knowledge Distillation
by: Lee, Jungsoo, et al.
Published: (2025)
by: Lee, Jungsoo, et al.
Published: (2025)
Perceptio: Perception Enhanced Vision Language Models via Spatial Token Generation
by: Li, Yuchen, et al.
Published: (2026)
by: Li, Yuchen, et al.
Published: (2026)
TokenSwap: Backdoor Attack on the Compositional Understanding of Large Vision-Language Models
by: Zhang, Zhifang, et al.
Published: (2025)
by: Zhang, Zhifang, et al.
Published: (2025)
Similar Items
-
Linear Spaces of Meanings: Compositional Structures in Vision-Language Models
by: Trager, Matthew, et al.
Published: (2023) -
Multi-Modal Hallucination Control by Visual Information Grounding
by: Favero, Alessandro, et al.
Published: (2024) -
Compositional Structures in Neural Embedding and Interaction Decompositions
by: Trager, Matthew, et al.
Published: (2024) -
NeRF-Insert: 3D Local Editing with Multimodal Control Signals
by: Sabat, Benet Oriol, et al.
Published: (2024) -
CPR: Retrieval Augmented Generation for Copyright Protection
by: Golatkar, Aditya, et al.
Published: (2024)