Subobject-level Image Tokenization
Fuente:
arXiv
Salvato in:
| Autori principali: | Chen, Delong, Cahyawijaya, Samuel, Liu, Jianfeng, Wang, Baoyuan, Fung, Pascale |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Negative Object Presence Evaluation (NOPE) to Measure Object Hallucination in Vision-Language Models
di: Lovenia, Holy, et al.
Pubblicazione: (2023)
di: Lovenia, Holy, et al.
Pubblicazione: (2023)
What Makes for Good Image Captions?
di: Chen, Delong, et al.
Pubblicazione: (2024)
di: Chen, Delong, et al.
Pubblicazione: (2024)
WorldPrediction: A Benchmark for High-level World Modeling and Long-horizon Procedural Planning
di: Chen, Delong, et al.
Pubblicazione: (2025)
di: Chen, Delong, et al.
Pubblicazione: (2025)
Towards Efficient and Robust VQA-NLE Data Generation with Large Vision-Language Models
di: Irawan, Patrick Amadeus, et al.
Pubblicazione: (2024)
di: Irawan, Patrick Amadeus, et al.
Pubblicazione: (2024)
Chain-of-Talkers (CoTalk): Fast Human Annotation of Dense Image Captions
di: Shen, Yijun, et al.
Pubblicazione: (2025)
di: Shen, Yijun, et al.
Pubblicazione: (2025)
Text-guided Image Restoration and Semantic Enhancement for Text-to-Image Person Retrieval
di: Liu, Delong, et al.
Pubblicazione: (2023)
di: Liu, Delong, et al.
Pubblicazione: (2023)
LLMs Are Few-Shot In-Context Low-Resource Language Learners
di: Cahyawijaya, Samuel, et al.
Pubblicazione: (2024)
di: Cahyawijaya, Samuel, et al.
Pubblicazione: (2024)
CultureCLIP: Empowering CLIP with Cultural Awareness through Synthetic Images and Contextualized Captions
di: Huang, Yuchen, et al.
Pubblicazione: (2025)
di: Huang, Yuchen, et al.
Pubblicazione: (2025)
LongCat-Next: Lexicalizing Modalities as Discrete Tokens
di: Meituan LongCat Team, et al.
Pubblicazione: (2026)
di: Meituan LongCat Team, et al.
Pubblicazione: (2026)
StrokeNUWA: Tokenizing Strokes for Vector Graphic Synthesis
di: Tang, Zecheng, et al.
Pubblicazione: (2024)
di: Tang, Zecheng, et al.
Pubblicazione: (2024)
UTPTrack: Towards Simple and Unified Token Pruning for Visual Tracking
di: Wu, Hao, et al.
Pubblicazione: (2026)
di: Wu, Hao, et al.
Pubblicazione: (2026)
Multi-Modal Multi-Granularity Tokenizer for Chu Bamboo Slip Scripts
di: Chen, Yingfa, et al.
Pubblicazione: (2024)
di: Chen, Yingfa, et al.
Pubblicazione: (2024)
DualToken: Towards Unifying Visual Understanding and Generation with Dual Visual Vocabularies
di: Song, Wei, et al.
Pubblicazione: (2025)
di: Song, Wei, et al.
Pubblicazione: (2025)
See the Text: From Tokenization to Visual Reading
di: Xing, Ling, et al.
Pubblicazione: (2025)
di: Xing, Ling, et al.
Pubblicazione: (2025)
Dynamic Adapter with Semantics Disentangling for Cross-lingual Cross-modal Retrieval
di: Cai, Rui, et al.
Pubblicazione: (2024)
di: Cai, Rui, et al.
Pubblicazione: (2024)
LLM Internal States Reveal Hallucination Risk Faced With a Query
di: Ji, Ziwei, et al.
Pubblicazione: (2024)
di: Ji, Ziwei, et al.
Pubblicazione: (2024)
Action100M: A Large-scale Video Action Dataset
di: Chen, Delong, et al.
Pubblicazione: (2026)
di: Chen, Delong, et al.
Pubblicazione: (2026)
From Pixels to Tokens: Revisiting Object Hallucinations in Large Vision-Language Models
di: Shang, Yuying, et al.
Pubblicazione: (2024)
di: Shang, Yuying, et al.
Pubblicazione: (2024)
Prioritizing Image-Related Tokens Enhances Vision-Language Pre-Training
di: Chen, Yangyi, et al.
Pubblicazione: (2025)
di: Chen, Yangyi, et al.
Pubblicazione: (2025)
Vision-centric Token Compression in Large Language Model
di: Xing, Ling, et al.
Pubblicazione: (2025)
di: Xing, Ling, et al.
Pubblicazione: (2025)
Why 1 + 1 < 1 in Visual Token Pruning: Beyond Naive Integration via Multi-Objective Balanced Covering
di: Li, Yangfu, et al.
Pubblicazione: (2025)
di: Li, Yangfu, et al.
Pubblicazione: (2025)
Dynamic Token Reweighting for Robust Vision-Language Models
di: Jiang, Tanqiu, et al.
Pubblicazione: (2025)
di: Jiang, Tanqiu, et al.
Pubblicazione: (2025)
Efficient Whole Slide Pathology VQA via Token Compression
di: Lyu, Weimin, et al.
Pubblicazione: (2025)
di: Lyu, Weimin, et al.
Pubblicazione: (2025)
InfoMerge: Information-aware Token Compression for Efficient Video Large Language Models
di: Liu, Xinxin, et al.
Pubblicazione: (2026)
di: Liu, Xinxin, et al.
Pubblicazione: (2026)
ShortV: Efficient Multimodal Large Language Models by Freezing Visual Tokens in Ineffective Layers
di: Yuan, Qianhao, et al.
Pubblicazione: (2025)
di: Yuan, Qianhao, et al.
Pubblicazione: (2025)
Video-LaVIT: Unified Video-Language Pre-training with Decoupled Visual-Motional Tokenization
di: Jin, Yang, et al.
Pubblicazione: (2024)
di: Jin, Yang, et al.
Pubblicazione: (2024)
DynTok: Dynamic Compression of Visual Tokens for Efficient and Effective Video Understanding
di: Zhang, Hongzhi, et al.
Pubblicazione: (2025)
di: Zhang, Hongzhi, et al.
Pubblicazione: (2025)
High-Dimensional Interlingual Representations of Large Language Models
di: Wilie, Bryan, et al.
Pubblicazione: (2025)
di: Wilie, Bryan, et al.
Pubblicazione: (2025)
Llama Guard 3 Vision: Safeguarding Human-AI Image Understanding Conversations
di: Chi, Jianfeng, et al.
Pubblicazione: (2024)
di: Chi, Jianfeng, et al.
Pubblicazione: (2024)
CrossGET: Cross-Guided Ensemble of Tokens for Accelerating Vision-Language Transformers
di: Shi, Dachuan, et al.
Pubblicazione: (2023)
di: Shi, Dachuan, et al.
Pubblicazione: (2023)
Stop Looking for Important Tokens in Multimodal Language Models: Duplication Matters More
di: Wen, Zichen, et al.
Pubblicazione: (2025)
di: Wen, Zichen, et al.
Pubblicazione: (2025)
MetaToken: Detecting Hallucination in Image Descriptions by Meta Classification
di: Fieback, Laura, et al.
Pubblicazione: (2024)
di: Fieback, Laura, et al.
Pubblicazione: (2024)
Efficient Document Parsing via Parallel Token Prediction
di: Li, Lei, et al.
Pubblicazione: (2026)
di: Li, Lei, et al.
Pubblicazione: (2026)
An Image is Worth 1/2 Tokens After Layer 2: Plug-and-Play Inference Acceleration for Large Vision-Language Models
di: Chen, Liang, et al.
Pubblicazione: (2024)
di: Chen, Liang, et al.
Pubblicazione: (2024)
Open-Source Image Editing Models Are Zero-Shot Vision Learners
di: Liu, Wei, et al.
Pubblicazione: (2026)
di: Liu, Wei, et al.
Pubblicazione: (2026)
TokenCompose: Text-to-Image Diffusion with Token-level Supervision
di: Wang, Zirui, et al.
Pubblicazione: (2023)
di: Wang, Zirui, et al.
Pubblicazione: (2023)
Hiding Faces in Plain Sight: Defending DeepFakes by Disrupting Face Detection
di: Zhu, Delong, et al.
Pubblicazione: (2024)
di: Zhu, Delong, et al.
Pubblicazione: (2024)
Beyond Filtering: Adaptive Image-Text Quality Enhancement for MLLM Pretraining
di: Huang, Han, et al.
Pubblicazione: (2024)
di: Huang, Han, et al.
Pubblicazione: (2024)
Interleaved Scene Graphs for Interleaved Text-and-Image Generation Assessment
di: Chen, Dongping, et al.
Pubblicazione: (2024)
di: Chen, Dongping, et al.
Pubblicazione: (2024)
ReGATE: Learning Faster and Better with Fewer Tokens in MLLMs
di: Li, Chaoyu, et al.
Pubblicazione: (2025)
di: Li, Chaoyu, et al.
Pubblicazione: (2025)
Documenti analoghi
-
Negative Object Presence Evaluation (NOPE) to Measure Object Hallucination in Vision-Language Models
di: Lovenia, Holy, et al.
Pubblicazione: (2023) -
What Makes for Good Image Captions?
di: Chen, Delong, et al.
Pubblicazione: (2024) -
WorldPrediction: A Benchmark for High-level World Modeling and Long-horizon Procedural Planning
di: Chen, Delong, et al.
Pubblicazione: (2025) -
Towards Efficient and Robust VQA-NLE Data Generation with Large Vision-Language Models
di: Irawan, Patrick Amadeus, et al.
Pubblicazione: (2024) -
Chain-of-Talkers (CoTalk): Fast Human Annotation of Dense Image Captions
di: Shen, Yijun, et al.
Pubblicazione: (2025)