Superpixel Tokenization for Vision Transformers: Preserving Semantic Integrity in Visual Tokens
Fuente:
arXiv
Salvato in:
| Autori principali: | Lew, Jaihyun, Jang, Soohyuk, Lee, Jaehoon, Yoo, Seungryong, Kim, Eunji, Lee, Saehyung, Mok, Jisoo, Kim, Siwon, Yoon, Sungroh |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Balancing Saliency and Coverage: Semantic Prominence-Aware Budgeting for Visual Token Compression in VLMs
di: Lee, Jaehoon, et al.
Pubblicazione: (2026)
di: Lee, Jaehoon, et al.
Pubblicazione: (2026)
SAVE: Sparse Autoencoder-Driven Visual Information Enhancement for Mitigating Object Hallucination
di: Park, Sangha, et al.
Pubblicazione: (2025)
di: Park, Sangha, et al.
Pubblicazione: (2025)
Visual Attention Never Fades: Selective Progressive Attention ReCalibration for Detailed Image Captioning in Multimodal Large Language Models
di: Jung, Mingi, et al.
Pubblicazione: (2025)
di: Jung, Mingi, et al.
Pubblicazione: (2025)
HeSS: Head Sensitivity Score for Sparsity Redistribution in VGGT
di: Kim, Yongsung, et al.
Pubblicazione: (2026)
di: Kim, Yongsung, et al.
Pubblicazione: (2026)
Textual Training for the Hassle-Free Removal of Unwanted Visual Data: Case Studies on OOD and Hateful Image Detection
di: Lee, Saehyung, et al.
Pubblicazione: (2024)
di: Lee, Saehyung, et al.
Pubblicazione: (2024)
Causality-Aware Contrastive Learning for Robust Multivariate Time-Series Anomaly Detection
di: Kim, HyunGi, et al.
Pubblicazione: (2025)
di: Kim, HyunGi, et al.
Pubblicazione: (2025)
Semantic Token Reweighting for Interpretable and Controllable Text Embeddings in CLIP
di: Kim, Eunji, et al.
Pubblicazione: (2024)
di: Kim, Eunji, et al.
Pubblicazione: (2024)
Unsupervised Homography Estimation on Multimodal Image Pair via Alternating Optimization
di: Song, Sanghyeob, et al.
Pubblicazione: (2024)
di: Song, Sanghyeob, et al.
Pubblicazione: (2024)
Diagnosing and Correcting Concept Omission in Multimodal Diffusion Transformers
di: Baek, Kanghyun, et al.
Pubblicazione: (2026)
di: Baek, Kanghyun, et al.
Pubblicazione: (2026)
Battling the Non-stationarity in Time Series Forecasting via Test-time Adaptation
di: Kim, HyunGi, et al.
Pubblicazione: (2025)
di: Kim, HyunGi, et al.
Pubblicazione: (2025)
Disentangled Motion Modeling for Video Frame Interpolation
di: Lew, Jaihyun, et al.
Pubblicazione: (2024)
di: Lew, Jaihyun, et al.
Pubblicazione: (2024)
On mitigating stability-plasticity dilemma in CLIP-guided image morphing via geodesic distillation loss
di: Oh, Yeongtak, et al.
Pubblicazione: (2024)
di: Oh, Yeongtak, et al.
Pubblicazione: (2024)
Contextualized Visual Personalization in Vision-Language Models
di: Oh, Yeongtak, et al.
Pubblicazione: (2026)
di: Oh, Yeongtak, et al.
Pubblicazione: (2026)
Toward Robust Hyper-Detailed Image Captioning: A Multiagent Approach and Dual Evaluation Metrics for Factuality and Coverage
di: Lee, Saehyung, et al.
Pubblicazione: (2024)
di: Lee, Saehyung, et al.
Pubblicazione: (2024)
Rethinking Training for De-biasing Text-to-Image Generation: Unlocking the Potential of Stable Diffusion
di: Kim, Eunji, et al.
Pubblicazione: (2024)
di: Kim, Eunji, et al.
Pubblicazione: (2024)
Interactive Text-to-Image Retrieval with Large Language Models: A Plug-and-Play Approach
di: Lee, Saehyung, et al.
Pubblicazione: (2024)
di: Lee, Saehyung, et al.
Pubblicazione: (2024)
A Spitting Image: Modular Superpixel Tokenization in Vision Transformers
di: Aasan, Marius, et al.
Pubblicazione: (2024)
di: Aasan, Marius, et al.
Pubblicazione: (2024)
DefectFill: Realistic Defect Generation with Inpainting Diffusion Model for Visual Inspection
di: Song, Jaewoo, et al.
Pubblicazione: (2025)
di: Song, Jaewoo, et al.
Pubblicazione: (2025)
CANDI: Curated Test-Time Adaptation for Multivariate Time-Series Anomaly Detection Under Distribution Shift
di: Kim, HyunGi, et al.
Pubblicazione: (2026)
di: Kim, HyunGi, et al.
Pubblicazione: (2026)
ERASE: Eliminating Redundant Visual Tokens via Adaptive Two-Stage Token Pruning
di: Lee, Yuna, et al.
Pubblicazione: (2026)
di: Lee, Yuna, et al.
Pubblicazione: (2026)
Entropy is not Enough for Test-Time Adaptation: From the Perspective of Disentangled Factors
di: Lee, Jonghyun, et al.
Pubblicazione: (2024)
di: Lee, Jonghyun, et al.
Pubblicazione: (2024)
Exploring the Potential of LLMs as Personalized Assistants: Dataset, Evaluation, and Analysis
di: Mok, Jisoo, et al.
Pubblicazione: (2025)
di: Mok, Jisoo, et al.
Pubblicazione: (2025)
Frequency-Aware Token Reduction for Efficient Vision Transformer
di: Lee, Dong-Jae, et al.
Pubblicazione: (2025)
di: Lee, Dong-Jae, et al.
Pubblicazione: (2025)
Multi-criteria Token Fusion with One-step-ahead Attention for Efficient Vision Transformers
di: Lee, Sanghyeok, et al.
Pubblicazione: (2024)
di: Lee, Sanghyeok, et al.
Pubblicazione: (2024)
Talk3D: High-Fidelity Talking Portrait Synthesis via Personalized 3D Generative Prior
di: Ko, Jaehoon, et al.
Pubblicazione: (2024)
di: Ko, Jaehoon, et al.
Pubblicazione: (2024)
Unifying Feature and Cost Aggregation with Transformers for Semantic and Visual Correspondence
di: Hong, Sunghwan, et al.
Pubblicazione: (2024)
di: Hong, Sunghwan, et al.
Pubblicazione: (2024)
STAG: Structural Test-time Alignment of Gradients for Online Adaptation
di: Shin, Juhyeon, et al.
Pubblicazione: (2024)
di: Shin, Juhyeon, et al.
Pubblicazione: (2024)
Guiding What Not to Generate: Automated Negative Prompting for Text-Image Alignment
di: Park, Sangha, et al.
Pubblicazione: (2025)
di: Park, Sangha, et al.
Pubblicazione: (2025)
DAFA: Distance-Aware Fair Adversarial Training
di: Lee, Hyungyu, et al.
Pubblicazione: (2024)
di: Lee, Hyungyu, et al.
Pubblicazione: (2024)
Visual-Word Tokenizer: Beyond Fixed Sets of Tokens in Vision Transformers
di: Gee, Leonidas, et al.
Pubblicazione: (2024)
di: Gee, Leonidas, et al.
Pubblicazione: (2024)
RePIC: Reinforced Post-Training for Personalizing Multi-Modal Language Models
di: Oh, Yeongtak, et al.
Pubblicazione: (2025)
di: Oh, Yeongtak, et al.
Pubblicazione: (2025)
Revisit What You See: Revealing Visual Semantics in Vision Tokens to Guide LVLM Decoding
di: Cho, Beomsik, et al.
Pubblicazione: (2025)
di: Cho, Beomsik, et al.
Pubblicazione: (2025)
AT-SNN: Adaptive Tokens for Vision Transformer on Spiking Neural Network
di: Kang, Donghwa, et al.
Pubblicazione: (2024)
di: Kang, Donghwa, et al.
Pubblicazione: (2024)
GaussianTalker: Real-Time High-Fidelity Talking Head Synthesis with Audio-Driven 3D Gaussian Splatting
di: Cho, Kyusun, et al.
Pubblicazione: (2024)
di: Cho, Kyusun, et al.
Pubblicazione: (2024)
Self-Supervised Time-Series Anomaly Detection Using Learnable Data Augmentation
di: Choi, Kukjin, et al.
Pubblicazione: (2024)
di: Choi, Kukjin, et al.
Pubblicazione: (2024)
Verbal-R3: Verbal Reranker as the Missing Bridge between Retrieval and Reasoning
di: Park, Sangkwon, et al.
Pubblicazione: (2026)
di: Park, Sangkwon, et al.
Pubblicazione: (2026)
MVFormer: Diversifying Feature Normalization and Token Mixing for Efficient Vision Transformers
di: Bae, Jongseong, et al.
Pubblicazione: (2024)
di: Bae, Jongseong, et al.
Pubblicazione: (2024)
TabFlash: Efficient Table Understanding with Progressive Question Conditioning and Token Focusing
di: Kim, Jongha, et al.
Pubblicazione: (2025)
di: Kim, Jongha, et al.
Pubblicazione: (2025)
Why and When Visual Token Pruning Fails? A Study on Relevant Visual Information Shift in MLLMs Decoding
di: Kim, Jiwan, et al.
Pubblicazione: (2026)
di: Kim, Jiwan, et al.
Pubblicazione: (2026)
Structured State-Space Regularization for Generation-Friendly Image Tokenization
di: Lee, Jinsung, et al.
Pubblicazione: (2026)
di: Lee, Jinsung, et al.
Pubblicazione: (2026)
Documenti analoghi
-
Balancing Saliency and Coverage: Semantic Prominence-Aware Budgeting for Visual Token Compression in VLMs
di: Lee, Jaehoon, et al.
Pubblicazione: (2026) -
SAVE: Sparse Autoencoder-Driven Visual Information Enhancement for Mitigating Object Hallucination
di: Park, Sangha, et al.
Pubblicazione: (2025) -
Visual Attention Never Fades: Selective Progressive Attention ReCalibration for Detailed Image Captioning in Multimodal Large Language Models
di: Jung, Mingi, et al.
Pubblicazione: (2025) -
HeSS: Head Sensitivity Score for Sparsity Redistribution in VGGT
di: Kim, Yongsung, et al.
Pubblicazione: (2026) -
Textual Training for the Hassle-Free Removal of Unwanted Visual Data: Case Studies on OOD and Hateful Image Detection
di: Lee, Saehyung, et al.
Pubblicazione: (2024)