Token Bottleneck: One Token to Remember Dynamics
Fuente:
arXiv
Saved in:
| Main Authors: | Kim, Taekyung, Han, Dongyoon, Heo, Byeongho, Park, Jeongeun, Yun, Sangdoo |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Morphing Tokens Draw Strong Masked Image Models
by: Kim, Taekyung, et al.
Published: (2023)
by: Kim, Taekyung, et al.
Published: (2023)
Masking meets Supervision: A Strong Learning Alliance
by: Heo, Byeongho, et al.
Published: (2023)
by: Heo, Byeongho, et al.
Published: (2023)
Learning with Unmasked Tokens Drives Stronger Vision Learners
by: Kim, Taekyung, et al.
Published: (2023)
by: Kim, Taekyung, et al.
Published: (2023)
Rotary Position Embedding for Vision Transformer
by: Heo, Byeongho, et al.
Published: (2024)
by: Heo, Byeongho, et al.
Published: (2024)
Match me if you can: Semi-Supervised Semantic Correspondence Learning with Unpaired Images
by: Kim, Jiwon, et al.
Published: (2023)
by: Kim, Jiwon, et al.
Published: (2023)
SeiT++: Masked Token Modeling Improves Storage-efficient Training
by: Lee, Minhyun, et al.
Published: (2023)
by: Lee, Minhyun, et al.
Published: (2023)
Exploring Conditions for Diffusion models in Robotic Control
by: Shin, Heeseong, et al.
Published: (2025)
by: Shin, Heeseong, et al.
Published: (2025)
RL makes MLLMs see better than SFT
by: Song, Junha, et al.
Published: (2025)
by: Song, Junha, et al.
Published: (2025)
Learning to See What You Need: Gaze Attention for Multimodal Large Language Models
by: Song, Junha, et al.
Published: (2026)
by: Song, Junha, et al.
Published: (2026)
HYPE: Hyperbolic Entailment Filtering for Underspecified Images and Texts
by: Kim, Wonjae, et al.
Published: (2024)
by: Kim, Wonjae, et al.
Published: (2024)
DenseNets Reloaded: Paradigm Shift Beyond ResNets and ViTs
by: Kim, Donghyun, et al.
Published: (2024)
by: Kim, Donghyun, et al.
Published: (2024)
DNNs May Determine Major Properties of Their Outputs Early, with Timing Possibly Driven by Bias
by: Park, Song, et al.
Published: (2025)
by: Park, Song, et al.
Published: (2025)
Oops, Wait: Token-Level Signals as a Lens into LLM Reasoning
by: Hwang, Jaehui, et al.
Published: (2026)
by: Hwang, Jaehui, et al.
Published: (2026)
Aligned Novel View Image and Geometry Synthesis via Cross-modal Attention Instillation
by: Kwak, Min-Seop, et al.
Published: (2025)
by: Kwak, Min-Seop, et al.
Published: (2025)
MaskRIS: Semantic Distortion-aware Data Augmentation for Referring Image Segmentation
by: Lee, Minhyun, et al.
Published: (2024)
by: Lee, Minhyun, et al.
Published: (2024)
Leveraging Temporal Contextualization for Video Action Recognition
by: Kim, Minji, et al.
Published: (2024)
by: Kim, Minji, et al.
Published: (2024)
Model Stock: All we need is just a few fine-tuned models
by: Jang, Dong-Hwan, et al.
Published: (2024)
by: Jang, Dong-Hwan, et al.
Published: (2024)
DaWin: Training-free Dynamic Weight Interpolation for Robust Adaptation
by: Oh, Changdae, et al.
Published: (2024)
by: Oh, Changdae, et al.
Published: (2024)
Similarity of Neural Architectures using Adversarial Attack Transferability
by: Hwang, Jaehui, et al.
Published: (2022)
by: Hwang, Jaehui, et al.
Published: (2022)
Emergence of Text Readability in Vision Language Models
by: Park, Jaeyoo, et al.
Published: (2025)
by: Park, Jaeyoo, et al.
Published: (2025)
Scratching Visual Transformer's Back with Uniform Attention
by: Hyeon-Woo, Nam, et al.
Published: (2022)
by: Hyeon-Woo, Nam, et al.
Published: (2022)
Personalized Reward Modeling for Text-to-Image Generation
by: Lee, Jeongeun, et al.
Published: (2025)
by: Lee, Jeongeun, et al.
Published: (2025)
Lipsum-FT: Robust Fine-Tuning of Zero-Shot Models Using Random Text Guidance
by: Nam, Giung, et al.
Published: (2024)
by: Nam, Giung, et al.
Published: (2024)
Probabilistic Language-Image Pre-Training
by: Chun, Sanghyuk, et al.
Published: (2024)
by: Chun, Sanghyuk, et al.
Published: (2024)
Map the Flow: Revealing Hidden Pathways of Information in VideoLLMs
by: Kim, Minji, et al.
Published: (2025)
by: Kim, Minji, et al.
Published: (2025)
HalLoc: Token-level Localization of Hallucinations for Vision Language Models
by: Park, Eunkyu, et al.
Published: (2025)
by: Park, Eunkyu, et al.
Published: (2025)
V2C-CBM: Building Concept Bottlenecks with Vision-to-Concept Tokenizer
by: He, Hangzhou, et al.
Published: (2025)
by: He, Hangzhou, et al.
Published: (2025)
A Frame is Worth One Token: Efficient Generative World Modeling with Delta Tokens
by: Kerssies, Tommie, et al.
Published: (2026)
by: Kerssies, Tommie, et al.
Published: (2026)
OneChart: Purify the Chart Structural Extraction via One Auxiliary Token
by: Chen, Jinyue, et al.
Published: (2024)
by: Chen, Jinyue, et al.
Published: (2024)
Multi-Granular Spatio-Temporal Token Merging for Training-Free Acceleration of Video LLMs
by: Hyun, Jeongseok, et al.
Published: (2025)
by: Hyun, Jeongseok, et al.
Published: (2025)
Towards Efficient Vision State Space Models via Token Merging
by: Park, Jinyoung, et al.
Published: (2025)
by: Park, Jinyoung, et al.
Published: (2025)
Can Natural Image Autoencoders Compactly Tokenize fMRI Volumes for Long-Range Dynamics Modeling?
by: Kim, Peter Yongho, et al.
Published: (2026)
by: Kim, Peter Yongho, et al.
Published: (2026)
Which Concepts to Forget and How to Refuse? Decomposing Concepts for Continual Unlearning in Large Vision-Language Models
by: Jin, Hyundong, et al.
Published: (2026)
by: Jin, Hyundong, et al.
Published: (2026)
Superpixel Tokenization for Vision Transformers: Preserving Semantic Integrity in Visual Tokens
by: Lew, Jaihyun, et al.
Published: (2024)
by: Lew, Jaihyun, et al.
Published: (2024)
LongProLIP: A Probabilistic Vision-Language Model with Long Context Text
by: Chun, Sanghyuk, et al.
Published: (2025)
by: Chun, Sanghyuk, et al.
Published: (2025)
Towards Calibrated Robust Fine-Tuning of Vision-Language Models
by: Oh, Changdae, et al.
Published: (2023)
by: Oh, Changdae, et al.
Published: (2023)
Multi-criteria Token Fusion with One-step-ahead Attention for Efficient Vision Transformers
by: Lee, Sanghyeok, et al.
Published: (2024)
by: Lee, Sanghyeok, et al.
Published: (2024)
Semantic One-Dimensional Tokenizer for Image Reconstruction and Generation
by: Qu, Yunpeng, et al.
Published: (2026)
by: Qu, Yunpeng, et al.
Published: (2026)
One Token, Two Fates: A Unified Framework via Vision Token Manipulation Against MLLMs Hallucination
by: Fa, Zhan, et al.
Published: (2026)
by: Fa, Zhan, et al.
Published: (2026)
ERASE: Eliminating Redundant Visual Tokens via Adaptive Two-Stage Token Pruning
by: Lee, Yuna, et al.
Published: (2026)
by: Lee, Yuna, et al.
Published: (2026)
Similar Items
-
Morphing Tokens Draw Strong Masked Image Models
by: Kim, Taekyung, et al.
Published: (2023) -
Masking meets Supervision: A Strong Learning Alliance
by: Heo, Byeongho, et al.
Published: (2023) -
Learning with Unmasked Tokens Drives Stronger Vision Learners
by: Kim, Taekyung, et al.
Published: (2023) -
Rotary Position Embedding for Vision Transformer
by: Heo, Byeongho, et al.
Published: (2024) -
Match me if you can: Semi-Supervised Semantic Correspondence Learning with Unpaired Images
by: Kim, Jiwon, et al.
Published: (2023)