Multi-Token Prediction Needs Registers
Fuente:
arXiv
Saved in:
| Main Authors: | Gerontopoulos, Anastasios, Gidaris, Spyros, Komodakis, Nikos |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
MOCA: Self-supervised Representation Learning by Predicting Masked Online Codebook Assignments
by: Gidaris, Spyros, et al.
Published: (2023)
by: Gidaris, Spyros, et al.
Published: (2023)
Representations Before Pixels: Semantics-Guided Hierarchical Video Prediction
by: Karypidis, Efstathios, et al.
Published: (2026)
by: Karypidis, Efstathios, et al.
Published: (2026)
Coevolving Representations in Joint Image-Feature Diffusion
by: Kouzelis, Theodoros, et al.
Published: (2026)
by: Kouzelis, Theodoros, et al.
Published: (2026)
Advancing Semantic Future Prediction through Multimodal Visual Sequence Transformers
by: Karypidis, Efstathios, et al.
Published: (2025)
by: Karypidis, Efstathios, et al.
Published: (2025)
SPOT: Self-Training with Patch-Order Permutation for Object-Centric Learning with Autoregressive Transformers
by: Kakogeorgiou, Ioannis, et al.
Published: (2023)
by: Kakogeorgiou, Ioannis, et al.
Published: (2023)
DINO-Foresight: Looking into the Future with DINO
by: Karypidis, Efstathios, et al.
Published: (2024)
by: Karypidis, Efstathios, et al.
Published: (2024)
No Train, all Gain: Self-Supervised Gradients Improve Deep Frozen Representations
by: Simoncini, Walter, et al.
Published: (2024)
by: Simoncini, Walter, et al.
Published: (2024)
Boosting Generative Image Modeling via Joint Image-Feature Synthesis
by: Kouzelis, Theodoros, et al.
Published: (2025)
by: Kouzelis, Theodoros, et al.
Published: (2025)
xGen-MM-Vid (BLIP-3-Video): You Only Need 32 Tokens to Represent a Video Even in VLMs
by: Ryoo, Michael S., et al.
Published: (2024)
by: Ryoo, Michael S., et al.
Published: (2024)
Analyzing The Language of Visual Tokens
by: Chan, David M., et al.
Published: (2024)
by: Chan, David M., et al.
Published: (2024)
A General and Efficient Training for Transformer via Token Expansion
by: Huang, Wenxuan, et al.
Published: (2024)
by: Huang, Wenxuan, et al.
Published: (2024)
A High-Level Survey of Optical Remote Sensing
by: Koletsis, Panagiotis, et al.
Published: (2026)
by: Koletsis, Panagiotis, et al.
Published: (2026)
Squeeze Out Tokens from Sample for Finer-Grained Data Governance
by: Lin, Weixiong, et al.
Published: (2025)
by: Lin, Weixiong, et al.
Published: (2025)
Groma: Localized Visual Tokenization for Grounding Multimodal Large Language Models
by: Ma, Chuofan, et al.
Published: (2024)
by: Ma, Chuofan, et al.
Published: (2024)
Efficient Vision-Language Models by Summarizing Visual Tokens into Compact Registers
by: Wen, Yuxin, et al.
Published: (2024)
by: Wen, Yuxin, et al.
Published: (2024)
Is Less More? Exploring Token Condensation as Training-free Test-time Adaptation
by: Wang, Zixin, et al.
Published: (2024)
by: Wang, Zixin, et al.
Published: (2024)
LagKV: Lag-Relative Information of the KV Cache Tells Which Tokens Are Important
by: Liang, Manlai, et al.
Published: (2025)
by: Liang, Manlai, et al.
Published: (2025)
Less Is More, but Where? Dynamic Token Compression via LLM-Guided Keyframe Prior
by: Li, Yulin, et al.
Published: (2025)
by: Li, Yulin, et al.
Published: (2025)
T2I-R1: Reinforcing Image Generation with Collaborative Semantic-level and Token-level CoT
by: Jiang, Dongzhi, et al.
Published: (2025)
by: Jiang, Dongzhi, et al.
Published: (2025)
FlashVID: Efficient Video Large Language Models via Training-free Tree-based Spatiotemporal Token Merging
by: Fan, Ziyang, et al.
Published: (2026)
by: Fan, Ziyang, et al.
Published: (2026)
Driving on Registers
by: Kirby, Ellington, et al.
Published: (2026)
by: Kirby, Ellington, et al.
Published: (2026)
Inference Optimal VLMs Need Fewer Visual Tokens and More Parameters
by: Li, Kevin Y., et al.
Published: (2024)
by: Li, Kevin Y., et al.
Published: (2024)
MuMA-ToM: Multi-modal Multi-Agent Theory of Mind
by: Shi, Haojun, et al.
Published: (2024)
by: Shi, Haojun, et al.
Published: (2024)
Moto: Latent Motion Token as the Bridging Language for Learning Robot Manipulation from Videos
by: Chen, Yi, et al.
Published: (2024)
by: Chen, Yi, et al.
Published: (2024)
GeoToken: Hierarchical Geolocalization of Images via Next Token Prediction
by: Ghasemi, Narges, et al.
Published: (2025)
by: Ghasemi, Narges, et al.
Published: (2025)
Thought Flow Nets: From Single Predictions to Trains of Model Thought
by: Schuff, Hendrik, et al.
Published: (2021)
by: Schuff, Hendrik, et al.
Published: (2021)
Do We Really Need Curated Malicious Data for Safety Alignment in Multi-modal Large Language Models?
by: Wang, Yanbo, et al.
Published: (2025)
by: Wang, Yanbo, et al.
Published: (2025)
MultiDelete for Multimodal Machine Unlearning
by: Cheng, Jiali, et al.
Published: (2023)
by: Cheng, Jiali, et al.
Published: (2023)
PQPP: A Joint Benchmark for Text-to-Image Prompt and Query Performance Prediction
by: Poesina, Eduard, et al.
Published: (2024)
by: Poesina, Eduard, et al.
Published: (2024)
MobileCLIP2: Improving Multi-Modal Reinforced Training
by: Faghri, Fartash, et al.
Published: (2025)
by: Faghri, Fartash, et al.
Published: (2025)
The Labyrinth of Links: Navigating the Associative Maze of Multi-modal LLMs
by: Li, Hong, et al.
Published: (2024)
by: Li, Hong, et al.
Published: (2024)
MM-LIMA: Less Is More for Alignment in Multi-Modal Datasets
by: Wei, Lai, et al.
Published: (2023)
by: Wei, Lai, et al.
Published: (2023)
Explaining Black-box Model Predictions via Two-level Nested Feature Attributions with Consistency Property
by: Yoshikawa, Yuya, et al.
Published: (2024)
by: Yoshikawa, Yuya, et al.
Published: (2024)
GET: Unlocking the Multi-modal Potential of CLIP for Generalized Category Discovery
by: Wang, Enguang, et al.
Published: (2024)
by: Wang, Enguang, et al.
Published: (2024)
MMedAgent-RL: Optimizing Multi-Agent Collaboration for Multimodal Medical Reasoning
by: Xia, Peng, et al.
Published: (2025)
by: Xia, Peng, et al.
Published: (2025)
Weighted Multi-Prompt Learning with Description-free Large Language Model Distillation
by: Lee, Sua, et al.
Published: (2025)
by: Lee, Sua, et al.
Published: (2025)
Mimicking or Reasoning: Rethinking Multi-Modal In-Context Learning in Vision-Language Models
by: Huang, Chengyue, et al.
Published: (2025)
by: Huang, Chengyue, et al.
Published: (2025)
SignMusketeers: An Efficient Multi-Stream Approach for Sign Language Translation at Scale
by: Gueuwou, Shester, et al.
Published: (2024)
by: Gueuwou, Shester, et al.
Published: (2024)
MMCTAgent: Multi-modal Critical Thinking Agent Framework for Complex Visual Reasoning
by: Kumar, Somnath, et al.
Published: (2024)
by: Kumar, Somnath, et al.
Published: (2024)
Mixture-of-Mamba: Enhancing Multi-Modal State-Space Models with Modality-Aware Sparsity
by: Liang, Weixin, et al.
Published: (2025)
by: Liang, Weixin, et al.
Published: (2025)
Similar Items
-
MOCA: Self-supervised Representation Learning by Predicting Masked Online Codebook Assignments
by: Gidaris, Spyros, et al.
Published: (2023) -
Representations Before Pixels: Semantics-Guided Hierarchical Video Prediction
by: Karypidis, Efstathios, et al.
Published: (2026) -
Coevolving Representations in Joint Image-Feature Diffusion
by: Kouzelis, Theodoros, et al.
Published: (2026) -
Advancing Semantic Future Prediction through Multimodal Visual Sequence Transformers
by: Karypidis, Efstathios, et al.
Published: (2025) -
SPOT: Self-Training with Patch-Order Permutation for Object-Centric Learning with Autoregressive Transformers
by: Kakogeorgiou, Ioannis, et al.
Published: (2023)