Saved in:
| Main Authors: | Park, Minyoung, Kong, Taehun, Ahn, Sangjun |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2605.19322 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
AdaTok: Adaptive Token Compression with Object-Aware Representations for Efficient Multimodal LLMs
by: Zhang, Xinliang, et al.
Published: (2025)
by: Zhang, Xinliang, et al.
Published: (2025)
InfoTok: Adaptive Discrete Video Tokenizer via Information-Theoretic Compression
by: Ye, Haotian, et al.
Published: (2025)
by: Ye, Haotian, et al.
Published: (2025)
SweetTok: Semantic-Aware Spatial-Temporal Tokenizer for Compact Video Discretization
by: Tan, Zhentao, et al.
Published: (2024)
by: Tan, Zhentao, et al.
Published: (2024)
Learning Adaptive Pseudo-Label Selection for Semi-Supervised 3D Object Detection
by: Kong, Taehun, et al.
Published: (2025)
by: Kong, Taehun, et al.
Published: (2025)
DynTok: Dynamic Compression of Visual Tokens for Efficient and Effective Video Understanding
by: Zhang, Hongzhi, et al.
Published: (2025)
by: Zhang, Hongzhi, et al.
Published: (2025)
TrajTok: Learning Trajectory Tokens enables better Video Understanding
by: Zheng, Chenhao, et al.
Published: (2026)
by: Zheng, Chenhao, et al.
Published: (2026)
RefTok: Reference-Based Tokenization for Video Generation
by: Fan, Xiang, et al.
Published: (2025)
by: Fan, Xiang, et al.
Published: (2025)
Cardiac Segmentation on CT Images through Shape-Aware Contour Attentions
by: Park, Sanguk, et al.
Published: (2021)
by: Park, Sanguk, et al.
Published: (2021)
VideoFlexTok: Flexible-Length Coarse-to-Fine Video Tokenization
by: Atanov, Andrei, et al.
Published: (2026)
by: Atanov, Andrei, et al.
Published: (2026)
Unified Spatiotemporal Token Compression for Video-LLMs at Ultra-Low Retention
by: Du, Junhao, et al.
Published: (2026)
by: Du, Junhao, et al.
Published: (2026)
Semi-Supervised 3D Object Detection with Channel Augmentation using Transformation Equivariance
by: Kang, Minju, et al.
Published: (2024)
by: Kang, Minju, et al.
Published: (2024)
DINO-Tok: Adapting DINO for Visual Tokenizers
by: Jia, Mingkai, et al.
Published: (2025)
by: Jia, Mingkai, et al.
Published: (2025)
Extreme Model Compression for Edge Vision-Language Models: Sparse Temporal Token Fusion and Adaptive Neural Compression
by: Tanvir, Md Tasnin, et al.
Published: (2025)
by: Tanvir, Md Tasnin, et al.
Published: (2025)
Enhancing Multi-Image Understanding through Delimiter Token Scaling
by: Lee, Minyoung, et al.
Published: (2026)
by: Lee, Minyoung, et al.
Published: (2026)
PyraTok: Language-Aligned Pyramidal Tokenizer for Video Understanding and Generation
by: Susladkar, Onkar, et al.
Published: (2026)
by: Susladkar, Onkar, et al.
Published: (2026)
Adaptive-VoCo: Complexity-Aware Visual Token Compression for Vision-Language Models
by: Guo, Xiaoyang, et al.
Published: (2025)
by: Guo, Xiaoyang, et al.
Published: (2025)
MergeTok: Unified Continuous and Discrete Visual Tokenization via Token Merging
by: Zhang, Luyuan, et al.
Published: (2026)
by: Zhang, Luyuan, et al.
Published: (2026)
VidTok: A Versatile and Open-Source Video Tokenizer
by: Tang, Anni, et al.
Published: (2024)
by: Tang, Anni, et al.
Published: (2024)
STORM: Token-Efficient Long Video Understanding for Multimodal LLMs
by: Jiang, Jindong, et al.
Published: (2025)
by: Jiang, Jindong, et al.
Published: (2025)
SceneTok: A Compressed, Diffusable Token Space for 3D Scenes
by: Asim, Mohammad, et al.
Published: (2026)
by: Asim, Mohammad, et al.
Published: (2026)
TAG: A Simple Yet Effective Temporal-Aware Approach for Zero-Shot Video Temporal Grounding
by: Lee, Jin-Seop, et al.
Published: (2025)
by: Lee, Jin-Seop, et al.
Published: (2025)
Learning Adaptive and Temporally Causal Video Tokenization in a 1D Latent Space
by: Li, Yan, et al.
Published: (2025)
by: Li, Yan, et al.
Published: (2025)
NativeTok: Native Visual Tokenization for Improved Image Generation
by: Wu, Bin, et al.
Published: (2026)
by: Wu, Bin, et al.
Published: (2026)
GloTok: Global Perspective Tokenizer for Image Reconstruction and Generation
by: Zhao, Xuan, et al.
Published: (2025)
by: Zhao, Xuan, et al.
Published: (2025)
FlowTok: Flowing Seamlessly Across Text and Image Tokens
by: He, Ju, et al.
Published: (2025)
by: He, Ju, et al.
Published: (2025)
See and Fix the Flaws: Enabling VLMs and Diffusion Models to Comprehend Visual Artifacts via Agentic Data Synthesis
by: Park, Jaehyun, et al.
Published: (2026)
by: Park, Jaehyun, et al.
Published: (2026)
Token Transforming: A Unified and Training-Free Token Compression Framework for Vision Transformer Acceleration
by: Zeng, Fanhu, et al.
Published: (2025)
by: Zeng, Fanhu, et al.
Published: (2025)
Mitigating Hallucination in VideoLLMs via Temporal-Aware Activation Engineering
by: Cai, Jianfeng, et al.
Published: (2025)
by: Cai, Jianfeng, et al.
Published: (2025)
Spatial Degradation-Aware and Temporal Consistent Diffusion Model for Compressed Video Super-Resolution
by: An, Hongyu, et al.
Published: (2025)
by: An, Hongyu, et al.
Published: (2025)
OTT-Vid: Optimal Transport Temporal Token Compression for Video Large Language Models
by: Kang, Minseok, et al.
Published: (2026)
by: Kang, Minseok, et al.
Published: (2026)
DTTNet: Improving Video Shadow Detection via Dark-Aware Guidance and Tokenized Temporal Modeling
by: Li, Zhicheng, et al.
Published: (2025)
by: Li, Zhicheng, et al.
Published: (2025)
WeTok: Powerful Discrete Tokenization for High-Fidelity Visual Reconstruction
by: Zhuang, Shaobin, et al.
Published: (2025)
by: Zhuang, Shaobin, et al.
Published: (2025)
AlignTok: Aligning Visual Foundation Encoders to Tokenizers for Diffusion Models
by: Chen, Bowei, et al.
Published: (2025)
by: Chen, Bowei, et al.
Published: (2025)
V-LynX: Token Interface Alignment for Video+X LLMs
by: Park, Jungin, et al.
Published: (2026)
by: Park, Jungin, et al.
Published: (2026)
TemporalVLM: Video LLMs for Temporal Reasoning in Long Videos
by: Fateh, Fawad Javed, et al.
Published: (2024)
by: Fateh, Fawad Javed, et al.
Published: (2024)
MacTok: Robust Continuous Tokenization for Image Generation
by: Zeng, Hengyu, et al.
Published: (2026)
by: Zeng, Hengyu, et al.
Published: (2026)
Video-XL-Pro: Reconstructive Token Compression for Extremely Long Video Understanding
by: Liu, Xiangrui, et al.
Published: (2025)
by: Liu, Xiangrui, et al.
Published: (2025)
Video Compression with Hierarchical Temporal Neural Representation
by: Zhu, Jun, et al.
Published: (2026)
by: Zhu, Jun, et al.
Published: (2026)
CaTok: Taming Mean Flows for One-Dimensional Causal Image Tokenization
by: Chen, Yitong, et al.
Published: (2026)
by: Chen, Yitong, et al.
Published: (2026)
WinTok: A Win-Win Hybrid Tokenizer via Decomposing Visual Understanding and Generation with Transferable Tokens
by: Guo, Yiwei, et al.
Published: (2026)
by: Guo, Yiwei, et al.
Published: (2026)
Similar Items
-
AdaTok: Adaptive Token Compression with Object-Aware Representations for Efficient Multimodal LLMs
by: Zhang, Xinliang, et al.
Published: (2025) -
InfoTok: Adaptive Discrete Video Tokenizer via Information-Theoretic Compression
by: Ye, Haotian, et al.
Published: (2025) -
SweetTok: Semantic-Aware Spatial-Temporal Tokenizer for Compact Video Discretization
by: Tan, Zhentao, et al.
Published: (2024) -
Learning Adaptive Pseudo-Label Selection for Semi-Supervised 3D Object Detection
by: Kong, Taehun, et al.
Published: (2025) -
DynTok: Dynamic Compression of Visual Tokens for Efficient and Effective Video Understanding
by: Zhang, Hongzhi, et al.
Published: (2025)