A More Word-like Image Tokenization for MLLMs
Fuente:
arXiv
Saved in:
| Main Authors: | Lee, Hyun, Jeong, Hyemin, Kim, Yejin, Choi, Hyungwook, Cho, Hyunsoo, Kim, Soo Kyung, Lee, Joonseok |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
TripleSumm: Adaptive Triple-Modality Fusion for Video Summarization
by: Kim, Sumin, et al.
Published: (2026)
by: Kim, Sumin, et al.
Published: (2026)
Latent Diffusion Models with Masked AutoEncoders
by: Lee, Junho, et al.
Published: (2025)
by: Lee, Junho, et al.
Published: (2025)
Local Representative Token Guided Merging for Text-to-Image Generation
by: Lee, Min-Jeong, et al.
Published: (2025)
by: Lee, Min-Jeong, et al.
Published: (2025)
Mining Instance-Centric Vision-Language Contexts for Human-Object Interaction Detection
by: Seo, Soo Won, et al.
Published: (2026)
by: Seo, Soo Won, et al.
Published: (2026)
CLIP-KOA: Enhancing Knee Osteoarthritis Diagnosis with Multi-Modal Learning and Symmetry-Aware Loss Functions
by: Jeong, Yejin, et al.
Published: (2025)
by: Jeong, Yejin, et al.
Published: (2025)
DisCoRD: Discrete Tokens to Continuous Motion via Rectified Flow Decoding
by: Cho, Jungbin, et al.
Published: (2024)
by: Cho, Jungbin, et al.
Published: (2024)
Identifiable Token Correspondence for World Models
by: Kim, Youngin, et al.
Published: (2026)
by: Kim, Youngin, et al.
Published: (2026)
ArtSplat: Feed-Forward Articulated 3D Gaussian Splatting from Sparse Multi-State Uncalibrated Views
by: Lee, Inseo, et al.
Published: (2026)
by: Lee, Inseo, et al.
Published: (2026)
IFCap: Image-like Retrieval and Frequency-based Entity Filtering for Zero-shot Captioning
by: Lee, Soeun, et al.
Published: (2024)
by: Lee, Soeun, et al.
Published: (2024)
Latent Expression Generation for Referring Image Segmentation and Grounding
by: Yu, Seonghoon, et al.
Published: (2025)
by: Yu, Seonghoon, et al.
Published: (2025)
OTT-Vid: Optimal Transport Temporal Token Compression for Video Large Language Models
by: Kang, Minseok, et al.
Published: (2026)
by: Kang, Minseok, et al.
Published: (2026)
Robust 3D Shape Reconstruction in Zero-Shot from a Single Image in the Wild
by: Cho, Junhyeong, et al.
Published: (2024)
by: Cho, Junhyeong, et al.
Published: (2024)
Learning Equi-angular Representations for Online Continual Learning
by: Seo, Minhyuk, et al.
Published: (2024)
by: Seo, Minhyuk, et al.
Published: (2024)
Learning to Explore for Stochastic Gradient MCMC
by: Kim, SeungHyun, et al.
Published: (2024)
by: Kim, SeungHyun, et al.
Published: (2024)
Towards Motion-aware Referring Image Segmentation
by: Kim, Chaeyun, et al.
Published: (2026)
by: Kim, Chaeyun, et al.
Published: (2026)
MIRROR: Multimodal Cognitive Reframing Therapy for Rolling with Resistance
by: Kim, Subin, et al.
Published: (2025)
by: Kim, Subin, et al.
Published: (2025)
Text-Aware Image Restoration with Diffusion Models
by: Min, Jaewon, et al.
Published: (2025)
by: Min, Jaewon, et al.
Published: (2025)
Planning in 8 Tokens: A Compact Discrete Tokenizer for Latent World Model
by: Kim, Dongwon, et al.
Published: (2026)
by: Kim, Dongwon, et al.
Published: (2026)
GenOL: Generating Diverse Examples for Name-only Online Learning
by: Seo, Minhyuk, et al.
Published: (2024)
by: Seo, Minhyuk, et al.
Published: (2024)
DIAMOND: An LLM-Driven Agent for Context-Aware Baseball Highlight Summarization
by: Kang, Jeonghun, et al.
Published: (2025)
by: Kang, Jeonghun, et al.
Published: (2025)
Training-Free Restoration of Pruned Neural Networks
by: Lee, Keonho, et al.
Published: (2025)
by: Lee, Keonho, et al.
Published: (2025)
Negative Token Merging: Image-based Adversarial Feature Guidance
by: Singh, Jaskirat, et al.
Published: (2024)
by: Singh, Jaskirat, et al.
Published: (2024)
From Drop-off to Recovery: A Mechanistic Analysis of Segmentation in MLLMs
by: Wu, Boyong, et al.
Published: (2026)
by: Wu, Boyong, et al.
Published: (2026)
Cooperative Inference for Real-Time 3D Human Pose Estimation in Multi-Device Edge Networks
by: Choi, Hyun-Ho, et al.
Published: (2025)
by: Choi, Hyun-Ho, et al.
Published: (2025)
Efficient Policy Adaptation with Contrastive Prompt Ensemble for Embodied Agents
by: Choi, Wonje, et al.
Published: (2024)
by: Choi, Wonje, et al.
Published: (2024)
Domain-Invariant Per-Frame Feature Extraction for Cross-Domain Imitation Learning with Visual Observations
by: Kim, Minung, et al.
Published: (2025)
by: Kim, Minung, et al.
Published: (2025)
A Training-Free, Task-Agnostic Framework for Enhancing MLLM Performance on High-Resolution Images
by: Lee, Jaeseong, et al.
Published: (2025)
by: Lee, Jaeseong, et al.
Published: (2025)
AdaRank: Adaptive Rank Pruning for Enhanced Model Merging
by: Lee, Chanhyuk, et al.
Published: (2025)
by: Lee, Chanhyuk, et al.
Published: (2025)
Improving Diffusion-Based Image Editing Faithfulness via Guidance and Scheduling
by: Cho, Hansam, et al.
Published: (2025)
by: Cho, Hansam, et al.
Published: (2025)
Geometrical Properties of Text Token Embeddings for Strong Semantic Binding in Text-to-Image Generation
by: Seo, Hoigi, et al.
Published: (2025)
by: Seo, Hoigi, et al.
Published: (2025)
EdgeFusion: On-Device Text-to-Image Generation
by: Castells, Thibault, et al.
Published: (2024)
by: Castells, Thibault, et al.
Published: (2024)
CRiM-GS: Continuous Rigid Motion-Aware Gaussian Splatting from Motion-Blurred Images
by: Lee, Jungho, et al.
Published: (2024)
by: Lee, Jungho, et al.
Published: (2024)
LogicQA: Logical Anomaly Detection with Vision Language Model Generated Questions
by: Kwon, Yejin, et al.
Published: (2025)
by: Kwon, Yejin, et al.
Published: (2025)
Boost Your Human Image Generation Model via Direct Preference Optimization
by: Na, Sanghyeon, et al.
Published: (2024)
by: Na, Sanghyeon, et al.
Published: (2024)
CountSteer: Steering Attention for Object Counting in Diffusion Models
by: Boo, Hyemin, et al.
Published: (2025)
by: Boo, Hyemin, et al.
Published: (2025)
Frequency-Aware Token Reduction for Efficient Vision Transformer
by: Lee, Dong-Jae, et al.
Published: (2025)
by: Lee, Dong-Jae, et al.
Published: (2025)
DiffBlender: Composable and Versatile Multimodal Text-to-Image Diffusion Models
by: Kim, Sungnyun, et al.
Published: (2023)
by: Kim, Sungnyun, et al.
Published: (2023)
Bidirectional Multimodal Prompt Learning with Scale-Aware Training for Few-Shot Multi-Class Anomaly Detection
by: Lee, Yujin, et al.
Published: (2024)
by: Lee, Yujin, et al.
Published: (2024)
Text Change Detection in Multilingual Documents Using Image Comparison
by: Park, Doyoung, et al.
Published: (2024)
by: Park, Doyoung, et al.
Published: (2024)
Learned Image Compression and Restoration for Digital Pathology
by: Lee, SeonYeong, et al.
Published: (2025)
by: Lee, SeonYeong, et al.
Published: (2025)
Similar Items
-
TripleSumm: Adaptive Triple-Modality Fusion for Video Summarization
by: Kim, Sumin, et al.
Published: (2026) -
Latent Diffusion Models with Masked AutoEncoders
by: Lee, Junho, et al.
Published: (2025) -
Local Representative Token Guided Merging for Text-to-Image Generation
by: Lee, Min-Jeong, et al.
Published: (2025) -
Mining Instance-Centric Vision-Language Contexts for Human-Object Interaction Detection
by: Seo, Soo Won, et al.
Published: (2026) -
CLIP-KOA: Enhancing Knee Osteoarthritis Diagnosis with Multi-Modal Learning and Symmetry-Aware Loss Functions
by: Jeong, Yejin, et al.
Published: (2025)