Slot-MLLM: Object-Centric Visual Tokenization for Multimodal LLM
Fuente:
arXiv
Saved in:
| Main Authors: | Chi, Donghwan, Kim, Hyomin, Oh, Yoonjin, Kim, Yongjin, Lee, Donghoon, Jo, Daejin, Kim, Jongmin, Baek, Junyeob, Ahn, Sungjin, Kim, Sungwoong |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
OSPO: Object-Centric Self-Improving Preference Optimization for Text-to-Image Generation
by: Oh, Yoonjin, et al.
Published: (2025)
by: Oh, Yoonjin, et al.
Published: (2025)
LM-SPT: LM-Aligned Semantic Distillation for Speech Tokenization
by: Jo, Daejin, et al.
Published: (2025)
by: Jo, Daejin, et al.
Published: (2025)
SGPO: Self-Generated Preference Optimization based on Self-Improver
by: Lee, Hyeonji, et al.
Published: (2025)
by: Lee, Hyeonji, et al.
Published: (2025)
Generative Recursive Reasoning
by: Baek, Junyeob, et al.
Published: (2026)
by: Baek, Junyeob, et al.
Published: (2026)
Discrete JEPA: Learning Discrete Token Representations without Reconstruction
by: Baek, Junyeob, et al.
Published: (2025)
by: Baek, Junyeob, et al.
Published: (2025)
EraseLoRA: MLLM-Driven Foreground Exclusion and Background Subtype Aggregation for Dataset-Free Object Removal
by: Jo, Sanghyun, et al.
Published: (2025)
by: Jo, Sanghyun, et al.
Published: (2025)
Learning to Theorize the World from Observation
by: Baek, Doojin, et al.
Published: (2026)
by: Baek, Doojin, et al.
Published: (2026)
SlotVTG: Object-Centric Adapter for Generalizable Video Temporal Grounding
by: Han, Jiwook, et al.
Published: (2026)
by: Han, Jiwook, et al.
Published: (2026)
Term Structure and Risk Premiums of Commodity Futures With Linear Regressions
by: Daejin Kim
Published: (2024)
by: Daejin Kim
Published: (2024)
HyPHEN: A Hybrid Packing Method and Optimizations for Homomorphic Encryption-Based Neural Networks
by: Kim, Donghwan, et al.
Published: (2023)
by: Kim, Donghwan, et al.
Published: (2023)
Dreamweaver: Learning Compositional World Models from Pixels
by: Baek, Junyeob, et al.
Published: (2025)
by: Baek, Junyeob, et al.
Published: (2025)
A non-Hopfian ascending HNN-extension of a finitely presented Hopfian group
by: Kim, Jan, et al.
Published: (2025)
by: Kim, Jan, et al.
Published: (2025)
Hexa: Self-Improving for Knowledge-Grounded Dialogue System
by: Jo, Daejin, et al.
Published: (2023)
by: Jo, Daejin, et al.
Published: (2023)
Low‐Surface‐Energy Capped Hydrogel Micropillar Arrays for Transparent Superhydrophobic Antifogging Surfaces
by: Hyeongjeong Kim, et al.
Published: (2025)
by: Hyeongjeong Kim, et al.
Published: (2025)
Improving Visual Token Reduction via Rectifying Distortions for Efficient Multimodal LLM Inference
by: Cho, Hyeonwoo, et al.
Published: (2026)
by: Cho, Hyeonwoo, et al.
Published: (2026)
MT-Mol:Multi Agent System with Tool-based Reasoning for Molecular Optimization
by: Kim, Hyomin, et al.
Published: (2025)
by: Kim, Hyomin, et al.
Published: (2025)
DNACHUNKER: Learnable Tokenization for DNA Language Models
by: Kim, Taewon, et al.
Published: (2026)
by: Kim, Taewon, et al.
Published: (2026)
PhysHanDI: Physics-Based Reconstruction of Hand-Deformable Object Interactions
by: Lee, Jihyun, et al.
Published: (2026)
by: Lee, Jihyun, et al.
Published: (2026)
Hybrid Neural Representations for Spherical Data
by: Kim, Hyomin, et al.
Published: (2024)
by: Kim, Hyomin, et al.
Published: (2024)
FastSTAR: Spatiotemporal Token Pruning for Efficient Autoregressive Video Synthesis
by: Yune, Sungwoong, et al.
Published: (2026)
by: Yune, Sungwoong, et al.
Published: (2026)
CuxS back contact for CdTe solar cells
by: Donghwan Kim
Published: (2007)
by: Donghwan Kim
Published: (2007)
Numerical Analysis of Therapeutic Effects by Varying Slot Numbers and Slot‐to‐Slot Distance in Microwave Ablation Using Multislot Coaxial Antenna
by: Donghyuk Kim, et al.
Published: (2024)
by: Donghyuk Kim, et al.
Published: (2024)
Optimizing Korean-Centric LLMs via Token Pruning
by: Kim, Hoyeol, et al.
Published: (2026)
by: Kim, Hoyeol, et al.
Published: (2026)
PlanDQ: Hierarchical Plan Orchestration via D-Conductor and Q-Performer
by: Chen, Chang, et al.
Published: (2024)
by: Chen, Chang, et al.
Published: (2024)
Slot State Space Models
by: Jiang, Jindong, et al.
Published: (2024)
by: Jiang, Jindong, et al.
Published: (2024)
SeDi-Instruct: Enhancing Alignment of Language Models through Self-Directed Instruction Generation
by: Kim, Jungwoo, et al.
Published: (2025)
by: Kim, Jungwoo, et al.
Published: (2025)
v1: Learning to Point Visual Tokens for Multimodal Grounded Reasoning
by: Chung, Jiwan, et al.
Published: (2025)
by: Chung, Jiwan, et al.
Published: (2025)
Modulating Water Activity in Electrochemical Systems: Key Factors and Strategies
by: Suyun Lee, et al.
Published: (2026)
by: Suyun Lee, et al.
Published: (2026)
An empirical investigation of the mitigating effect of debt on overinvestment as shareholder rights vary
by: Chune Young Chung, et al.
Published: (2024)
by: Chune Young Chung, et al.
Published: (2024)
MEVG: Multi-event Video Generation with Text-to-Video Models
by: Oh, Gyeongrok, et al.
Published: (2023)
by: Oh, Gyeongrok, et al.
Published: (2023)
HELEA: Hard-Negative Benchmark and LLM-based Reranking for Robust Entity Alignment
by: Jang, Yoonjin, et al.
Published: (2026)
by: Jang, Yoonjin, et al.
Published: (2026)
CiFHER: A Chiplet-Based FHE Accelerator with a Resizable Structure
by: Kim, Sangpyo, et al.
Published: (2023)
by: Kim, Sangpyo, et al.
Published: (2023)
BAT: Balancing Agility and Stability via Online Policy Switching for Long-Horizon Whole-Body Humanoid Control
by: Baek, Donghoon, et al.
Published: (2026)
by: Baek, Donghoon, et al.
Published: (2026)
Inlier-Centric Post-Training Quantization for Object Detection Models
by: Kim, Minsu, et al.
Published: (2026)
by: Kim, Minsu, et al.
Published: (2026)
LieHMR: Autoregressive Human Mesh Recovery with $SO(3)$ Diffusion
by: Kim, Donghwan, et al.
Published: (2025)
by: Kim, Donghwan, et al.
Published: (2025)
Multi-hypotheses Conditioned Point Cloud Diffusion for 3D Human Reconstruction from Occluded Images
by: Kim, Donghwan, et al.
Published: (2024)
by: Kim, Donghwan, et al.
Published: (2024)
Learning to Compose: Improving Object Centric Learning by Injecting Compositionality
by: Jung, Whie, et al.
Published: (2024)
by: Jung, Whie, et al.
Published: (2024)
Discontinuity-preserving Normal Integration with Auxiliary Edges
by: Kim, Hyomin, et al.
Published: (2024)
by: Kim, Hyomin, et al.
Published: (2024)
Rethinking LLM Inference Bottlenecks: Insights from Latent Attention and Mixture-of-Experts
by: Yun, Sungmin, et al.
Published: (2025)
by: Yun, Sungmin, et al.
Published: (2025)
NeuJeans: Private Neural Network Inference with Joint Optimization of Convolution and FHE Bootstrapping
by: Ju, Jae Hyung, et al.
Published: (2023)
by: Ju, Jae Hyung, et al.
Published: (2023)
Similar Items
-
OSPO: Object-Centric Self-Improving Preference Optimization for Text-to-Image Generation
by: Oh, Yoonjin, et al.
Published: (2025) -
LM-SPT: LM-Aligned Semantic Distillation for Speech Tokenization
by: Jo, Daejin, et al.
Published: (2025) -
SGPO: Self-Generated Preference Optimization based on Self-Improver
by: Lee, Hyeonji, et al.
Published: (2025) -
Generative Recursive Reasoning
by: Baek, Junyeob, et al.
Published: (2026) -
Discrete JEPA: Learning Discrete Token Representations without Reconstruction
by: Baek, Junyeob, et al.
Published: (2025)