Planning in 8 Tokens: A Compact Discrete Tokenizer for Latent World Model
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Kim, Dongwon, Seo, Gawon, Lee, Jinsung, Cho, Minsu, Kwak, Suha |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Structured State-Space Regularization for Generation-Friendly Image Tokenization
von: Lee, Jinsung, et al.
Veröffentlicht: (2026)
von: Lee, Jinsung, et al.
Veröffentlicht: (2026)
SToRM: Supervised Token Reduction for Multi-modal LLMs toward efficient end-to-end autonomous driving
von: Kim, Seo Hyun, et al.
Veröffentlicht: (2026)
von: Kim, Seo Hyun, et al.
Veröffentlicht: (2026)
Democratizing Text-to-Image Masked Generative Models with Compact Text-Aware One-Dimensional Tokens
von: Kim, Dongwon, et al.
Veröffentlicht: (2025)
von: Kim, Dongwon, et al.
Veröffentlicht: (2025)
DEF-oriCORN: efficient 3D scene understanding for robust language-directed manipulation without demonstrations
von: Son, Dongwon, et al.
Veröffentlicht: (2024)
von: Son, Dongwon, et al.
Veröffentlicht: (2024)
Sparse Imagination for Efficient Visual World Model Planning
von: Chun, Junha, et al.
Veröffentlicht: (2025)
von: Chun, Junha, et al.
Veröffentlicht: (2025)
GST-VLA: Structured Gaussian Spatial Tokens for 3D Depth-Aware Vision-Language-Action Models
von: Sarowar, Md Selim, et al.
Veröffentlicht: (2026)
von: Sarowar, Md Selim, et al.
Veröffentlicht: (2026)
Chain of World: World Model Thinking in Latent Motion
von: Yang, Fuxiang, et al.
Veröffentlicht: (2026)
von: Yang, Fuxiang, et al.
Veröffentlicht: (2026)
Tempered Self-Similarity Alignment for Physically Plausible Video Generation
von: Kim, Manjin, et al.
Veröffentlicht: (2026)
von: Kim, Manjin, et al.
Veröffentlicht: (2026)
Part-Aware Bottom-Up Group Reasoning for Fine-Grained Social Interaction Detection
von: Kim, Dongkeun, et al.
Veröffentlicht: (2025)
von: Kim, Dongkeun, et al.
Veröffentlicht: (2025)
Pri4R: Learning World Dynamics for Vision-Language-Action Models with Privileged 4D Representation
von: Kim, Jisoo, et al.
Veröffentlicht: (2026)
von: Kim, Jisoo, et al.
Veröffentlicht: (2026)
DiLA: Disentangled Latent Action World Models
von: Zhang, Tianqiu, et al.
Veröffentlicht: (2026)
von: Zhang, Tianqiu, et al.
Veröffentlicht: (2026)
Classification Matters: Improving Video Action Detection with Class-Specific Attention
von: Lee, Jinsung, et al.
Veröffentlicht: (2024)
von: Lee, Jinsung, et al.
Veröffentlicht: (2024)
FreqPolicy: Frequency Autoregressive Visuomotor Policy with Continuous Tokens
von: Zhong, Yiming, et al.
Veröffentlicht: (2025)
von: Zhong, Yiming, et al.
Veröffentlicht: (2025)
Image Tokens Matter: Mitigating Hallucination in Discrete Tokenizer-based Large Vision-Language Models via Latent Editing
von: Wang, Weixing, et al.
Veröffentlicht: (2025)
von: Wang, Weixing, et al.
Veröffentlicht: (2025)
KITE: Keyframe-Indexed Tokenized Evidence for VLM-Based Robot Failure Analysis
von: Hosseinzadeh, Mehdi, et al.
Veröffentlicht: (2026)
von: Hosseinzadeh, Mehdi, et al.
Veröffentlicht: (2026)
Towards More Practical Group Activity Detection: A New Benchmark and Model
von: Kim, Dongkeun, et al.
Veröffentlicht: (2023)
von: Kim, Dongkeun, et al.
Veröffentlicht: (2023)
Interactive Spatiotemporal Token Attention Network for Skeleton-based General Interactive Action Recognition
von: Wen, Yuhang, et al.
Veröffentlicht: (2023)
von: Wen, Yuhang, et al.
Veröffentlicht: (2023)
SweetTok: Semantic-Aware Spatial-Temporal Tokenizer for Compact Video Discretization
von: Tan, Zhentao, et al.
Veröffentlicht: (2024)
von: Tan, Zhentao, et al.
Veröffentlicht: (2024)
Moto: Latent Motion Token as the Bridging Language for Learning Robot Manipulation from Videos
von: Chen, Yi, et al.
Veröffentlicht: (2024)
von: Chen, Yi, et al.
Veröffentlicht: (2024)
Faster or Stronger: Towards Flexible Visual Place Recognition via Weighted Aggregation and Token Pruning
von: Zeng, Zichao, et al.
Veröffentlicht: (2026)
von: Zeng, Zichao, et al.
Veröffentlicht: (2026)
Embodied Tree of Thoughts: Deliberate Manipulation Planning with Embodied World Model
von: Xu, Wenjiang, et al.
Veröffentlicht: (2025)
von: Xu, Wenjiang, et al.
Veröffentlicht: (2025)
Pixel-level Scene Understanding in One Token: Visual States Need What-is-Where Composition
von: Lee, Seokmin, et al.
Veröffentlicht: (2026)
von: Lee, Seokmin, et al.
Veröffentlicht: (2026)
ActFusion: a Unified Diffusion Model for Action Segmentation and Anticipation
von: Gong, Dayoung, et al.
Veröffentlicht: (2024)
von: Gong, Dayoung, et al.
Veröffentlicht: (2024)
Conditioning Latent-Space Clusters for Real-World Anomaly Classification
von: Bogdoll, Daniel, et al.
Veröffentlicht: (2023)
von: Bogdoll, Daniel, et al.
Veröffentlicht: (2023)
GroupCoOp: Group-robust Fine-tuning via Group Prompt Learning
von: Kim, Nayeong, et al.
Veröffentlicht: (2025)
von: Kim, Nayeong, et al.
Veröffentlicht: (2025)
Efficient Robotic Policy Learning via Latent Space Backward Planning
von: Liu, Dongxiu, et al.
Veröffentlicht: (2025)
von: Liu, Dongxiu, et al.
Veröffentlicht: (2025)
Driver-WM: A Driver-Centric Traffic-Conditioned Latent World Model for In-Cabin Dynamics Rollout
von: Chi, Haozhuang, et al.
Veröffentlicht: (2026)
von: Chi, Haozhuang, et al.
Veröffentlicht: (2026)
AdaWorld: Learning Adaptable World Models with Latent Actions
von: Gao, Shenyuan, et al.
Veröffentlicht: (2025)
von: Gao, Shenyuan, et al.
Veröffentlicht: (2025)
Towards Unified World Models for Visual Navigation via Memory-Augmented Planning and Foresight
von: Dong, Yifei, et al.
Veröffentlicht: (2025)
von: Dong, Yifei, et al.
Veröffentlicht: (2025)
Good Token Hunting: A Hitchhiker's Guide to Token Selection for Visual Geometry Transformers
von: Zheng, Shuhong, et al.
Veröffentlicht: (2026)
von: Zheng, Shuhong, et al.
Veröffentlicht: (2026)
Online Temporal Action Localization with Memory-Augmented Transformer
von: Song, Youngkil, et al.
Veröffentlicht: (2024)
von: Song, Youngkil, et al.
Veröffentlicht: (2024)
Learning Unified Distance Metric Across Diverse Data Distributions with Parameter-Efficient Transfer Learning
von: Kim, Sungyeon, et al.
Veröffentlicht: (2023)
von: Kim, Sungyeon, et al.
Veröffentlicht: (2023)
On Epistemic Uncertainty of Visual Tokens for Object Hallucinations in Large Vision-Language Models
von: Seo, Hoigi, et al.
Veröffentlicht: (2025)
von: Seo, Hoigi, et al.
Veröffentlicht: (2025)
FreeAction: Training-Free Techniques for Enhanced Fidelity of Trajectory-to-Video Generation
von: Kim, Seungwook, et al.
Veröffentlicht: (2025)
von: Kim, Seungwook, et al.
Veröffentlicht: (2025)
DriveDreamer-Policy: A Geometry-Grounded World-Action Model for Unified Generation and Planning
von: Zhou, Yang, et al.
Veröffentlicht: (2026)
von: Zhou, Yang, et al.
Veröffentlicht: (2026)
Object-Centric World Model for Language-Guided Manipulation
von: Jeong, Youngjoon, et al.
Veröffentlicht: (2025)
von: Jeong, Youngjoon, et al.
Veröffentlicht: (2025)
Adaptive Length Image Tokenization via Recurrent Allocation
von: Duggal, Shivam, et al.
Veröffentlicht: (2024)
von: Duggal, Shivam, et al.
Veröffentlicht: (2024)
Identifiable Token Correspondence for World Models
von: Kim, Youngin, et al.
Veröffentlicht: (2026)
von: Kim, Youngin, et al.
Veröffentlicht: (2026)
Generating Accurate and Detailed Captions for High-Resolution Images
von: Lee, Hankyeol, et al.
Veröffentlicht: (2025)
von: Lee, Hankyeol, et al.
Veröffentlicht: (2025)
WAM-Flow: Parallel Coarse-to-Fine Motion Planning via Discrete Flow Matching for Autonomous Driving
von: Xu, Yifang, et al.
Veröffentlicht: (2025)
von: Xu, Yifang, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Structured State-Space Regularization for Generation-Friendly Image Tokenization
von: Lee, Jinsung, et al.
Veröffentlicht: (2026) -
SToRM: Supervised Token Reduction for Multi-modal LLMs toward efficient end-to-end autonomous driving
von: Kim, Seo Hyun, et al.
Veröffentlicht: (2026) -
Democratizing Text-to-Image Masked Generative Models with Compact Text-Aware One-Dimensional Tokens
von: Kim, Dongwon, et al.
Veröffentlicht: (2025) -
DEF-oriCORN: efficient 3D scene understanding for robust language-directed manipulation without demonstrations
von: Son, Dongwon, et al.
Veröffentlicht: (2024) -
Sparse Imagination for Efficient Visual World Model Planning
von: Chun, Junha, et al.
Veröffentlicht: (2025)