Identifiable Token Correspondence for World Models
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Kim, Youngin, Sun, Ray, Kim, Inho, Park, Bumsoo, Song, Hyun Oh |
|---|---|
| Format: | Preprint |
| Publié: |
2026
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
DisCoRD: Discrete Tokens to Continuous Motion via Rectified Flow Decoding
par: Cho, Jungbin, et autres
Publié: (2024)
par: Cho, Jungbin, et autres
Publié: (2024)
MultiFloodSynth: Multi-Annotated Flood Synthetic Dataset Generation
par: Kang, YoonJe, et autres
Publié: (2025)
par: Kang, YoonJe, et autres
Publié: (2025)
ToonAging: Face Re-Aging upon Artistic Portrait Style Transfer
par: Kim, Bumsoo, et autres
Publié: (2024)
par: Kim, Bumsoo, et autres
Publié: (2024)
A More Word-like Image Tokenization for MLLMs
par: Lee, Hyun, et autres
Publié: (2026)
par: Lee, Hyun, et autres
Publié: (2026)
Make VLM Recognize Visual Hallucination on Cartoon Character Image with Pose Information
par: Kim, Bumsoo, et autres
Publié: (2024)
par: Kim, Bumsoo, et autres
Publié: (2024)
Robust 3D Shape Reconstruction in Zero-Shot from a Single Image in the Wild
par: Cho, Junhyeong, et autres
Publié: (2024)
par: Cho, Junhyeong, et autres
Publié: (2024)
Video Face Re-Aging: Toward Temporally Consistent Face Re-Aging
par: Muqeet, Abdul, et autres
Publié: (2023)
par: Muqeet, Abdul, et autres
Publié: (2023)
Efficient World Models with Context-Aware Tokenization
par: Micheli, Vincent, et autres
Publié: (2024)
par: Micheli, Vincent, et autres
Publié: (2024)
Extending Foundational Monocular Depth Estimators to Fisheye Cameras with Calibration Tokens
par: Gangopadhyay, Suchisrit, et autres
Publié: (2025)
par: Gangopadhyay, Suchisrit, et autres
Publié: (2025)
DarkQA: Benchmarking Vision-Language Models on Visual-Primitive Question Answering in Low-Light Indoor Scenes
par: Park, Yohan, et autres
Publié: (2025)
par: Park, Yohan, et autres
Publié: (2025)
Scratching Visual Transformer's Back with Uniform Attention
par: Hyeon-Woo, Nam, et autres
Publié: (2022)
par: Hyeon-Woo, Nam, et autres
Publié: (2022)
Text-Aware Image Restoration with Diffusion Models
par: Min, Jaewon, et autres
Publié: (2025)
par: Min, Jaewon, et autres
Publié: (2025)
Task Prototype-Based Knowledge Retrieval for Multi-Task Learning from Partially Annotated Data
par: Oh, Youngmin, et autres
Publié: (2026)
par: Oh, Youngmin, et autres
Publié: (2026)
Toward Stable World Models: Measuring and Addressing World Instability in Generative Environments
par: Kwon, Soonwoo, et autres
Publié: (2025)
par: Kwon, Soonwoo, et autres
Publié: (2025)
Transferable Model-agnostic Vision-Language Model Adaptation for Efficient Weak-to-Strong Generalization
par: Park, Jihwan, et autres
Publié: (2025)
par: Park, Jihwan, et autres
Publié: (2025)
Mean-field Chaos Diffusion Models
par: Park, Sungwoo, et autres
Publié: (2024)
par: Park, Sungwoo, et autres
Publié: (2024)
Exploring the Spectrum of Visio-Linguistic Compositionality and Recognition
par: Oh, Youngtaek, et autres
Publié: (2024)
par: Oh, Youngtaek, et autres
Publié: (2024)
Minecraft-ify: Minecraft Style Image Generation with Text-guided Image Editing for In-Game Application
par: Kim, Bumsoo, et autres
Publié: (2024)
par: Kim, Bumsoo, et autres
Publié: (2024)
Learning to Explore for Stochastic Gradient MCMC
par: Kim, SeungHyun, et autres
Publié: (2024)
par: Kim, SeungHyun, et autres
Publié: (2024)
End-to-End Breast Cancer Radiotherapy Planning via LMMs with Consistency Embedding
par: Kim, Kwanyoung, et autres
Publié: (2023)
par: Kim, Kwanyoung, et autres
Publié: (2023)
Interpreting Attention Heads for Image-to-Text Information Flow in Large Vision-Language Models
par: Kim, Jinyeong, et autres
Publié: (2025)
par: Kim, Jinyeong, et autres
Publié: (2025)
DeepRepViz: Identifying Confounders in Deep Learning Model Predictions
par: Rane, Roshan Prakash, et autres
Publié: (2023)
par: Rane, Roshan Prakash, et autres
Publié: (2023)
MIMIC: Masked Image Modeling with Image Correspondences
par: Marathe, Kalyani, et autres
Publié: (2023)
par: Marathe, Kalyani, et autres
Publié: (2023)
Deep Variational Bayesian Modeling of Haze Degradation Process
par: Im, Eun Woo, et autres
Publié: (2024)
par: Im, Eun Woo, et autres
Publié: (2024)
Self-Rectifying Diffusion Sampling with Perturbed-Attention Guidance
par: Ahn, Donghoon, et autres
Publié: (2024)
par: Ahn, Donghoon, et autres
Publié: (2024)
OTSeg: Multi-prompt Sinkhorn Attention for Zero-Shot Semantic Segmentation
par: Kim, Kwanyoung, et autres
Publié: (2024)
par: Kim, Kwanyoung, et autres
Publié: (2024)
Diffusion-Based Offline RL for Improved Decision-Making in Augmented ARC Task
par: Kim, Yunho, et autres
Publié: (2024)
par: Kim, Yunho, et autres
Publié: (2024)
LLM4SGG: Large Language Models for Weakly Supervised Scene Graph Generation
par: Kim, Kibum, et autres
Publié: (2023)
par: Kim, Kibum, et autres
Publié: (2023)
TripleSumm: Adaptive Triple-Modality Fusion for Video Summarization
par: Kim, Sumin, et autres
Publié: (2026)
par: Kim, Sumin, et autres
Publié: (2026)
SynRES: Towards Referring Expression Segmentation in the Wild via Synthetic Data
par: Kim, Dong-Hee, et autres
Publié: (2025)
par: Kim, Dong-Hee, et autres
Publié: (2025)
Learning Relative Representations for Fine-Grained Multimodal Alignment with Limited Data
par: Kim, Shiwon, et autres
Publié: (2026)
par: Kim, Shiwon, et autres
Publié: (2026)
DefectFill: Realistic Defect Generation with Inpainting Diffusion Model for Visual Inspection
par: Song, Jaewoo, et autres
Publié: (2025)
par: Song, Jaewoo, et autres
Publié: (2025)
Deep learning for precipitation nowcasting: A survey from the perspective of time series forecasting
par: An, Sojung, et autres
Publié: (2024)
par: An, Sojung, et autres
Publié: (2024)
VDInstruct: Zero-Shot Key Information Extraction via Content-Aware Vision Tokenization
par: Nguyen, Son, et autres
Publié: (2025)
par: Nguyen, Son, et autres
Publié: (2025)
Generalized Consistency Trajectory Models for Image Manipulation
par: Kim, Beomsu, et autres
Publié: (2024)
par: Kim, Beomsu, et autres
Publié: (2024)
Test-time Alignment of Diffusion Models without Reward Over-optimization
par: Kim, Sunwoo, et autres
Publié: (2025)
par: Kim, Sunwoo, et autres
Publié: (2025)
Few Shot Part Segmentation Reveals Compositional Logic for Industrial Anomaly Detection
par: Kim, Soopil, et autres
Publié: (2023)
par: Kim, Soopil, et autres
Publié: (2023)
VLM-SubtleBench: How Far Are VLMs from Human-Level Subtle Comparative Reasoning?
par: Kim, Minkyu, et autres
Publié: (2026)
par: Kim, Minkyu, et autres
Publié: (2026)
Overcoming Data Inequality across Domains with Semi-Supervised Domain Generalization
par: Park, Jinha, et autres
Publié: (2024)
par: Park, Jinha, et autres
Publié: (2024)
Decoupling Augmentation Bias in Prompt Learning for Vision-Language Models
par: Kim, Gahyeon, et autres
Publié: (2025)
par: Kim, Gahyeon, et autres
Publié: (2025)
Documents similaires
-
DisCoRD: Discrete Tokens to Continuous Motion via Rectified Flow Decoding
par: Cho, Jungbin, et autres
Publié: (2024) -
MultiFloodSynth: Multi-Annotated Flood Synthetic Dataset Generation
par: Kang, YoonJe, et autres
Publié: (2025) -
ToonAging: Face Re-Aging upon Artistic Portrait Style Transfer
par: Kim, Bumsoo, et autres
Publié: (2024) -
A More Word-like Image Tokenization for MLLMs
par: Lee, Hyun, et autres
Publié: (2026) -
Make VLM Recognize Visual Hallucination on Cartoon Character Image with Pose Information
par: Kim, Bumsoo, et autres
Publié: (2024)