Generative Visual Code Mobile World Models
Fuente:
arXiv
Salvato in:
| Autori principali: | Koh, Woosung, Han, Sungjun, Lee, Segyu, Yun, Se-Young, Shin, Jamin |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Predicting LLM Reasoning Performance with Small Proxy Model
di: Koh, Woosung, et al.
Pubblicazione: (2025)
di: Koh, Woosung, et al.
Pubblicazione: (2025)
Disentangled Sparse Representations for Concept-Separated Diffusion Unlearning
di: Kim, Hyeonjin, et al.
Pubblicazione: (2026)
di: Kim, Hyeonjin, et al.
Pubblicazione: (2026)
FedSOL: Stabilized Orthogonal Learning with Proximal Restrictions in Federated Learning
di: Lee, Gihun, et al.
Pubblicazione: (2023)
di: Lee, Gihun, et al.
Pubblicazione: (2023)
Toward Stable World Models: Measuring and Addressing World Instability in Generative Environments
di: Kwon, Soonwoo, et al.
Pubblicazione: (2025)
di: Kwon, Soonwoo, et al.
Pubblicazione: (2025)
Learning Equi-angular Representations for Online Continual Learning
di: Seo, Minhyuk, et al.
Pubblicazione: (2024)
di: Seo, Minhyuk, et al.
Pubblicazione: (2024)
Skrr: Skip and Re-use Text Encoder Layers for Memory Efficient Text-to-Image Generation
di: Seo, Hoigi, et al.
Pubblicazione: (2025)
di: Seo, Hoigi, et al.
Pubblicazione: (2025)
RECODE: Reasoning Through Code Generation for Visual Question Answering
di: Shen, Junhong, et al.
Pubblicazione: (2025)
di: Shen, Junhong, et al.
Pubblicazione: (2025)
Tuning-Free Multi-Event Long Video Generation via Synchronized Coupled Sampling
di: Kim, Subin, et al.
Pubblicazione: (2025)
di: Kim, Subin, et al.
Pubblicazione: (2025)
Learning Transformer-based World Models with Contrastive Predictive Coding
di: Burchi, Maxime, et al.
Pubblicazione: (2025)
di: Burchi, Maxime, et al.
Pubblicazione: (2025)
Learning and Leveraging World Models in Visual Representation Learning
di: Garrido, Quentin, et al.
Pubblicazione: (2024)
di: Garrido, Quentin, et al.
Pubblicazione: (2024)
Diffusion for World Modeling: Visual Details Matter in Atari
di: Alonso, Eloi, et al.
Pubblicazione: (2024)
di: Alonso, Eloi, et al.
Pubblicazione: (2024)
PhyGround: Benchmarking Physical Reasoning in Generative World Models
di: Lin, Juyi, et al.
Pubblicazione: (2026)
di: Lin, Juyi, et al.
Pubblicazione: (2026)
UniSAFE: A Comprehensive Benchmark for Safety Evaluation of Unified Multimodal Models
di: Lee, Segyu, et al.
Pubblicazione: (2026)
di: Lee, Segyu, et al.
Pubblicazione: (2026)
Improving Visual Representation Alignment Generation with GRPO
di: Mo, Shentong, et al.
Pubblicazione: (2026)
di: Mo, Shentong, et al.
Pubblicazione: (2026)
DefectFill: Realistic Defect Generation with Inpainting Diffusion Model for Visual Inspection
di: Song, Jaewoo, et al.
Pubblicazione: (2025)
di: Song, Jaewoo, et al.
Pubblicazione: (2025)
HART: Efficient Visual Generation with Hybrid Autoregressive Transformer
di: Tang, Haotian, et al.
Pubblicazione: (2024)
di: Tang, Haotian, et al.
Pubblicazione: (2024)
Machine Learning Modeling for Multi-order Human Visual Motion Processing
di: Sun, Zitang, et al.
Pubblicazione: (2025)
di: Sun, Zitang, et al.
Pubblicazione: (2025)
Adversarial Wear and Tear: Exploiting Natural Damage for Generating Physical-World Adversarial Examples
di: Irshad, Samra, et al.
Pubblicazione: (2025)
di: Irshad, Samra, et al.
Pubblicazione: (2025)
Astra: General Interactive World Model with Autoregressive Denoising
di: Zhu, Yixuan, et al.
Pubblicazione: (2025)
di: Zhu, Yixuan, et al.
Pubblicazione: (2025)
Subtask-Aware Visual Reward Learning from Segmented Demonstrations
di: Kim, Changyeon, et al.
Pubblicazione: (2025)
di: Kim, Changyeon, et al.
Pubblicazione: (2025)
AGDC: Autoregressive Generation of Variable-Length Sequences with Joint Discrete and Continuous Spaces
di: Shin, Yeonsang, et al.
Pubblicazione: (2026)
di: Shin, Yeonsang, et al.
Pubblicazione: (2026)
LatticeWorld: A Multimodal Large Language Model-Empowered Framework for Interactive Complex World Generation
di: Duan, Yinglin, et al.
Pubblicazione: (2025)
di: Duan, Yinglin, et al.
Pubblicazione: (2025)
DMQ: Dissecting Outliers of Diffusion Models for Post-Training Quantization
di: Lee, Dongyeun, et al.
Pubblicazione: (2025)
di: Lee, Dongyeun, et al.
Pubblicazione: (2025)
EfficientTrain++: Generalized Curriculum Learning for Efficient Visual Backbone Training
di: Wang, Yulin, et al.
Pubblicazione: (2024)
di: Wang, Yulin, et al.
Pubblicazione: (2024)
MobilePlantViT: A Mobile-friendly Hybrid ViT for Generalized Plant Disease Image Classification
di: Tonmoy, Moshiur Rahman, et al.
Pubblicazione: (2025)
di: Tonmoy, Moshiur Rahman, et al.
Pubblicazione: (2025)
Retrieval Visual Contrastive Decoding to Mitigate Object Hallucinations in Large Vision-Language Models
di: Lee, Jihoon, et al.
Pubblicazione: (2025)
di: Lee, Jihoon, et al.
Pubblicazione: (2025)
Jodi: Unification of Visual Generation and Understanding via Joint Modeling
di: Xu, Yifeng, et al.
Pubblicazione: (2025)
di: Xu, Yifeng, et al.
Pubblicazione: (2025)
Simulating the Real World: A Unified Survey of Multimodal Generative Models
di: Hu, Yuqi, et al.
Pubblicazione: (2025)
di: Hu, Yuqi, et al.
Pubblicazione: (2025)
Owl-1: Omni World Model for Consistent Long Video Generation
di: Huang, Yuanhui, et al.
Pubblicazione: (2024)
di: Huang, Yuanhui, et al.
Pubblicazione: (2024)
Make VLM Recognize Visual Hallucination on Cartoon Character Image with Pose Information
di: Kim, Bumsoo, et al.
Pubblicazione: (2024)
di: Kim, Bumsoo, et al.
Pubblicazione: (2024)
What Drives Compositional Generalization? The Importance of Continuous Training Objectives in Visual Generative Models
di: Farid, Karim, et al.
Pubblicazione: (2025)
di: Farid, Karim, et al.
Pubblicazione: (2025)
FedWSQ: Efficient Federated Learning with Weight Standardization and Distribution-Aware Non-Uniform Quantization
di: Kim, Seung-Wook, et al.
Pubblicazione: (2025)
di: Kim, Seung-Wook, et al.
Pubblicazione: (2025)
Pix2Code: Learning to Compose Neural Visual Concepts as Programs
di: Wüst, Antonia, et al.
Pubblicazione: (2024)
di: Wüst, Antonia, et al.
Pubblicazione: (2024)
Real2Code: Reconstruct Articulated Objects via Code Generation
di: Mandi, Zhao, et al.
Pubblicazione: (2024)
di: Mandi, Zhao, et al.
Pubblicazione: (2024)
No Alignment Needed for Generation: Learning Linearly Separable Representations in Diffusion Models
di: Yun, Junno, et al.
Pubblicazione: (2025)
di: Yun, Junno, et al.
Pubblicazione: (2025)
World Modeling with Probabilistic Structure Integration
di: Kotar, Klemen, et al.
Pubblicazione: (2025)
di: Kotar, Klemen, et al.
Pubblicazione: (2025)
Domain-Invariant Per-Frame Feature Extraction for Cross-Domain Imitation Learning with Visual Observations
di: Kim, Minung, et al.
Pubblicazione: (2025)
di: Kim, Minung, et al.
Pubblicazione: (2025)
PAN: A World Model for General, Interactable, and Long-Horizon World Simulation
di: PAN Team, et al.
Pubblicazione: (2025)
di: PAN Team, et al.
Pubblicazione: (2025)
Cooperative Meta-Learning with Gradient Augmentation
di: Shin, Jongyun, et al.
Pubblicazione: (2024)
di: Shin, Jongyun, et al.
Pubblicazione: (2024)
ERGO: Efficient High-Resolution Visual Understanding for Vision-Language Models
di: Lee, Jewon, et al.
Pubblicazione: (2025)
di: Lee, Jewon, et al.
Pubblicazione: (2025)
Documenti analoghi
-
Predicting LLM Reasoning Performance with Small Proxy Model
di: Koh, Woosung, et al.
Pubblicazione: (2025) -
Disentangled Sparse Representations for Concept-Separated Diffusion Unlearning
di: Kim, Hyeonjin, et al.
Pubblicazione: (2026) -
FedSOL: Stabilized Orthogonal Learning with Proximal Restrictions in Federated Learning
di: Lee, Gihun, et al.
Pubblicazione: (2023) -
Toward Stable World Models: Measuring and Addressing World Instability in Generative Environments
di: Kwon, Soonwoo, et al.
Pubblicazione: (2025) -
Learning Equi-angular Representations for Online Continual Learning
di: Seo, Minhyuk, et al.
Pubblicazione: (2024)