Pri4R: Learning World Dynamics for Vision-Language-Action Models with Privileged 4D Representation
Fuente:
arXiv
Saved in:
| Main Authors: | Kim, Jisoo, Cho, Jungbin, Chu, Sanghyeok, Bal, Ananya, Kim, Jinhyung, Lee, Gunhee, Lee, Sihaeng, Kim, Seung Hwan, Han, Bohyung, Lee, Hyunmin, Jeni, Laszlo A., Kim, Seungryong |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
MASH-VLM: Mitigating Action-Scene Hallucination in Video-LLMs through Disentangled Spatial-Temporal Representations
by: Bae, Kyungho, et al.
Published: (2025)
by: Bae, Kyungho, et al.
Published: (2025)
Enhancing Mixture-of-Experts Specialization via Cluster-Aware Upcycling
by: Chu, Sanghyeok, et al.
Published: (2026)
by: Chu, Sanghyeok, et al.
Published: (2026)
SceneAdapt: Scene-aware Adaptation of Human Motion Diffusion
by: Cho, Jungbin, et al.
Published: (2025)
by: Cho, Jungbin, et al.
Published: (2025)
See It All: Contextualized Late Aggregation for 3D Dense Captioning
by: Kim, Minjung, et al.
Published: (2024)
by: Kim, Minjung, et al.
Published: (2024)
AVIN-Chat: An Audio-Visual Interactive Chatbot System with Emotional State Tuning
by: Park, Chanhyuk, et al.
Published: (2024)
by: Park, Chanhyuk, et al.
Published: (2024)
TAG: Tangential Amplifying Guidance for Hallucination-Resistant Sampling
by: Cho, Hyunmin, et al.
Published: (2025)
by: Cho, Hyunmin, et al.
Published: (2025)
A Proof of the Exact Convergence Rate of Gradient Descent
by: Kim, Jungbin
Published: (2024)
by: Kim, Jungbin
Published: (2024)
A Proof of Exact Convergence Rate of Gradient Descent. Part I. Performance Criterion $\Vert \nabla f(x_N)\Vert^2/(f(x_0)-f_*)$
by: Kim, Jungbin
Published: (2024)
by: Kim, Jungbin
Published: (2024)
ReSpec: Relevance and Specificity Grounded Online Filtering for Learning on Video-Text Data Streams
by: Kim, Chris Dongjoo, et al.
Published: (2025)
by: Kim, Chris Dongjoo, et al.
Published: (2025)
GP-4DGS: Probabilistic 4D Gaussian Splatting from Monocular Video via Variational Gaussian Processes
by: Kim, Mijeong, et al.
Published: (2026)
by: Kim, Mijeong, et al.
Published: (2026)
DEEPTalk: Dynamic Emotion Embedding for Probabilistic Speech-Driven 3D Face Animation
by: Kim, Jisoo, et al.
Published: (2024)
by: Kim, Jisoo, et al.
Published: (2024)
PhysGaia: A Physics-Aware Benchmark with Multi-Body Interactions for Dynamic Novel View Synthesis
by: Kim, Mijeong, et al.
Published: (2025)
by: Kim, Mijeong, et al.
Published: (2025)
Hybrid Video Diffusion Models with 2D Triplane and 3D Wavelet Representation
by: Kim, Kihong, et al.
Published: (2024)
by: Kim, Kihong, et al.
Published: (2024)
EgoSpeak: Learning When to Speak for Egocentric Conversational Agents in the Wild
by: Kim, Junhyeok, et al.
Published: (2025)
by: Kim, Junhyeok, et al.
Published: (2025)
ParamISP: Learned Forward and Inverse ISPs using Camera Parameters
by: Kim, Woohyeok, et al.
Published: (2023)
by: Kim, Woohyeok, et al.
Published: (2023)
IB-GAN: Disentangled Representation Learning with Information Bottleneck Generative Adversarial Networks
by: Jeon, Insu, et al.
Published: (2025)
by: Jeon, Insu, et al.
Published: (2025)
Diffusion Model for Dense Matching
by: Nam, Jisu, et al.
Published: (2023)
by: Nam, Jisu, et al.
Published: (2023)
DisCoRD: Discrete Tokens to Continuous Motion via Rectified Flow Decoding
by: Cho, Jungbin, et al.
Published: (2024)
by: Cho, Jungbin, et al.
Published: (2024)
InterRVOS: Interaction-aware Referring Video Object Segmentation
by: Jin, Woojeong, et al.
Published: (2025)
by: Jin, Woojeong, et al.
Published: (2025)
Fine-Grained Captioning of Long Videos through Scene Graph Consolidation
by: Chu, Sanghyeok, et al.
Published: (2025)
by: Chu, Sanghyeok, et al.
Published: (2025)
Leveraging Temporal Contextualization for Video Action Recognition
by: Kim, Minji, et al.
Published: (2024)
by: Kim, Minji, et al.
Published: (2024)
Fourier-Guided Attention Upsampling for Image Super-Resolution
by: Choi, Daejune, et al.
Published: (2025)
by: Choi, Daejune, et al.
Published: (2025)
Recasting Continual Learning as Sequence Modeling
by: Lee, Soochan, et al.
Published: (2023)
by: Lee, Soochan, et al.
Published: (2023)
Compositional Conservatism: A Transductive Approach in Offline Reinforcement Learning
by: Song, Yeda, et al.
Published: (2024)
by: Song, Yeda, et al.
Published: (2024)
Distilling Reinforcement Learning Algorithms for In-Context Model-Based Planning
by: Son, Jaehyeon, et al.
Published: (2025)
by: Son, Jaehyeon, et al.
Published: (2025)
When Meta-Learning Meets Online and Continual Learning: A Survey
by: Son, Jaehyeon, et al.
Published: (2023)
by: Son, Jaehyeon, et al.
Published: (2023)
Can Language Models Laugh at YouTube Short-form Videos?
by: Ko, Dayoon, et al.
Published: (2023)
by: Ko, Dayoon, et al.
Published: (2023)
AnthroTAP: Learning Point Tracking with Real-World Motion
by: Kim, Inès Hyeonsu, et al.
Published: (2025)
by: Kim, Inès Hyeonsu, et al.
Published: (2025)
A Diffusion-Based Framework for Configurable and Realistic Multi-Storage Trace Generation
by: Kim, Seohyun, et al.
Published: (2025)
by: Kim, Seohyun, et al.
Published: (2025)
Representation Shift: Unifying Token Compression with FlashAttention
by: Choi, Joonmyung, et al.
Published: (2025)
by: Choi, Joonmyung, et al.
Published: (2025)
Superpixel Tokenization for Vision Transformers: Preserving Semantic Integrity in Visual Tokens
by: Lew, Jaihyun, et al.
Published: (2024)
by: Lew, Jaihyun, et al.
Published: (2024)
Accelerating Vision Transformers with Adaptive Patch Sizes
by: Choudhury, Rohan, et al.
Published: (2025)
by: Choudhury, Rohan, et al.
Published: (2025)
MORPHOS: Autoregressive 4D Generation with Temporal Structured Latents
by: Kwon, Minkyung, et al.
Published: (2026)
by: Kwon, Minkyung, et al.
Published: (2026)
Seurat: From Moving Points to Depth
by: Cho, Seokju, et al.
Published: (2025)
by: Cho, Seokju, et al.
Published: (2025)
Layer-wise Update Aggregation with Recycling for Communication-Efficient Federated Learning
by: Kim, Jisoo, et al.
Published: (2025)
by: Kim, Jisoo, et al.
Published: (2025)
Near-field Meta-optics
by: Lee, Dongyoung, et al.
Published: (2026)
by: Lee, Dongyoung, et al.
Published: (2026)
Multi-criteria Token Fusion with One-step-ahead Attention for Efficient Vision Transformers
by: Lee, Sanghyeok, et al.
Published: (2024)
by: Lee, Sanghyeok, et al.
Published: (2024)
EfficientViM: Efficient Vision Mamba with Hidden State Mixer based State Space Duality
by: Lee, Sanghyeok, et al.
Published: (2024)
by: Lee, Sanghyeok, et al.
Published: (2024)
Prompt Flow Integrity to Prevent Privilege Escalation in LLM Agents
by: Kim, Juhee, et al.
Published: (2025)
by: Kim, Juhee, et al.
Published: (2025)
ALPS: Automated Least-Privilege Enforcement for Securing Serverless Functions
by: Shin, Changhee, et al.
Published: (2026)
by: Shin, Changhee, et al.
Published: (2026)
Similar Items
-
MASH-VLM: Mitigating Action-Scene Hallucination in Video-LLMs through Disentangled Spatial-Temporal Representations
by: Bae, Kyungho, et al.
Published: (2025) -
Enhancing Mixture-of-Experts Specialization via Cluster-Aware Upcycling
by: Chu, Sanghyeok, et al.
Published: (2026) -
SceneAdapt: Scene-aware Adaptation of Human Motion Diffusion
by: Cho, Jungbin, et al.
Published: (2025) -
See It All: Contextualized Late Aggregation for 3D Dense Captioning
by: Kim, Minjung, et al.
Published: (2024) -
AVIN-Chat: An Audio-Visual Interactive Chatbot System with Emotional State Tuning
by: Park, Chanhyuk, et al.
Published: (2024)