Exploring Conditions for Diffusion models in Robotic Control
Fuente:
arXiv
Salvato in:
| Autori principali: | Shin, Heeseong, Heo, Byeongho, Han, Dongyoon, Kim, Seungryong, Kim, Taekyung |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Morphing Tokens Draw Strong Masked Image Models
di: Kim, Taekyung, et al.
Pubblicazione: (2023)
di: Kim, Taekyung, et al.
Pubblicazione: (2023)
Learning with Unmasked Tokens Drives Stronger Vision Learners
di: Kim, Taekyung, et al.
Pubblicazione: (2023)
di: Kim, Taekyung, et al.
Pubblicazione: (2023)
Masking meets Supervision: A Strong Learning Alliance
di: Heo, Byeongho, et al.
Pubblicazione: (2023)
di: Heo, Byeongho, et al.
Pubblicazione: (2023)
Token Bottleneck: One Token to Remember Dynamics
di: Kim, Taekyung, et al.
Pubblicazione: (2025)
di: Kim, Taekyung, et al.
Pubblicazione: (2025)
Match me if you can: Semi-Supervised Semantic Correspondence Learning with Unpaired Images
di: Kim, Jiwon, et al.
Pubblicazione: (2023)
di: Kim, Jiwon, et al.
Pubblicazione: (2023)
DenseNets Reloaded: Paradigm Shift Beyond ResNets and ViTs
di: Kim, Donghyun, et al.
Pubblicazione: (2024)
di: Kim, Donghyun, et al.
Pubblicazione: (2024)
AgentRVOS: Reasoning over Object Tracks for Zero-Shot Referring Video Object Segmentation
di: Jin, Woojeong, et al.
Pubblicazione: (2026)
di: Jin, Woojeong, et al.
Pubblicazione: (2026)
Leveraging Temporal Contextualization for Video Action Recognition
di: Kim, Minji, et al.
Pubblicazione: (2024)
di: Kim, Minji, et al.
Pubblicazione: (2024)
Aligned Novel View Image and Geometry Synthesis via Cross-modal Attention Instillation
di: Kwak, Min-Seop, et al.
Pubblicazione: (2025)
di: Kwak, Min-Seop, et al.
Pubblicazione: (2025)
DNNs May Determine Major Properties of Their Outputs Early, with Timing Possibly Driven by Bias
di: Park, Song, et al.
Pubblicazione: (2025)
di: Park, Song, et al.
Pubblicazione: (2025)
Rotary Position Embedding for Vision Transformer
di: Heo, Byeongho, et al.
Pubblicazione: (2024)
di: Heo, Byeongho, et al.
Pubblicazione: (2024)
SeiT++: Masked Token Modeling Improves Storage-efficient Training
di: Lee, Minhyun, et al.
Pubblicazione: (2023)
di: Lee, Minhyun, et al.
Pubblicazione: (2023)
HYPE: Hyperbolic Entailment Filtering for Underspecified Images and Texts
di: Kim, Wonjae, et al.
Pubblicazione: (2024)
di: Kim, Wonjae, et al.
Pubblicazione: (2024)
Seg4Diff: Unveiling Open-Vocabulary Segmentation in Text-to-Image Diffusion Transformers
di: Kim, Chaehyun, et al.
Pubblicazione: (2025)
di: Kim, Chaehyun, et al.
Pubblicazione: (2025)
Unifying Correspondence, Pose and NeRF for Pose-Free Novel View Synthesis from Stereo Pairs
di: Hong, Sunghwan, et al.
Pubblicazione: (2023)
di: Hong, Sunghwan, et al.
Pubblicazione: (2023)
DirecT2V: Large Language Models are Frame-Level Directors for Zero-Shot Text-to-Video Generation
di: Hong, Susung, et al.
Pubblicazione: (2023)
di: Hong, Susung, et al.
Pubblicazione: (2023)
PF3plat: Pose-Free Feed-Forward 3D Gaussian Splatting
di: Hong, Sunghwan, et al.
Pubblicazione: (2024)
di: Hong, Sunghwan, et al.
Pubblicazione: (2024)
RL makes MLLMs see better than SFT
di: Song, Junha, et al.
Pubblicazione: (2025)
di: Song, Junha, et al.
Pubblicazione: (2025)
Adapting Pretrained ViTs with Convolution Injector for Visuo-Motor Control
di: Hwang, Dongyoon, et al.
Pubblicazione: (2024)
di: Hwang, Dongyoon, et al.
Pubblicazione: (2024)
Learning to See What You Need: Gaze Attention for Multimodal Large Language Models
di: Song, Junha, et al.
Pubblicazione: (2026)
di: Song, Junha, et al.
Pubblicazione: (2026)
MaskRIS: Semantic Distortion-aware Data Augmentation for Referring Image Segmentation
di: Lee, Minhyun, et al.
Pubblicazione: (2024)
di: Lee, Minhyun, et al.
Pubblicazione: (2024)
CAT-Seg: Cost Aggregation for Open-Vocabulary Semantic Segmentation
di: Cho, Seokju, et al.
Pubblicazione: (2023)
di: Cho, Seokju, et al.
Pubblicazione: (2023)
Towards Open-Vocabulary Semantic Segmentation Without Semantic Labels
di: Shin, Heeseong, et al.
Pubblicazione: (2024)
di: Shin, Heeseong, et al.
Pubblicazione: (2024)
METAVerse: Meta-Learning Traversability Cost Map for Off-Road Navigation
di: Seo, Junwon, et al.
Pubblicazione: (2023)
di: Seo, Junwon, et al.
Pubblicazione: (2023)
S^4M: Boosting Semi-Supervised Instance Segmentation with SAM
di: Yoon, Heeji, et al.
Pubblicazione: (2025)
di: Yoon, Heeji, et al.
Pubblicazione: (2025)
OmniRobotHome: A Multi-Camera Platform for Real-Time Multiadic Human-Robot Interaction
di: Lee, Junyoung, et al.
Pubblicazione: (2026)
di: Lee, Junyoung, et al.
Pubblicazione: (2026)
A Dual Process VLA: Efficient Robotic Manipulation Leveraging VLM
di: Han, ByungOk, et al.
Pubblicazione: (2024)
di: Han, ByungOk, et al.
Pubblicazione: (2024)
Dual-Stream Diffusion for World-Model Augmented Vision-Language-Action Model
di: Won, John, et al.
Pubblicazione: (2025)
di: Won, John, et al.
Pubblicazione: (2025)
VENTURA: Adapting Image Diffusion Models for Unified Task Conditioned Navigation
di: Zhang, Arthur, et al.
Pubblicazione: (2025)
di: Zhang, Arthur, et al.
Pubblicazione: (2025)
Pixel Motion Diffusion is What We Need for Robot Control
di: Nguyen, E-Ro, et al.
Pubblicazione: (2025)
di: Nguyen, E-Ro, et al.
Pubblicazione: (2025)
HD Maps are Lane Detection Generalizers: A Novel Generative Framework for Single-Source Domain Generalization
di: Lee, Daeun, et al.
Pubblicazione: (2023)
di: Lee, Daeun, et al.
Pubblicazione: (2023)
Similarity of Neural Architectures using Adversarial Attack Transferability
di: Hwang, Jaehui, et al.
Pubblicazione: (2022)
di: Hwang, Jaehui, et al.
Pubblicazione: (2022)
RoboCurate: Harnessing Diversity with Action-Verified Neural Trajectory for Robot Learning
di: Kim, Seungku, et al.
Pubblicazione: (2026)
di: Kim, Seungku, et al.
Pubblicazione: (2026)
PeLiCal: Targetless Extrinsic Calibration via Penetrating Lines for RGB-D Cameras with Limited Co-visibility
di: Shin, Jaeho, et al.
Pubblicazione: (2024)
di: Shin, Jaeho, et al.
Pubblicazione: (2024)
RoEL: Robust Event-based 3D Line Reconstruction
di: Bae, Gwangtak, et al.
Pubblicazione: (2026)
di: Bae, Gwangtak, et al.
Pubblicazione: (2026)
Space-Aware Instruction Tuning: Dataset and Benchmark for Guide Dog Robots Assisting the Visually Impaired
di: Han, ByungOk, et al.
Pubblicazione: (2025)
di: Han, ByungOk, et al.
Pubblicazione: (2025)
Camera Agnostic Two-Head Network for Ego-Lane Inference
di: Song, Chaehyeon, et al.
Pubblicazione: (2024)
di: Song, Chaehyeon, et al.
Pubblicazione: (2024)
BRIC: Bridging Kinematic Plans and Physical Control at Test Time
di: Lim, Dohun, et al.
Pubblicazione: (2025)
di: Lim, Dohun, et al.
Pubblicazione: (2025)
DexVLA: Vision-Language Model with Plug-In Diffusion Expert for General Robot Control
di: Wen, Junjie, et al.
Pubblicazione: (2025)
di: Wen, Junjie, et al.
Pubblicazione: (2025)
Pri4R: Learning World Dynamics for Vision-Language-Action Models with Privileged 4D Representation
di: Kim, Jisoo, et al.
Pubblicazione: (2026)
di: Kim, Jisoo, et al.
Pubblicazione: (2026)
Documenti analoghi
-
Morphing Tokens Draw Strong Masked Image Models
di: Kim, Taekyung, et al.
Pubblicazione: (2023) -
Learning with Unmasked Tokens Drives Stronger Vision Learners
di: Kim, Taekyung, et al.
Pubblicazione: (2023) -
Masking meets Supervision: A Strong Learning Alliance
di: Heo, Byeongho, et al.
Pubblicazione: (2023) -
Token Bottleneck: One Token to Remember Dynamics
di: Kim, Taekyung, et al.
Pubblicazione: (2025) -
Match me if you can: Semi-Supervised Semantic Correspondence Learning with Unpaired Images
di: Kim, Jiwon, et al.
Pubblicazione: (2023)