RoboAlign: Learning Test-Time Reasoning for Language-Action Alignment in Vision-Language-Action Models
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Kim, Dongyoung, Park, Sumin, Song, Woomin, Kim, Seungku, Kim, Taeyoung, Jang, Huiwon, Shin, Jinwoo, Kim, Jaehyung, Seo, Younggyo |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Robot-R1: Reinforcement Learning for Enhanced Embodied Reasoning in Robotics
von: Kim, Dongyoung, et al.
Veröffentlicht: (2025)
von: Kim, Dongyoung, et al.
Veröffentlicht: (2025)
RoboCurate: Harnessing Diversity with Action-Verified Neural Trajectory for Robot Learning
von: Kim, Seungku, et al.
Veröffentlicht: (2026)
von: Kim, Seungku, et al.
Veröffentlicht: (2026)
Verifier-free Test-Time Sampling for Vision Language Action Models
von: Jang, Suhyeok, et al.
Veröffentlicht: (2025)
von: Jang, Suhyeok, et al.
Veröffentlicht: (2025)
Dual-Stream Diffusion for World-Model Augmented Vision-Language-Action Model
von: Won, John, et al.
Veröffentlicht: (2025)
von: Won, John, et al.
Veröffentlicht: (2025)
Visual Representation Learning with Stochastic Frame Prediction
von: Jang, Huiwon, et al.
Veröffentlicht: (2024)
von: Jang, Huiwon, et al.
Veröffentlicht: (2024)
SpatialBoost: Enhancing Visual Representation through Language-Guided Reasoning
von: Jeon, Byungwoo, et al.
Veröffentlicht: (2026)
von: Jeon, Byungwoo, et al.
Veröffentlicht: (2026)
HAMLET: Switch your Vision-Language-Action Model into a History-Aware Policy
von: Koo, Myungkyu, et al.
Veröffentlicht: (2025)
von: Koo, Myungkyu, et al.
Veröffentlicht: (2025)
ContextVLA: Vision-Language-Action Model with Amortized Multi-Frame Context
von: Jang, Huiwon, et al.
Veröffentlicht: (2025)
von: Jang, Huiwon, et al.
Veröffentlicht: (2025)
Accelerating Reinforcement Learning with Value-Conditional State Entropy Exploration
von: Kim, Dongyoung, et al.
Veröffentlicht: (2023)
von: Kim, Dongyoung, et al.
Veröffentlicht: (2023)
Spread Preference Annotation: Direct Preference Judgment for Efficient LLM Alignment
von: Kim, Dongyoung, et al.
Veröffentlicht: (2024)
von: Kim, Dongyoung, et al.
Veröffentlicht: (2024)
Modular Sensory Stream for Integrating Physical Feedback in Vision-Language-Action Models
von: Lee, Jimin, et al.
Veröffentlicht: (2026)
von: Lee, Jimin, et al.
Veröffentlicht: (2026)
Debiasing Online Preference Learning via Preference Feature Preservation
von: Kim, Dongyoung, et al.
Veröffentlicht: (2025)
von: Kim, Dongyoung, et al.
Veröffentlicht: (2025)
Efficient Long Video Tokenization via Coordinate-based Patch Reconstruction
von: Jang, Huiwon, et al.
Veröffentlicht: (2024)
von: Jang, Huiwon, et al.
Veröffentlicht: (2024)
Personalized Language Models via Privacy-Preserving Evolutionary Model Merging
von: Kim, Kyuyoung, et al.
Veröffentlicht: (2025)
von: Kim, Kyuyoung, et al.
Veröffentlicht: (2025)
Learning to Correct for QA Reasoning with Black-box LLMs
von: Kim, Jaehyung, et al.
Veröffentlicht: (2024)
von: Kim, Jaehyung, et al.
Veröffentlicht: (2024)
Test-Time Training for Visual Foresight Vision-Language-Action Models
von: Park, Sangwu, et al.
Veröffentlicht: (2026)
von: Park, Sangwu, et al.
Veröffentlicht: (2026)
Efficient LLM Collaboration via Planning
von: Lee, Byeongchan, et al.
Veröffentlicht: (2025)
von: Lee, Byeongchan, et al.
Veröffentlicht: (2025)
RoboAlign-R1: Distilled Multimodal Reward Alignment for Robot Video World Models
von: Wu, Hao, et al.
Veröffentlicht: (2026)
von: Wu, Hao, et al.
Veröffentlicht: (2026)
DEAS: DEtached value learning with Action Sequence for Scalable Offline RL
von: Kim, Changyeon, et al.
Veröffentlicht: (2025)
von: Kim, Changyeon, et al.
Veröffentlicht: (2025)
ReVISE: Learning to Refine at Test-Time via Intrinsic Self-Verification
von: Lee, Hyunseok, et al.
Veröffentlicht: (2025)
von: Lee, Hyunseok, et al.
Veröffentlicht: (2025)
Hierarchical Context Merging: Better Long Context Understanding for Pre-trained LLMs
von: Song, Woomin, et al.
Veröffentlicht: (2024)
von: Song, Woomin, et al.
Veröffentlicht: (2024)
Explainable Adversarial-Robust Vision-Language-Action Model for Robotic Manipulation
von: Kim, Ju-Young, et al.
Veröffentlicht: (2025)
von: Kim, Ju-Young, et al.
Veröffentlicht: (2025)
Training-free LLM Verification via Recycling Few-shot Examples
von: Lee, Dongseok, et al.
Veröffentlicht: (2025)
von: Lee, Dongseok, et al.
Veröffentlicht: (2025)
ExComm: Exploration-Stage Communication for Error-Resilient Agentic Test-Time Scaling
von: Song, Woomin, et al.
Veröffentlicht: (2026)
von: Song, Woomin, et al.
Veröffentlicht: (2026)
Restoration-Aligned Generative Flow Models for Blind Motion Deblurring
von: Kim, Insoo, et al.
Veröffentlicht: (2026)
von: Kim, Insoo, et al.
Veröffentlicht: (2026)
Render and Diffuse: Aligning Image and Action Spaces for Diffusion-based Behaviour Cloning
von: Vosylius, Vitalis, et al.
Veröffentlicht: (2024)
von: Vosylius, Vitalis, et al.
Veröffentlicht: (2024)
Tabular Transfer Learning via Prompting LLMs
von: Nam, Jaehyun, et al.
Veröffentlicht: (2024)
von: Nam, Jaehyun, et al.
Veröffentlicht: (2024)
Optimized Feature Generation for Tabular Data via LLMs with Decision Tree Reasoning
von: Nam, Jaehyun, et al.
Veröffentlicht: (2024)
von: Nam, Jaehyun, et al.
Veröffentlicht: (2024)
Pedagogical Alignment for Vision-Language-Action Models: A Comprehensive Framework for Data, Architecture, and Evaluation in Education
von: Lee, Unggi, et al.
Veröffentlicht: (2026)
von: Lee, Unggi, et al.
Veröffentlicht: (2026)
Coarse-to-fine Q-Network with Action Sequence for Data-Efficient Reinforcement Learning
von: Seo, Younggyo, et al.
Veröffentlicht: (2024)
von: Seo, Younggyo, et al.
Veröffentlicht: (2024)
ACG: Action Coherence Guidance for Flow-based Vision-Language-Action models
von: Park, Minho, et al.
Veröffentlicht: (2025)
von: Park, Minho, et al.
Veröffentlicht: (2025)
Attentive Illumination Decomposition Model for Multi-Illuminant White Balancing
von: Kim, Dongyoung, et al.
Veröffentlicht: (2024)
von: Kim, Dongyoung, et al.
Veröffentlicht: (2024)
Reasoning or Fluency? Dissecting Probabilistic Confidence in Best-of-N Selection
von: Kim, Hojin, et al.
Veröffentlicht: (2026)
von: Kim, Hojin, et al.
Veröffentlicht: (2026)
RPM: Reasoning-Level Personalization for Black-Box Large Language Models
von: Kim, Jieyong, et al.
Veröffentlicht: (2025)
von: Kim, Jieyong, et al.
Veröffentlicht: (2025)
RoboMonkey: Scaling Test-Time Sampling and Verification for Vision-Language-Action Models
von: Kwok, Jacky, et al.
Veröffentlicht: (2025)
von: Kwok, Jacky, et al.
Veröffentlicht: (2025)
Structural Reasoning Improves Molecular Understanding of LLM
von: Jang, Yunhui, et al.
Veröffentlicht: (2024)
von: Jang, Yunhui, et al.
Veröffentlicht: (2024)
Probabilistic Vision-Language Representation for Weakly Supervised Temporal Action Localization
von: Lim, Geuntaek, et al.
Veröffentlicht: (2024)
von: Lim, Geuntaek, et al.
Veröffentlicht: (2024)
Align to Misalign: Automatic LLM Jailbreak with Meta-Optimized LLM Judges
von: Koo, Hamin, et al.
Veröffentlicht: (2025)
von: Koo, Hamin, et al.
Veröffentlicht: (2025)
Mixed-Session Conversation with Egocentric Memory
von: Jang, Jihyoung, et al.
Veröffentlicht: (2024)
von: Jang, Jihyoung, et al.
Veröffentlicht: (2024)
VLM2Rec: Resolving Modality Collapse in Vision-Language Model Embedders for Multimodal Sequential Recommendation
von: Kim, Junyoung, et al.
Veröffentlicht: (2026)
von: Kim, Junyoung, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
Robot-R1: Reinforcement Learning for Enhanced Embodied Reasoning in Robotics
von: Kim, Dongyoung, et al.
Veröffentlicht: (2025) -
RoboCurate: Harnessing Diversity with Action-Verified Neural Trajectory for Robot Learning
von: Kim, Seungku, et al.
Veröffentlicht: (2026) -
Verifier-free Test-Time Sampling for Vision Language Action Models
von: Jang, Suhyeok, et al.
Veröffentlicht: (2025) -
Dual-Stream Diffusion for World-Model Augmented Vision-Language-Action Model
von: Won, John, et al.
Veröffentlicht: (2025) -
Visual Representation Learning with Stochastic Frame Prediction
von: Jang, Huiwon, et al.
Veröffentlicht: (2024)