RADAR: Benchmarking Vision-Language-Action Generalization via Real-World Dynamics, Spatial-Physical Intelligence, and Autonomous Evaluation
Fuente:
arXiv
Salvato in:
| Autori principali: | Chen, Yuhao, Zhan, Zhihao, Lin, Xiaoxin, Song, Zijian, Liu, Hao, Lyu, Qinhan, Zu, Yubo, Chen, Xiao, Liu, Zhiyuan, Pu, Tao, Chen, Tianshui, Wang, Keze, Lin, Liang, Wang, Guangrun |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Stable Language Guidance for Vision-Language-Action Models
di: Zhan, Zhihao, et al.
Pubblicazione: (2026)
di: Zhan, Zhihao, et al.
Pubblicazione: (2026)
TAG: Target-Agnostic Guidance for Stable Object-Centric Inference in Vision-Language-Action Models
di: Zhou, Jiaying, et al.
Pubblicazione: (2026)
di: Zhou, Jiaying, et al.
Pubblicazione: (2026)
Physical Autoregressive Model for Robotic Manipulation without Action Pretraining
di: Song, Zijian, et al.
Pubblicazione: (2025)
di: Song, Zijian, et al.
Pubblicazione: (2025)
Learning Physics from Pretrained Video Models: A Multimodal Continuous and Sequential World Interaction Models for Robotic Manipulation
di: Song, Zijian, et al.
Pubblicazione: (2026)
di: Song, Zijian, et al.
Pubblicazione: (2026)
E0: Enhancing Generalization and Fine-Grained Control in VLA Models via Tweedie Discrete Diffusion
di: Zhan, Zhihao, et al.
Pubblicazione: (2025)
di: Zhan, Zhihao, et al.
Pubblicazione: (2025)
SIRI-Bench: Challenging VLMs' Spatial Intelligence through Complex Reasoning Tasks
di: Song, Zijian, et al.
Pubblicazione: (2025)
di: Song, Zijian, et al.
Pubblicazione: (2025)
Human-Centric Open-Future Task Discovery: Formulation, Benchmark, and Scalable Tree-Based Search
di: Song, Zijian, et al.
Pubblicazione: (2025)
di: Song, Zijian, et al.
Pubblicazione: (2025)
Learning Spatial-Temporal Coherent Correlations for Speech-Preserving Facial Expression Manipulation
di: Chen, Tianshui, et al.
Pubblicazione: (2026)
di: Chen, Tianshui, et al.
Pubblicazione: (2026)
SQLNet: Scale-Modulated Query and Localization Network for Few-Shot Class-Agnostic Counting
di: Wu, Hefeng, et al.
Pubblicazione: (2023)
di: Wu, Hefeng, et al.
Pubblicazione: (2023)
GS: Generative Segmentation via Label Diffusion
di: Chen, Yuhao, et al.
Pubblicazione: (2025)
di: Chen, Yuhao, et al.
Pubblicazione: (2025)
Robotic Manipulation is Vision-to-Geometry Mapping ($f(v) \rightarrow G$): Vision-Geometry Backbones over Language and Video Models
di: Song, Zijian, et al.
Pubblicazione: (2026)
di: Song, Zijian, et al.
Pubblicazione: (2026)
Rational ANOVA Networks
di: Zhang, Jusheng, et al.
Pubblicazione: (2026)
di: Zhang, Jusheng, et al.
Pubblicazione: (2026)
OOWM: Structuring Embodied Reasoning and Planning via Object-Oriented Programmatic World Modeling
di: Chen, Hongyu, et al.
Pubblicazione: (2026)
di: Chen, Hongyu, et al.
Pubblicazione: (2026)
Bridging the Discrete-Continuous Gap: Unified Multimodal Generation via Coupled Manifold Discrete Absorbing Diffusion
di: Xu, Yuanfeng, et al.
Pubblicazione: (2026)
di: Xu, Yuanfeng, et al.
Pubblicazione: (2026)
In-Situ Tweedie Discrete Diffusion Models
di: Li, Xiao, et al.
Pubblicazione: (2025)
di: Li, Xiao, et al.
Pubblicazione: (2025)
Exploiting Temporal Audio-Visual Correlation Embedding for Audio-Driven One-Shot Talking Head Animation
di: Xu, Zhihua, et al.
Pubblicazione: (2025)
di: Xu, Zhihua, et al.
Pubblicazione: (2025)
CoAgent: Collaborative Planning and Consistency Agent for Coherent Video Generation
di: Zeng, Qinglin, et al.
Pubblicazione: (2025)
di: Zeng, Qinglin, et al.
Pubblicazione: (2025)
Towards CausalGPT: A Multi-Agent Approach for Faithful Knowledge Reasoning via Promoting Causal Consistency in LLMs
di: Tang, Ziyi, et al.
Pubblicazione: (2023)
di: Tang, Ziyi, et al.
Pubblicazione: (2023)
AnimateZoo: Zero-shot Video Generation of Cross-Species Animation via Subject Alignment
di: Xu, Yuanfeng, et al.
Pubblicazione: (2024)
di: Xu, Yuanfeng, et al.
Pubblicazione: (2024)
AgriWorld:A World Tools Protocol Framework for Verifiable Agricultural Reasoning with Code-Executing LLM Agents
di: Zhang, Zhixing, et al.
Pubblicazione: (2026)
di: Zhang, Zhixing, et al.
Pubblicazione: (2026)
Robust Egocentric Referring Video Object Segmentation via Dual-Modal Causal Intervention
di: Liu, Haijing, et al.
Pubblicazione: (2025)
di: Liu, Haijing, et al.
Pubblicazione: (2025)
One Model for All: Unified Try-On and Try-Off in Any Pose via LLM-Inspired Bidirectional Tweedie Diffusion
di: Liu, Jinxi, et al.
Pubblicazione: (2025)
di: Liu, Jinxi, et al.
Pubblicazione: (2025)
MM-CoT:A Benchmark for Probing Visual Chain-of-Thought Reasoning in Multimodal Models
di: Zhang, Jusheng, et al.
Pubblicazione: (2025)
di: Zhang, Jusheng, et al.
Pubblicazione: (2025)
NeRF-VPT: Learning Novel View Representations with Neural Radiance Fields via View Prompt Tuning
di: Chen, Linsheng, et al.
Pubblicazione: (2024)
di: Chen, Linsheng, et al.
Pubblicazione: (2024)
UML-CoT: Structured Reasoning and Planning with Unified Modeling Language for Robotic Room Cleaning
di: Chen, Hongyu, et al.
Pubblicazione: (2025)
di: Chen, Hongyu, et al.
Pubblicazione: (2025)
DART: Dual Adaptive Refinement Transfer for Open-Vocabulary Multi-Label Recognition
di: Liu, Haijing, et al.
Pubblicazione: (2025)
di: Liu, Haijing, et al.
Pubblicazione: (2025)
Category-Adaptive Cross-Modal Semantic Refinement and Transfer for Open-Vocabulary Multi-Label Recognition
di: Liu, Haijing, et al.
Pubblicazione: (2024)
di: Liu, Haijing, et al.
Pubblicazione: (2024)
Heterogeneous Semantic Transfer for Multi-label Recognition with Partial Labels
di: Chen, Tianshui, et al.
Pubblicazione: (2022)
di: Chen, Tianshui, et al.
Pubblicazione: (2022)
WildVidFit: Video Virtual Try-On in the Wild via Image-Based Controlled Diffusion Models
di: He, Zijian, et al.
Pubblicazione: (2024)
di: He, Zijian, et al.
Pubblicazione: (2024)
Dynamic Correlation Learning and Regularization for Multi-Label Confidence Calibration
di: Chen, Tianshui, et al.
Pubblicazione: (2024)
di: Chen, Tianshui, et al.
Pubblicazione: (2024)
Adaptive Global-Local Representation Learning and Selection for Cross-Domain Facial Expression Recognition
di: Gao, Yuefang, et al.
Pubblicazione: (2024)
di: Gao, Yuefang, et al.
Pubblicazione: (2024)
RADAR: Defending RAG Dynamically against Retrieval Corruption
di: Chen, Ziyuan, et al.
Pubblicazione: (2026)
di: Chen, Ziyuan, et al.
Pubblicazione: (2026)
Towards Physics-informed Spatial Intelligence with Human Priors: An Autonomous Driving Pilot Study
di: Wu, Guanlin, et al.
Pubblicazione: (2025)
di: Wu, Guanlin, et al.
Pubblicazione: (2025)
STORM: Search-Guided Generative World Models for Robotic Manipulation
di: Lin, Wenjun, et al.
Pubblicazione: (2025)
di: Lin, Wenjun, et al.
Pubblicazione: (2025)
DriveWorld-VLA: Unified Latent-Space World Modeling with Vision-Language-Action for Autonomous Driving
di: jia, Feiyang, et al.
Pubblicazione: (2026)
di: jia, Feiyang, et al.
Pubblicazione: (2026)
MPDrive: Improving Spatial Understanding with Marker-Based Prompt Learning for Autonomous Driving
di: Zhang, Zhiyuan, et al.
Pubblicazione: (2025)
di: Zhang, Zhiyuan, et al.
Pubblicazione: (2025)
Exploring Talking Head Models With Adjacent Frame Prior for Speech-Preserving Facial Expression Manipulation
di: Lu, Zhenxuan, et al.
Pubblicazione: (2026)
di: Lu, Zhenxuan, et al.
Pubblicazione: (2026)
HiVA: Self-organized Hierarchical Variable Agent via Goal-driven Semantic-Topological Evolution
di: Tang, Jinzhou, et al.
Pubblicazione: (2025)
di: Tang, Jinzhou, et al.
Pubblicazione: (2025)
Language-free Compositional Action Generation via Decoupling Refinement
di: Liu, Xiao, et al.
Pubblicazione: (2023)
di: Liu, Xiao, et al.
Pubblicazione: (2023)
PanoWorld: Towards Spatial Supersensing in 360$^\circ$ Panorama World
di: Wang, Changpeng, et al.
Pubblicazione: (2026)
di: Wang, Changpeng, et al.
Pubblicazione: (2026)
Documenti analoghi
-
Stable Language Guidance for Vision-Language-Action Models
di: Zhan, Zhihao, et al.
Pubblicazione: (2026) -
TAG: Target-Agnostic Guidance for Stable Object-Centric Inference in Vision-Language-Action Models
di: Zhou, Jiaying, et al.
Pubblicazione: (2026) -
Physical Autoregressive Model for Robotic Manipulation without Action Pretraining
di: Song, Zijian, et al.
Pubblicazione: (2025) -
Learning Physics from Pretrained Video Models: A Multimodal Continuous and Sequential World Interaction Models for Robotic Manipulation
di: Song, Zijian, et al.
Pubblicazione: (2026) -
E0: Enhancing Generalization and Fine-Grained Control in VLA Models via Tweedie Discrete Diffusion
di: Zhan, Zhihao, et al.
Pubblicazione: (2025)