CrossVLA: Cross-Paradigm Post-Training and Inference Optimization for Vision-Language-Action Models
Fuente:
arXiv
Guardado en:
| Autor principal: | Liu, Zhi |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
ExpReS-VLA: Specializing Vision-Language-Action Models Through Experience Replay and Retrieval
por: Syed, Shahram Najam, et al.
Publicado: (2025)
por: Syed, Shahram Najam, et al.
Publicado: (2025)
Closed-Loop Neural Activation Control in Vision-Language-Action Models
por: Babu, Abhijith, et al.
Publicado: (2026)
por: Babu, Abhijith, et al.
Publicado: (2026)
VLA Foundry: A Unified Framework for Training Vision-Language-Action Models
por: Mercat, Jean, et al.
Publicado: (2026)
por: Mercat, Jean, et al.
Publicado: (2026)
VITA: Zero-Shot Value Functions via Test-Time Adaptation of Vision-Language Models
por: Ziakas, Christos, et al.
Publicado: (2025)
por: Ziakas, Christos, et al.
Publicado: (2025)
Dream to Fly: Model-Based Reinforcement Learning for Vision-Based Drone Flight
por: Romero, Angel, et al.
Publicado: (2025)
por: Romero, Angel, et al.
Publicado: (2025)
CLARE: Continual Learning for Vision-Language-Action Models via Autonomous Adapter Routing and Expansion
por: Römer, Ralf, et al.
Publicado: (2026)
por: Römer, Ralf, et al.
Publicado: (2026)
eStonefish-Scenes: A Sim-to-Real Validated and Robot-Centric Event-based Optical Flow Dataset for Underwater Vehicles
por: Mansour, Jad, et al.
Publicado: (2025)
por: Mansour, Jad, et al.
Publicado: (2025)
eCARLA-scenes: A synthetically generated dataset for event-based optical flow prediction
por: Mansour, Jad, et al.
Publicado: (2024)
por: Mansour, Jad, et al.
Publicado: (2024)
WayFASTER: a Self-Supervised Traversability Prediction for Increased Navigation Awareness
por: Gasparino, Mateus Valverde, et al.
Publicado: (2024)
por: Gasparino, Mateus Valverde, et al.
Publicado: (2024)
CulinaryCut-VLAP: A Vision-Language-Action-Physics Framework for Food Cutting via a Force-Aware Material Point Method
por: Koh, Hyunseo, et al.
Publicado: (2026)
por: Koh, Hyunseo, et al.
Publicado: (2026)
Pointing-Guided Target Estimation via Transformer-Based Attention
por: Müller, Luca, et al.
Publicado: (2025)
por: Müller, Luca, et al.
Publicado: (2025)
HumanEgo: Zero-Shot Robot Learning from Minutes of Human Egocentric Videos
por: Wang, Zhi, et al.
Publicado: (2026)
por: Wang, Zhi, et al.
Publicado: (2026)
Learning from Watching: Scalable Extraction of Manipulation Trajectories from Human Videos
por: Hu, X., et al.
Publicado: (2025)
por: Hu, X., et al.
Publicado: (2025)
Cortex 2.0: Grounding World Models in Real-World Industrial Deployment
por: Aida, Adriana, et al.
Publicado: (2026)
por: Aida, Adriana, et al.
Publicado: (2026)
Exposing Blindspots: Cultural Bias Evaluation in Generative Image Models
por: Seo, Huichan, et al.
Publicado: (2025)
por: Seo, Huichan, et al.
Publicado: (2025)
StemVLA:An Open-Source Vision-Language-Action Model with Future 3D Spatial Geometry Knowledge and 4D Historical Representation
por: Xiao, Jiasong, et al.
Publicado: (2026)
por: Xiao, Jiasong, et al.
Publicado: (2026)
Contextual Graph Representations for Task-Driven 3D Perception and Planning
por: Agia, Christopher
Publicado: (2026)
por: Agia, Christopher
Publicado: (2026)
Curb Your Attention: Causal Attention Gating for Robust Trajectory Prediction in Autonomous Driving
por: Ahmadi, Ehsan, et al.
Publicado: (2024)
por: Ahmadi, Ehsan, et al.
Publicado: (2024)
The Impact of 2D Segmentation Backbones on Point Cloud Predictions Using 4D Radar
por: Muckelroy III, William, et al.
Publicado: (2025)
por: Muckelroy III, William, et al.
Publicado: (2025)
Optimizing Neurorobot Policy under Limited Demonstration Data through Preference Regret
por: Nguyen, Viet Dung, et al.
Publicado: (2026)
por: Nguyen, Viet Dung, et al.
Publicado: (2026)
From Demonstrations to Safe Deployment: Path-Consistent Safety Filtering for Diffusion Policies
por: Römer, Ralf, et al.
Publicado: (2025)
por: Römer, Ralf, et al.
Publicado: (2025)
Convolutional Model Trees
por: Armstrong, William Ward, et al.
Publicado: (2025)
por: Armstrong, William Ward, et al.
Publicado: (2025)
Goal-Based Vision-Language Driving
por: Patapati, Santosh, et al.
Publicado: (2025)
por: Patapati, Santosh, et al.
Publicado: (2025)
RoboPack: Learning Tactile-Informed Dynamics Models for Dense Packing
por: Ai, Bo, et al.
Publicado: (2024)
por: Ai, Bo, et al.
Publicado: (2024)
Large Language Models and 3D Vision for Intelligent Robotic Perception and Autonomy
por: Mehta, Vinit, et al.
Publicado: (2025)
por: Mehta, Vinit, et al.
Publicado: (2025)
Adaptive Self-Training for Object Detection
por: Vandeghen, Renaud, et al.
Publicado: (2022)
por: Vandeghen, Renaud, et al.
Publicado: (2022)
Ego-Motion Aware Target Prediction Module for Robust Multi-Object Tracking
por: Mahdian, Navid, et al.
Publicado: (2024)
por: Mahdian, Navid, et al.
Publicado: (2024)
Locate 3D: Real-World Object Localization via Self-Supervised Learning in 3D
por: Arnaud, Sergio, et al.
Publicado: (2025)
por: Arnaud, Sergio, et al.
Publicado: (2025)
Efficient Image Pre-Training with Siamese Cropped Masked Autoencoders
por: Eymaël, Alexandre, et al.
Publicado: (2024)
por: Eymaël, Alexandre, et al.
Publicado: (2024)
Bayesian Data Augmentation and Training for Perception DNN in Autonomous Aerial Vehicles
por: Rasul, Ashik E, et al.
Publicado: (2024)
por: Rasul, Ashik E, et al.
Publicado: (2024)
CoT4AD: A Vision-Language-Action Model with Explicit Chain-of-Thought Reasoning for Autonomous Driving
por: Wang, Zhaohui, et al.
Publicado: (2025)
por: Wang, Zhaohui, et al.
Publicado: (2025)
RDPO: Real Data Preference Optimization for Physics Consistency Video Generation
por: Qian, Wenxu, et al.
Publicado: (2025)
por: Qian, Wenxu, et al.
Publicado: (2025)
RAE-NWM: Navigation World Model in Dense Visual Representation Space
por: Zhang, Mingkun, et al.
Publicado: (2026)
por: Zhang, Mingkun, et al.
Publicado: (2026)
ATAAT: Adaptive Threat-Aware Adversarial Tuning Framework against Backdoor Attacks on Vision-Language-Action Models
por: Chen, Kewei, et al.
Publicado: (2026)
por: Chen, Kewei, et al.
Publicado: (2026)
MaSC: A Masked Similarity Metric for Evaluating Concept-Driven Generation
por: Bartkowiak, Patryk, et al.
Publicado: (2026)
por: Bartkowiak, Patryk, et al.
Publicado: (2026)
A Survey on Vision-Language-Action Models for Embodied AI
por: Ma, Yueen, et al.
Publicado: (2024)
por: Ma, Yueen, et al.
Publicado: (2024)
SpectraNet: FFT-assisted Deep Learning Classifier for Deepfake Face Detection
por: Jayarathne, Nithira, et al.
Publicado: (2025)
por: Jayarathne, Nithira, et al.
Publicado: (2025)
Human-Centric Anomaly Detection in Surveillance Videos Using YOLO-World and Spatio-Temporal Deep Learning
por: Naeen, Mohammad Ali Etemadi, et al.
Publicado: (2025)
por: Naeen, Mohammad Ali Etemadi, et al.
Publicado: (2025)
FT-NCFM: An Influence-Aware Data Distillation Framework for Efficient VLA Models
por: Chen, Kewei, et al.
Publicado: (2025)
por: Chen, Kewei, et al.
Publicado: (2025)
Towards Localizing Structural Elements: Merging Geometrical Detection with Semantic Verification in RGB-D Data
por: Tourani, Ali, et al.
Publicado: (2024)
por: Tourani, Ali, et al.
Publicado: (2024)
Ejemplares similares
-
ExpReS-VLA: Specializing Vision-Language-Action Models Through Experience Replay and Retrieval
por: Syed, Shahram Najam, et al.
Publicado: (2025) -
Closed-Loop Neural Activation Control in Vision-Language-Action Models
por: Babu, Abhijith, et al.
Publicado: (2026) -
VLA Foundry: A Unified Framework for Training Vision-Language-Action Models
por: Mercat, Jean, et al.
Publicado: (2026) -
VITA: Zero-Shot Value Functions via Test-Time Adaptation of Vision-Language Models
por: Ziakas, Christos, et al.
Publicado: (2025) -
Dream to Fly: Model-Based Reinforcement Learning for Vision-Based Drone Flight
por: Romero, Angel, et al.
Publicado: (2025)