ROCKET: Residual-Oriented Multi-Layer Alignment for Spatially-Aware Vision-Language-Action Models
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Sun, Guoheng, Du, Tingting, Feng, Kaixi, Luo, Chenxiang, Ding, Xingguo, Shen, Zheyu, Wang, Ziyao, He, Yexiao, Li, Ang |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Vision-Language-Action in Robotics: A Survey of Datasets, Benchmarks, and Data Engines
von: Wang, Ziyao, et al.
Veröffentlicht: (2026)
von: Wang, Ziyao, et al.
Veröffentlicht: (2026)
Predictive Auditing of Hidden Tokens in LLM APIs via Reasoning Length Estimation
von: Wang, Ziyao, et al.
Veröffentlicht: (2025)
von: Wang, Ziyao, et al.
Veröffentlicht: (2025)
EdgeLoRA: An Efficient Multi-Tenant LLM Serving System on Edge Devices
von: Shen, Zheyu, et al.
Veröffentlicht: (2025)
von: Shen, Zheyu, et al.
Veröffentlicht: (2025)
FedMOA: Federated GRPO for Personalized Reasoning LLMs under Heterogeneous Rewards
von: Wang, Ziyao, et al.
Veröffentlicht: (2026)
von: Wang, Ziyao, et al.
Veröffentlicht: (2026)
Prada: Black-Box LLM Adaptation with Private Data on Resource-Constrained Devices
von: Wang, Ziyao, et al.
Veröffentlicht: (2025)
von: Wang, Ziyao, et al.
Veröffentlicht: (2025)
FLoRA: Federated Fine-Tuning Large Language Models with Heterogeneous Low-Rank Adaptations
von: Wang, Ziyao, et al.
Veröffentlicht: (2024)
von: Wang, Ziyao, et al.
Veröffentlicht: (2024)
Invisible Tokens, Visible Bills: The Urgent Need to Audit Hidden Operations in Opaque LLM Services
von: Sun, Guoheng, et al.
Veröffentlicht: (2025)
von: Sun, Guoheng, et al.
Veröffentlicht: (2025)
SHED: Shapley-Based Automated Dataset Refinement for Instruction Fine-Tuning
von: He, Yexiao, et al.
Veröffentlicht: (2024)
von: He, Yexiao, et al.
Veröffentlicht: (2024)
Revisiting Federated Fine-Tuning: A Single Communication Round is Enough for Foundation Models
von: Wang, Ziyao, et al.
Veröffentlicht: (2024)
von: Wang, Ziyao, et al.
Veröffentlicht: (2024)
CoIn: Counting the Invisible Reasoning Tokens in Commercial Opaque LLM APIs
von: Sun, Guoheng, et al.
Veröffentlicht: (2025)
von: Sun, Guoheng, et al.
Veröffentlicht: (2025)
What Matters in Transformers? Not All Attention is Needed
von: He, Shwai, et al.
Veröffentlicht: (2024)
von: He, Shwai, et al.
Veröffentlicht: (2024)
Fair Diagnosis: Leveraging Causal Modeling to Mitigate Medical Bias
von: Tian, Bowei, et al.
Veröffentlicht: (2024)
von: Tian, Bowei, et al.
Veröffentlicht: (2024)
NeuroSymAD: A Neuro-Symbolic Framework for Interpretable Alzheimer's Disease Diagnosis
von: He, Yexiao, et al.
Veröffentlicht: (2025)
von: He, Yexiao, et al.
Veröffentlicht: (2025)
SymRTLO: Enhancing RTL Code Optimization with LLMs and Neuron-Inspired Symbolic Reasoning
von: Wang, Yiting, et al.
Veröffentlicht: (2025)
von: Wang, Yiting, et al.
Veröffentlicht: (2025)
MindCraft: How Concept Trees Take Shape In Deep Models
von: Tian, Bowei, et al.
Veröffentlicht: (2025)
von: Tian, Bowei, et al.
Veröffentlicht: (2025)
HIT-ROCKET: Hadamard-vector Inner-product Transformer for ROCKET
von: Hao, Wang, et al.
Veröffentlicht: (2025)
von: Hao, Wang, et al.
Veröffentlicht: (2025)
CogniPair: From LLM Chatbots to Conscious AI Agents -- GNWT-Based Multi-Agent Digital Twins for Social Pairing -- Dating & Hiring Applications
von: Ye, Wanghao, et al.
Veröffentlicht: (2025)
von: Ye, Wanghao, et al.
Veröffentlicht: (2025)
Towards Building Non-Fine-Tunable Foundation Models
von: Wang, Ziyao, et al.
Veröffentlicht: (2026)
von: Wang, Ziyao, et al.
Veröffentlicht: (2026)
VEGA: Visual Encoder Grounding Alignment for Spatially-Aware Vision-Language-Action Models
von: Wang, Hao, et al.
Veröffentlicht: (2026)
von: Wang, Hao, et al.
Veröffentlicht: (2026)
EvoVerilog: Large Langugage Model Assisted Evolution of Verilog Code
von: Guo, Ping, et al.
Veröffentlicht: (2025)
von: Guo, Ping, et al.
Veröffentlicht: (2025)
MCP4EDA: LLM-Powered Model Context Protocol RTL-to-GDSII Automation with Backend Aware Synthesis Optimization
von: Wang, Yiting, et al.
Veröffentlicht: (2025)
von: Wang, Yiting, et al.
Veröffentlicht: (2025)
MedOrch: Medical Diagnosis with Tool-Augmented Reasoning Agents for Flexible Extensibility
von: He, Yexiao, et al.
Veröffentlicht: (2025)
von: He, Yexiao, et al.
Veröffentlicht: (2025)
ROCKET-2: Steering Visuomotor Policy via Cross-View Goal Alignment
von: Cai, Shaofei, et al.
Veröffentlicht: (2025)
von: Cai, Shaofei, et al.
Veröffentlicht: (2025)
CUROCKET: Optimizing ROCKET for GPU
von: Stüven, Ole, et al.
Veröffentlicht: (2026)
von: Stüven, Ole, et al.
Veröffentlicht: (2026)
Data-Efficient Spectral Classification of Hyperspectral Data Using MiniROCKET and HDC-MiniROCKET
von: Theisen, Nick, et al.
Veröffentlicht: (2025)
von: Theisen, Nick, et al.
Veröffentlicht: (2025)
Towards counterfactual fairness through auxiliary variables
von: Tian, Bowei, et al.
Veröffentlicht: (2024)
von: Tian, Bowei, et al.
Veröffentlicht: (2024)
Residual Supply and the Price of Risk Absorption
von: Wang, Ziyao
Veröffentlicht: (2026)
von: Wang, Ziyao
Veröffentlicht: (2026)
Arctic-Text2SQL-R1: Simple Rewards, Strong Reasoning in Text-to-SQL
von: Yao, Zhewei, et al.
Veröffentlicht: (2025)
von: Yao, Zhewei, et al.
Veröffentlicht: (2025)
Cognibit: From Digital Exhaustion to Real-World Connection Through Gamified Territory Control and LLM-Powered Twin Networking
von: Ye, Wanghao, et al.
Veröffentlicht: (2026)
von: Ye, Wanghao, et al.
Veröffentlicht: (2026)
Decentralized Time Series Classification with ROCKET Features
von: Casella, Bruno, et al.
Veröffentlicht: (2025)
von: Casella, Bruno, et al.
Veröffentlicht: (2025)
Global-local Spatial-temporal Aware Graph Attention Network for Network Traffic Forecasting
von: Xing, Jinming, et al.
Veröffentlicht: (2025)
von: Xing, Jinming, et al.
Veröffentlicht: (2025)
Demystifying When Pruning Works via Representation Hierarchies
von: He, Shwai, et al.
Veröffentlicht: (2026)
von: He, Shwai, et al.
Veröffentlicht: (2026)
Group Deliberation Oriented Multi-Agent Conversational Model for Complex Reasoning
von: Shi, Zheyu, et al.
Veröffentlicht: (2025)
von: Shi, Zheyu, et al.
Veröffentlicht: (2025)
Spatial Memory for Out-of-Vision Manipulation in Vision-Language-Action
von: Li, Pengteng, et al.
Veröffentlicht: (2026)
von: Li, Pengteng, et al.
Veröffentlicht: (2026)
Autoregressive Video Autoencoder with Decoupled Temporal and Spatial Context
von: Shen, Cuifeng, et al.
Veröffentlicht: (2025)
von: Shen, Cuifeng, et al.
Veröffentlicht: (2025)
ML-CLIPSim: Multi-Layer CLIP Similarity for Machine-Oriented Image Quality
von: Ding, Feng, et al.
Veröffentlicht: (2026)
von: Ding, Feng, et al.
Veröffentlicht: (2026)
InSpire: Vision-Language-Action Models with Intrinsic Spatial Reasoning
von: Zhang, Ji, et al.
Veröffentlicht: (2025)
von: Zhang, Ji, et al.
Veröffentlicht: (2025)
Self-Improving Vision-Language-Action Models with Data Generation via Residual RL
von: Xiao, Wenli, et al.
Veröffentlicht: (2025)
von: Xiao, Wenli, et al.
Veröffentlicht: (2025)
SPROCKET: Extending ROCKET to Distance-Based Time-Series Transformations With Prototypes
von: Harner, Nicholas
Veröffentlicht: (2025)
von: Harner, Nicholas
Veröffentlicht: (2025)
Spatial Forcing: Implicit Spatial Representation Alignment for Vision-language-action Model
von: Li, Fuhao, et al.
Veröffentlicht: (2025)
von: Li, Fuhao, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Vision-Language-Action in Robotics: A Survey of Datasets, Benchmarks, and Data Engines
von: Wang, Ziyao, et al.
Veröffentlicht: (2026) -
Predictive Auditing of Hidden Tokens in LLM APIs via Reasoning Length Estimation
von: Wang, Ziyao, et al.
Veröffentlicht: (2025) -
EdgeLoRA: An Efficient Multi-Tenant LLM Serving System on Edge Devices
von: Shen, Zheyu, et al.
Veröffentlicht: (2025) -
FedMOA: Federated GRPO for Personalized Reasoning LLMs under Heterogeneous Rewards
von: Wang, Ziyao, et al.
Veröffentlicht: (2026) -
Prada: Black-Box LLM Adaptation with Private Data on Resource-Constrained Devices
von: Wang, Ziyao, et al.
Veröffentlicht: (2025)