VTAM: Video-Tactile-Action Models for Complex Physical Interaction Beyond VLAs
Fuente:
arXiv
Guardado en:
| Autores principales: | Yuan, Haoran, Yi, Weigang, Zhang, Zhenyu, Chen, Wendi, Mo, Yuchen, Yin, Jiashi, Li, Xinzhuo, Zeng, Xiangyu, Wen, Chuan, Lu, Cewu, Driggs-Campbell, Katherine, Lourentzou, Ismini |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Commonsense for Zero-Shot Natural Language Video Localization
por: Holla, Meghana, et al.
Publicado: (2023)
por: Holla, Meghana, et al.
Publicado: (2023)
CoRe3D: Collaborative Reasoning as a Foundation for 3D Intelligence
por: Yu, Tianjiao, et al.
Publicado: (2025)
por: Yu, Tianjiao, et al.
Publicado: (2025)
DreamPartGen: Semantically Grounded Part-Level 3D Generation via Collaborative Latent Denoising
por: Yu, Tianjiao, et al.
Publicado: (2026)
por: Yu, Tianjiao, et al.
Publicado: (2026)
Tactile-Based Human Intent Recognition for Robot Assistive Navigation
por: Peng, Shaoting, et al.
Publicado: (2025)
por: Peng, Shaoting, et al.
Publicado: (2025)
Phantom: Physics-Infused Video Generation via Joint Modeling of Visual and Latent Physical Dynamics
por: Shen, Ying, et al.
Publicado: (2026)
por: Shen, Ying, et al.
Publicado: (2026)
Hierarchical Dataset Selection for High-Quality Data Sharing
por: Zhou, Xiaona, et al.
Publicado: (2025)
por: Zhou, Xiaona, et al.
Publicado: (2025)
FAIR: Facilitating Artificial Intelligence Resilience in Manufacturing Industrial Internet
por: Zeng, Yingyan, et al.
Publicado: (2025)
por: Zeng, Yingyan, et al.
Publicado: (2025)
3D-VCD: Hallucination Mitigation in 3D-LLM Embodied Agents through Visual Contrastive Decoding
por: Ogunleye, Makanjuola, et al.
Publicado: (2026)
por: Ogunleye, Makanjuola, et al.
Publicado: (2026)
mTSBench: Benchmarking Multivariate Time Series Anomaly Detection and Model Selection at Scale
por: Zhou, Xiaona, et al.
Publicado: (2025)
por: Zhou, Xiaona, et al.
Publicado: (2025)
mimic-video: Video-Action Models for Generalizable Robot Control Beyond VLAs
por: Pai, Jonas, et al.
Publicado: (2025)
por: Pai, Jonas, et al.
Publicado: (2025)
Do You Know the Way? Human-in-the-Loop Understanding for Fast Traversability Estimation in Mobile Robotics
por: Schreiber, Andre, et al.
Publicado: (2025)
por: Schreiber, Andre, et al.
Publicado: (2025)
A Brief Survey on Leveraging Large Scale Vision Models for Enhanced Robot Grasping
por: Kamboj, Abhi, et al.
Publicado: (2024)
por: Kamboj, Abhi, et al.
Publicado: (2024)
RAVEL: Rare Concept Generation and Editing via Graph-driven Relational Guidance
por: Venkatesh, Kavana, et al.
Publicado: (2024)
por: Venkatesh, Kavana, et al.
Publicado: (2024)
ELBA: Learning by Asking for Embodied Visual Navigation and Task Completion
por: Shen, Ying, et al.
Publicado: (2023)
por: Shen, Ying, et al.
Publicado: (2023)
An Expert Ensemble for Detecting Anomalous Scenes, Interactions, and Behaviors in Autonomous Driving
por: Ji, Tianchen, et al.
Publicado: (2025)
por: Ji, Tianchen, et al.
Publicado: (2025)
Counterfactual Segmentation Reasoning: Diagnosing and Mitigating Pixel-Grounding Hallucination
por: Li, Xinzhuo, et al.
Publicado: (2025)
por: Li, Xinzhuo, et al.
Publicado: (2025)
Uncertainty in Action: Confidence Elicitation in Embodied Agents
por: Yu, Tianjiao, et al.
Publicado: (2025)
por: Yu, Tianjiao, et al.
Publicado: (2025)
PRIMA: Multi-Image Vision-Language Models for Reasoning Segmentation
por: Wahed, Muntasir, et al.
Publicado: (2024)
por: Wahed, Muntasir, et al.
Publicado: (2024)
Beyond the Dashboard: Investigating Distracted Driver Communication Preferences for ADAS
por: Hasan, Aamir, et al.
Publicado: (2024)
por: Hasan, Aamir, et al.
Publicado: (2024)
Trust-Aware Embodied Bayesian Persuasion for Mixed-Autonomy
por: Peng, Shaoting, et al.
Publicado: (2025)
por: Peng, Shaoting, et al.
Publicado: (2025)
Towards Provable Log Density Policy Gradient
por: Katdare, Pulkit, et al.
Publicado: (2024)
por: Katdare, Pulkit, et al.
Publicado: (2024)
Towards Uncertainty Unification: A Case Study for Preference Learning
por: Peng, Shaoting, et al.
Publicado: (2025)
por: Peng, Shaoting, et al.
Publicado: (2025)
Hallucination Detection in Foundation Models for Decision-Making: A Flexible Definition and Review of the State of the Art
por: Chakraborty, Neeloy, et al.
Publicado: (2024)
por: Chakraborty, Neeloy, et al.
Publicado: (2024)
Evaluating Cognitive Age Alignment in Interactive AI Agents
por: Shen, Yifan, et al.
Publicado: (2026)
por: Shen, Yifan, et al.
Publicado: (2026)
MMPlanner: Zero-Shot Multimodal Procedural Planning with Chain-of-Thought Object State Reasoning
por: Tabassum, Afrina, et al.
Publicado: (2025)
por: Tabassum, Afrina, et al.
Publicado: (2025)
uaMix-MAE: Efficient Tuning of Pretrained Audio Transformers with Unsupervised Audio Mixtures
por: Tabassum, Afrina, et al.
Publicado: (2024)
por: Tabassum, Afrina, et al.
Publicado: (2024)
Tiny but Trusted: Efficient Vision-Language Reasoning for Time-Series Anomaly Detection
por: Zhou, Xiaona, et al.
Publicado: (2026)
por: Zhou, Xiaona, et al.
Publicado: (2026)
Rethinking Gaussian Trajectory Predictors: Calibrated Uncertainty for Safe Planning
por: Pouria, Fatemeh Cheraghi, et al.
Publicado: (2026)
por: Pouria, Fatemeh Cheraghi, et al.
Publicado: (2026)
In-Situ Soil-Property Estimation and Bayesian Mapping with a Simulated Compact Track Loader
por: Wagner, W. Jacob, et al.
Publicado: (2025)
por: Wagner, W. Jacob, et al.
Publicado: (2025)
Human-Agent Joint Learning for Efficient Robot Manipulation Skill Acquisition
por: Luo, Shengcheng, et al.
Publicado: (2024)
por: Luo, Shengcheng, et al.
Publicado: (2024)
DeformPAM: Data-Efficient Learning for Long-horizon Deformable Object Manipulation via Preference-based Action Alignment
por: Chen, Wendi, et al.
Publicado: (2024)
por: Chen, Wendi, et al.
Publicado: (2024)
CALICO: Part-Focused Semantic Co-Segmentation with Large Vision-Language Models
por: Nguyen, Kiet A., et al.
Publicado: (2024)
por: Nguyen, Kiet A., et al.
Publicado: (2024)
PyraTok: Language-Aligned Pyramidal Tokenizer for Video Understanding and Generation
por: Susladkar, Onkar, et al.
Publicado: (2026)
por: Susladkar, Onkar, et al.
Publicado: (2026)
Interaction-aware Conformal Prediction for Crowd Navigation
por: Huang, Zhe, et al.
Publicado: (2025)
por: Huang, Zhe, et al.
Publicado: (2025)
Characterizing the Robustness of Black-Box LLM Planners Under Perturbed Observations with Adaptive Stress Testing
por: Chakraborty, Neeloy, et al.
Publicado: (2025)
por: Chakraborty, Neeloy, et al.
Publicado: (2025)
LIT: Large Language Model Driven Intention Tracking for Proactive Human-Robot Collaboration -- A Robot Sous-Chef Application
por: Huang, Zhe, et al.
Publicado: (2024)
por: Huang, Zhe, et al.
Publicado: (2024)
Structured Graph Network for Constrained Robot Crowd Navigation with Low Fidelity Simulation
por: Liu, Shuijing, et al.
Publicado: (2024)
por: Liu, Shuijing, et al.
Publicado: (2024)
Neural Informed RRT*: Learning-based Path Planning with Point Cloud State Representations under Admissible Ellipsoidal Constraints
por: Huang, Zhe, et al.
Publicado: (2023)
por: Huang, Zhe, et al.
Publicado: (2023)
Reactive Diffusion Policy: Slow-Fast Visual-Tactile Policy Learning for Contact-Rich Manipulation
por: Xue, Han, et al.
Publicado: (2025)
por: Xue, Han, et al.
Publicado: (2025)
Sensor-Invariant Tactile Representation
por: Gupta, Harsh, et al.
Publicado: (2025)
por: Gupta, Harsh, et al.
Publicado: (2025)
Ejemplares similares
-
Commonsense for Zero-Shot Natural Language Video Localization
por: Holla, Meghana, et al.
Publicado: (2023) -
CoRe3D: Collaborative Reasoning as a Foundation for 3D Intelligence
por: Yu, Tianjiao, et al.
Publicado: (2025) -
DreamPartGen: Semantically Grounded Part-Level 3D Generation via Collaborative Latent Denoising
por: Yu, Tianjiao, et al.
Publicado: (2026) -
Tactile-Based Human Intent Recognition for Robot Assistive Navigation
por: Peng, Shaoting, et al.
Publicado: (2025) -
Phantom: Physics-Infused Video Generation via Joint Modeling of Visual and Latent Physical Dynamics
por: Shen, Ying, et al.
Publicado: (2026)