NoTVLA: Semantics-Preserving Robot Adaptation via Narrative Action Interfaces
Fuente:
arXiv
Saved in:
| Main Authors: | Huang, Zheng, Liu, Mingyu, Lin, Xiaoyi, Zhu, Muzhi, Zhao, Canyu, Du, Zongze, Lin, Ye, Li, Xiaoman, Jia, Yiduo, Zhong, Hao, Chen, Hao, Shen, Chunhua |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Bridge Thinking and Acting: Unleashing Physical Potential of VLM with Generalizable Action Expert
by: Liu, Mingyu, et al.
Published: (2025)
by: Liu, Mingyu, et al.
Published: (2025)
Omni-R1: Reinforcement Learning for Omnimodal Reasoning via Two-System Collaboration
by: Zhong, Hao, et al.
Published: (2025)
by: Zhong, Hao, et al.
Published: (2025)
Active-O3: Empowering Multimodal Large Language Models with Active Perception via GRPO
by: Zhu, Muzhi, et al.
Published: (2025)
by: Zhu, Muzhi, et al.
Published: (2025)
Preserving Source Video Realism: High-Fidelity Face Swapping for Cinematic Quality
by: Luo, Zekai, et al.
Published: (2025)
by: Luo, Zekai, et al.
Published: (2025)
OmniJigsaw: Enhancing Omni-Modal Reasoning via Modality-Orchestrated Reordering
by: Jia, Yiduo, et al.
Published: (2026)
by: Jia, Yiduo, et al.
Published: (2026)
DICEPTION: A Generalist Diffusion Model for Visual Perceptual Tasks
by: Zhao, Canyu, et al.
Published: (2025)
by: Zhao, Canyu, et al.
Published: (2025)
Tinker: Diffusion's Gift to 3D--Multi-View Consistent Editing From Sparse Inputs without Per-Scene Optimization
by: Zhao, Canyu, et al.
Published: (2025)
by: Zhao, Canyu, et al.
Published: (2025)
MoTVLA: A Vision-Language-Action Model with Unified Fast-Slow Reasoning
by: Huang, Wenhui, et al.
Published: (2025)
by: Huang, Wenhui, et al.
Published: (2025)
StaMo: Unsupervised Learning of Generalizable Robot Motion from Compact State Representation
by: Liu, Mingyu, et al.
Published: (2025)
by: Liu, Mingyu, et al.
Published: (2025)
A Three‐Level Meta‐Analysis on the Relation of Overweight or Obesity to Bullying Behavior Among Youths
by: Hao Chen, et al.
Published: (2025)
by: Hao Chen, et al.
Published: (2025)
Unleashing the Potential of the Diffusion Model in Few-shot Semantic Segmentation
by: Zhu, Muzhi, et al.
Published: (2024)
by: Zhu, Muzhi, et al.
Published: (2024)
MovieDreamer: Hierarchical Generation for Coherent Long Visual Sequence
by: Zhao, Canyu, et al.
Published: (2024)
by: Zhao, Canyu, et al.
Published: (2024)
MARBLE: Multi-Aspect Reward Balance for Diffusion RL
by: Zhao, Canyu, et al.
Published: (2026)
by: Zhao, Canyu, et al.
Published: (2026)
Matcher: Segment Anything with One Shot Using All-Purpose Feature Matching
by: Liu, Yang, et al.
Published: (2023)
by: Liu, Yang, et al.
Published: (2023)
Exploring Spatial Intelligence from a Generative Perspective
by: Zhu, Muzhi, et al.
Published: (2026)
by: Zhu, Muzhi, et al.
Published: (2026)
Where to Look: Can Foundation Models Reach a Target Viewpoint Through Active Exploration?
by: Li, Liyang, et al.
Published: (2026)
by: Li, Liyang, et al.
Published: (2026)
MMControl: Unified Multi-Modal Control for Joint Audio-Video Generation
by: Li, Liyang, et al.
Published: (2026)
by: Li, Liyang, et al.
Published: (2026)
FreeCustom: Tuning-Free Customized Image Generation for Multi-Concept Composition
by: Ding, Ganggui, et al.
Published: (2024)
by: Ding, Ganggui, et al.
Published: (2024)
DiverGen: Improving Instance Segmentation by Learning Wider Data Distribution with More Diverse Generative Data
by: Fan, Chengxiang, et al.
Published: (2024)
by: Fan, Chengxiang, et al.
Published: (2024)
Generative Active Learning for Long-tailed Instance Segmentation
by: Zhu, Muzhi, et al.
Published: (2024)
by: Zhu, Muzhi, et al.
Published: (2024)
A Simple Image Segmentation Framework via In-Context Examples
by: Liu, Yang, et al.
Published: (2024)
by: Liu, Yang, et al.
Published: (2024)
SurfaceSplat: Connecting Surface Reconstruction and Gaussian Splatting
by: Gao, Zihui, et al.
Published: (2025)
by: Gao, Zihui, et al.
Published: (2025)
Evo-1: Lightweight Vision-Language-Action Model with Preserved Semantic Alignment
by: Lin, Tao, et al.
Published: (2025)
by: Lin, Tao, et al.
Published: (2025)
Scalable Adaptation of 3D Geometric Foundation Models via Weak Supervision from Internet Video
by: Gao, Zihui, et al.
Published: (2026)
by: Gao, Zihui, et al.
Published: (2026)
SegAgent: Exploring Pixel Understanding Capabilities in MLLMs by Imitating Human Annotator Trajectories
by: Zhu, Muzhi, et al.
Published: (2025)
by: Zhu, Muzhi, et al.
Published: (2025)
Guiding Audio Editing with Audio Language Model
by: Lan, Zitong, et al.
Published: (2025)
by: Lan, Zitong, et al.
Published: (2025)
Resounding Acoustic Fields with Reciprocity
by: Lan, Zitong, et al.
Published: (2025)
by: Lan, Zitong, et al.
Published: (2025)
Unified Open-World Segmentation with Multi-Modal Prompts
by: Liu, Yang, et al.
Published: (2025)
by: Liu, Yang, et al.
Published: (2025)
GUI-Shepherd: Reliable Process Reward and Verification for Long-Sequence GUI Tasks
by: Chen, Cong, et al.
Published: (2025)
by: Chen, Cong, et al.
Published: (2025)
Higher-order multiscale method and its convergence analysis for nonlinear thermo-electric coupling problems of composite structures
by: Dong, Hao, et al.
Published: (2025)
by: Dong, Hao, et al.
Published: (2025)
Beyond TVLA: Anderson-Darling Leakage Assessment for Neural Network Side-Channel Leakage Detection
by: Mikulec, Ján, et al.
Published: (2026)
by: Mikulec, Ján, et al.
Published: (2026)
HieraTok: Multi-Scale Visual Tokenizer Improves Image Reconstruction and Generation
by: Chen, Cong, et al.
Published: (2025)
by: Chen, Cong, et al.
Published: (2025)
Towards Completeness: A Generalizable Action Proposal Generator for Zero-Shot Temporal Action Localization
by: Du, Jia-Run, et al.
Published: (2024)
by: Du, Jia-Run, et al.
Published: (2024)
TableDreamer: Progressive and Weakness-guided Data Synthesis from Scratch for Table Instruction Tuning
by: Zheng, Mingyu, et al.
Published: (2025)
by: Zheng, Mingyu, et al.
Published: (2025)
GUI Action Narrator: Where and When Did That Action Take Place?
by: Wu, Qinchen, et al.
Published: (2024)
by: Wu, Qinchen, et al.
Published: (2024)
Keeping the Evidence Chain: Semantic Evidence Allocation for Training-Free Token Pruning in Video Temporal Grounding
by: Li, Jiaqi, et al.
Published: (2026)
by: Li, Jiaqi, et al.
Published: (2026)
SPAR: Support-Preserving Action Rectification
by: Zhao, Jiaxin, et al.
Published: (2026)
by: Zhao, Jiaxin, et al.
Published: (2026)
FocusUI: Efficient UI Grounding via Position-Preserving Visual Token Selection
by: Ouyang, Mingyu, et al.
Published: (2026)
by: Ouyang, Mingyu, et al.
Published: (2026)
Polyphony: Diffusion-based Dual-Hand Action Segmentation with Alternating Vision Transformer and Semantic Conditioning
by: Zheng, Hao, et al.
Published: (2026)
by: Zheng, Hao, et al.
Published: (2026)
A Diachronic Investigation of the Change in Form and Formational‐Semantic Systematicity of the Chinese Sign Language Lexicon
by: Yue Zou, et al.
Published: (2025)
by: Yue Zou, et al.
Published: (2025)
Similar Items
-
Bridge Thinking and Acting: Unleashing Physical Potential of VLM with Generalizable Action Expert
by: Liu, Mingyu, et al.
Published: (2025) -
Omni-R1: Reinforcement Learning for Omnimodal Reasoning via Two-System Collaboration
by: Zhong, Hao, et al.
Published: (2025) -
Active-O3: Empowering Multimodal Large Language Models with Active Perception via GRPO
by: Zhu, Muzhi, et al.
Published: (2025) -
Preserving Source Video Realism: High-Fidelity Face Swapping for Cinematic Quality
by: Luo, Zekai, et al.
Published: (2025) -
OmniJigsaw: Enhancing Omni-Modal Reasoning via Modality-Orchestrated Reordering
by: Jia, Yiduo, et al.
Published: (2026)