Action Dubber: Timing Audible Actions via Inflectional Flow
Fuente:
arXiv
Guardado en:
| Autores principales: | Wan, Wenlong, Zheng, Weiying, Xiang, Tianyi, Li, Guiqing, He, Shengfeng |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Box2Flow: Instance-based Action Flow Graphs from Videos
por: Li, Jiatong, et al.
Publicado: (2024)
por: Li, Jiatong, et al.
Publicado: (2024)
EITNet: An IoT-Enhanced Framework for Real-Time Basketball Action Recognition
por: Liu, Jingyu, et al.
Publicado: (2024)
por: Liu, Jingyu, et al.
Publicado: (2024)
EaqVLA: Encoding-aligned Quantization for Vision-Language-Action Models
por: Jiang, Feng, et al.
Publicado: (2025)
por: Jiang, Feng, et al.
Publicado: (2025)
Taylor Videos for Action Recognition
por: Wang, Lei, et al.
Publicado: (2024)
por: Wang, Lei, et al.
Publicado: (2024)
Spanning Training Progress: Temporal Dual-Depth Scoring (TDDS) for Enhanced Dataset Pruning
por: Zhang, Xin, et al.
Publicado: (2023)
por: Zhang, Xin, et al.
Publicado: (2023)
Unfolding 3D Gaussian Splatting via Iterative Gaussian Synopsis
por: Lu, Yuqin, et al.
Publicado: (2026)
por: Lu, Yuqin, et al.
Publicado: (2026)
HFGCN:Hypergraph Fusion Graph Convolutional Networks for Skeleton-Based Action Recognition
por: Dong, Pengcheng, et al.
Publicado: (2025)
por: Dong, Pengcheng, et al.
Publicado: (2025)
Multi-State-Action Tokenisation in Decision Transformers for Multi-Discrete Action Spaces
por: Moodley, Perusha, et al.
Publicado: (2024)
por: Moodley, Perusha, et al.
Publicado: (2024)
Spatio-Temporal LLM: Reasoning about Environments and Actions
por: Zheng, Haozhen, et al.
Publicado: (2025)
por: Zheng, Haozhen, et al.
Publicado: (2025)
RALACs: Action Recognition in Autonomous Vehicles using Interaction Encoding and Optical Flow
por: Zhou, Eddy, et al.
Publicado: (2022)
por: Zhou, Eddy, et al.
Publicado: (2022)
Benchmarking Vision, Language, & Action Models in Procedurally Generated, Open Ended Action Environments
por: Guruprasad, Pranav, et al.
Publicado: (2025)
por: Guruprasad, Pranav, et al.
Publicado: (2025)
AD3: Implicit Action is the Key for World Models to Distinguish the Diverse Visual Distractors
por: Wang, Yucen, et al.
Publicado: (2024)
por: Wang, Yucen, et al.
Publicado: (2024)
Selective, Interpretable, and Motion Consistent Privacy Attribute Obfuscation for Action Recognition
por: Ilic, Filip, et al.
Publicado: (2024)
por: Ilic, Filip, et al.
Publicado: (2024)
Video RWKV:Video Action Recognition Based RWKV
por: Yin, Zhuowen, et al.
Publicado: (2024)
por: Yin, Zhuowen, et al.
Publicado: (2024)
FlowHijack: A Dynamics-Aware Backdoor Attack on Flow-Matching Vision-Language-Action Models
por: An, Xinyuan, et al.
Publicado: (2026)
por: An, Xinyuan, et al.
Publicado: (2026)
NinA: Normalizing Flows in Action. Training VLA Models with Normalizing Flows
por: Tarasov, Denis, et al.
Publicado: (2025)
por: Tarasov, Denis, et al.
Publicado: (2025)
Action-slot: Visual Action-centric Representations for Multi-label Atomic Activity Recognition in Traffic Scenes
por: Kung, Chi-Hsi, et al.
Publicado: (2023)
por: Kung, Chi-Hsi, et al.
Publicado: (2023)
Real-Time Human Action Recognition on Embedded Platforms
por: Wang, Ruiqi, et al.
Publicado: (2024)
por: Wang, Ruiqi, et al.
Publicado: (2024)
Detecting Informative Channels: ActionFormer
por: Zhao, Kunpeng, et al.
Publicado: (2025)
por: Zhao, Kunpeng, et al.
Publicado: (2025)
Human Action Anticipation: A Survey
por: Lai, Bolin, et al.
Publicado: (2024)
por: Lai, Bolin, et al.
Publicado: (2024)
SA-DVAE: Improving Zero-Shot Skeleton-Based Action Recognition by Disentangled Variational Autoencoders
por: Li, Sheng-Wei, et al.
Publicado: (2024)
por: Li, Sheng-Wei, et al.
Publicado: (2024)
Prototypical Calibrating Ambiguous Samples for Micro-Action Recognition
por: Li, Kun, et al.
Publicado: (2024)
por: Li, Kun, et al.
Publicado: (2024)
World Action Models are Zero-shot Policies
por: Ye, Seonghyeon, et al.
Publicado: (2026)
por: Ye, Seonghyeon, et al.
Publicado: (2026)
Test-Time Training for Visual Foresight Vision-Language-Action Models
por: Park, Sangwu, et al.
Publicado: (2026)
por: Park, Sangwu, et al.
Publicado: (2026)
Discrete Diffusion VLA: Bringing Discrete Diffusion to Action Decoding in Vision-Language-Action Policies
por: Liang, Zhixuan, et al.
Publicado: (2025)
por: Liang, Zhixuan, et al.
Publicado: (2025)
Improving Vision-Language-Action Model with Online Reinforcement Learning
por: Guo, Yanjiang, et al.
Publicado: (2025)
por: Guo, Yanjiang, et al.
Publicado: (2025)
Multi-level and Multi-modal Action Anticipation
por: Kim, Seulgi, et al.
Publicado: (2025)
por: Kim, Seulgi, et al.
Publicado: (2025)
When Spatial meets Temporal in Action Recognition
por: Chen, Huilin, et al.
Publicado: (2024)
por: Chen, Huilin, et al.
Publicado: (2024)
Classification of Tennis Actions Using Deep Learning
por: Hovad, Emil, et al.
Publicado: (2024)
por: Hovad, Emil, et al.
Publicado: (2024)
About Time: Advances, Challenges, and Outlooks of Action Understanding
por: Stergiou, Alexandros, et al.
Publicado: (2024)
por: Stergiou, Alexandros, et al.
Publicado: (2024)
Group Relative Augmentation for Data Efficient Action Detection
por: Patel, Deep Anil, et al.
Publicado: (2025)
por: Patel, Deep Anil, et al.
Publicado: (2025)
Action-Agnostic Point-Level Supervision for Temporal Action Detection
por: Yoshida, Shuhei M., et al.
Publicado: (2024)
por: Yoshida, Shuhei M., et al.
Publicado: (2024)
Live and Learn: Continual Action Clustering with Incremental Views
por: Yan, Xiaoqiang, et al.
Publicado: (2024)
por: Yan, Xiaoqiang, et al.
Publicado: (2024)
On the Utility of 3D Hand Poses for Action Recognition
por: Shamil, Md Salman, et al.
Publicado: (2024)
por: Shamil, Md Salman, et al.
Publicado: (2024)
EZ-CLIP: Efficient Zeroshot Video Action Recognition
por: Ahmad, Shahzad, et al.
Publicado: (2023)
por: Ahmad, Shahzad, et al.
Publicado: (2023)
Motus: A Unified Latent Action World Model
por: Bi, Hongzhe, et al.
Publicado: (2025)
por: Bi, Hongzhe, et al.
Publicado: (2025)
Universal Pose Pretraining for Generalizable Vision-Language-Action Policies
por: Lin, Haitao, et al.
Publicado: (2026)
por: Lin, Haitao, et al.
Publicado: (2026)
Dense Policy: Bidirectional Autoregressive Learning of Actions
por: Su, Yue, et al.
Publicado: (2025)
por: Su, Yue, et al.
Publicado: (2025)
FedVLMBench: Benchmarking Federated Fine-Tuning of Vision-Language Models
por: Zheng, Weiying, et al.
Publicado: (2025)
por: Zheng, Weiying, et al.
Publicado: (2025)
Modular Retrieval-Augmented Generalization for Human Action Recognition
por: Liao, Peng, et al.
Publicado: (2026)
por: Liao, Peng, et al.
Publicado: (2026)
Ejemplares similares
-
Box2Flow: Instance-based Action Flow Graphs from Videos
por: Li, Jiatong, et al.
Publicado: (2024) -
EITNet: An IoT-Enhanced Framework for Real-Time Basketball Action Recognition
por: Liu, Jingyu, et al.
Publicado: (2024) -
EaqVLA: Encoding-aligned Quantization for Vision-Language-Action Models
por: Jiang, Feng, et al.
Publicado: (2025) -
Taylor Videos for Action Recognition
por: Wang, Lei, et al.
Publicado: (2024) -
Spanning Training Progress: Temporal Dual-Depth Scoring (TDDS) for Enhanced Dataset Pruning
por: Zhang, Xin, et al.
Publicado: (2023)