USAD: End-to-End Human Activity Recognition via Diffusion Model with Spatiotemporal Attention
Fuente:
arXiv
Guardado en:
| Autores principales: | Xiao, Hang, Yu, Ying, Li, Jiarui, Yang, Zhifan, Tang, Haotian, Liu, Hanyu, Li, Chao |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Process Optimization and Deployment for Sensor-Based Human Activity Recognition Based on Deep Learning
por: Liu, Hanyu, et al.
Publicado: (2025)
por: Liu, Hanyu, et al.
Publicado: (2025)
CMD-HAR: Cross-Modal Disentanglement for Wearable Human Activity Recognition
por: Yu, Ying, et al.
Publicado: (2025)
por: Yu, Ying, et al.
Publicado: (2025)
Confidence-driven Gradient Modulation for Multimodal Human Activity Recognition: A Dynamic Contrastive Dual-Path Learning Approach
por: Ji, Panpan, et al.
Publicado: (2025)
por: Ji, Panpan, et al.
Publicado: (2025)
End-to-End Visual Autonomous Parking via Control-Aided Attention
por: Chen, Chao, et al.
Publicado: (2025)
por: Chen, Chao, et al.
Publicado: (2025)
Guiding Attention in End-to-End Driving Models
por: Porres, Diego, et al.
Publicado: (2024)
por: Porres, Diego, et al.
Publicado: (2024)
DiffusionDrive: Truncated Diffusion Model for End-to-End Autonomous Driving
por: Liao, Bencheng, et al.
Publicado: (2024)
por: Liao, Bencheng, et al.
Publicado: (2024)
From Category to Scenery: An End-to-End Framework for Multi-Person Human-Object Interaction Recognition in Videos
por: Qiao, Tanqiu, et al.
Publicado: (2024)
por: Qiao, Tanqiu, et al.
Publicado: (2024)
End-to-End Chess Recognition
por: Masouris, Athanasios, et al.
Publicado: (2023)
por: Masouris, Athanasios, et al.
Publicado: (2023)
End-to-End Training for Autoregressive Video Diffusion via Self-Resampling
por: Guo, Yuwei, et al.
Publicado: (2025)
por: Guo, Yuwei, et al.
Publicado: (2025)
CFVNet: An End-to-End Cancelable Finger Vein Network for Recognition
por: Wang, Yifan, et al.
Publicado: (2024)
por: Wang, Yifan, et al.
Publicado: (2024)
LPSNet: End-to-End Human Pose and Shape Estimation with Lensless Imaging
por: Ge, Haoyang, et al.
Publicado: (2024)
por: Ge, Haoyang, et al.
Publicado: (2024)
End-to-End 3D Spatiotemporal Perception with Multimodal Fusion and V2X Collaboration
por: Yang, Zhenwei, et al.
Publicado: (2025)
por: Yang, Zhenwei, et al.
Publicado: (2025)
End-to-End Driving with Online Trajectory Evaluation via BEV World Model
por: Li, Yingyan, et al.
Publicado: (2025)
por: Li, Yingyan, et al.
Publicado: (2025)
Diffusion As Self-Distillation: End-to-End Latent Diffusion In One Model
por: Wang, Xiyuan, et al.
Publicado: (2025)
por: Wang, Xiyuan, et al.
Publicado: (2025)
DeepSight: Long-Horizon World Modeling via Latent States Prediction for End-to-End Autonomous Driving
por: Zhang, Lingjun, et al.
Publicado: (2026)
por: Zhang, Lingjun, et al.
Publicado: (2026)
End-to-End Shared Attention Estimation via Group Detection with Feedback Refinement
por: Nakatani, Chihiro, et al.
Publicado: (2026)
por: Nakatani, Chihiro, et al.
Publicado: (2026)
An Effective End-to-End Solution for Multimodal Action Recognition
por: Wang, Songping, et al.
Publicado: (2025)
por: Wang, Songping, et al.
Publicado: (2025)
Exploring Disentangled and Controllable Human Image Synthesis: From End-to-End to Stage-by-Stage
por: Sun, Zhengwentai, et al.
Publicado: (2025)
por: Sun, Zhengwentai, et al.
Publicado: (2025)
An End-to-End Robust Point Cloud Semantic Segmentation Network with Single-Step Conditional Diffusion Models
por: Qu, Wentao, et al.
Publicado: (2024)
por: Qu, Wentao, et al.
Publicado: (2024)
SoGAR: Self-supervised Spatiotemporal Attention-based Social Group Activity Recognition
por: Chappa, Naga VS Raviteja, et al.
Publicado: (2023)
por: Chappa, Naga VS Raviteja, et al.
Publicado: (2023)
Conditional Diffusion Model with Anatomical-Dose Dual Constraints for End-to-End Multi-Tumor Dose Prediction
por: Xie, Hui, et al.
Publicado: (2025)
por: Xie, Hui, et al.
Publicado: (2025)
VividAnimator: An End-to-End Audio and Pose-driven Half-Body Human Animation Framework
por: Huang, Donglin, et al.
Publicado: (2025)
por: Huang, Donglin, et al.
Publicado: (2025)
Leverage Cross-Attention for End-to-End Open-Vocabulary Panoptic Reconstruction
por: Yu, Xuan, et al.
Publicado: (2025)
por: Yu, Xuan, et al.
Publicado: (2025)
FutureX: Enhance End-to-End Autonomous Driving via Latent Chain-of-Thought World Model
por: Lin, Hongbin, et al.
Publicado: (2025)
por: Lin, Hongbin, et al.
Publicado: (2025)
Hydra-MDP++: Advancing End-to-End Driving via Expert-Guided Hydra-Distillation
por: Li, Kailin, et al.
Publicado: (2025)
por: Li, Kailin, et al.
Publicado: (2025)
Active Learning from Scene Embeddings for End-to-End Autonomous Driving
por: Jiang, Wenhao, et al.
Publicado: (2025)
por: Jiang, Wenhao, et al.
Publicado: (2025)
FoundationSLAM: Unleashing the Power of Depth Foundation Models for End-to-End Dense Visual SLAM
por: Wu, Yuchen, et al.
Publicado: (2025)
por: Wu, Yuchen, et al.
Publicado: (2025)
E2E-GNet: An End-to-End Skeleton-based Geometric Deep Neural Network for Human Motion Recognition
por: Olaoluwa, Mubarak, et al.
Publicado: (2026)
por: Olaoluwa, Mubarak, et al.
Publicado: (2026)
NudgeVAD: Language-Nudged End-to-End Driving via FiLM Residuals
por: Yang, Chieh-Chi, et al.
Publicado: (2026)
por: Yang, Chieh-Chi, et al.
Publicado: (2026)
An End-to-End, Segmentation-Free, Arabic Handwritten Recognition Model on KHATT
por: Aabed, Sondos, et al.
Publicado: (2024)
por: Aabed, Sondos, et al.
Publicado: (2024)
Towards End-to-End Explainable Facial Action Unit Recognition via Vision-Language Joint Learning
por: Ge, Xuri, et al.
Publicado: (2024)
por: Ge, Xuri, et al.
Publicado: (2024)
HALO: Human-Aligned End-to-end Image Retargeting with Layered Transformations
por: Xu, Yiran, et al.
Publicado: (2025)
por: Xu, Yiran, et al.
Publicado: (2025)
FusionTrack: End-to-End Multi-Object Tracking in Arbitrary Multi-View Environment
por: Li, Xiaohe, et al.
Publicado: (2025)
por: Li, Xiaohe, et al.
Publicado: (2025)
End4: End-to-end Denoising Diffusion for Diffusion-Based Inpainting Detection
por: Wang, Fei, et al.
Publicado: (2025)
por: Wang, Fei, et al.
Publicado: (2025)
ScrewSplat: An End-to-End Method for Articulated Object Recognition
por: Kim, Seungyeon, et al.
Publicado: (2025)
por: Kim, Seungyeon, et al.
Publicado: (2025)
Towards Weakly Supervised End-to-end Learning for Long-video Action Recognition
por: Zhou, Jiaming, et al.
Publicado: (2023)
por: Zhou, Jiaming, et al.
Publicado: (2023)
OVTR: End-to-End Open-Vocabulary Multiple Object Tracking with Transformer
por: Li, Jinyang, et al.
Publicado: (2025)
por: Li, Jinyang, et al.
Publicado: (2025)
Noise2Map: End-to-End Diffusion Model for Semantic Segmentation and Change Detection
por: Shibli, Ali, et al.
Publicado: (2026)
por: Shibli, Ali, et al.
Publicado: (2026)
Multimodal Action Diffusion for Robust End-to-End Autonomous Driving
por: Rodríguez-Vidal, Jorge Daniel, et al.
Publicado: (2026)
por: Rodríguez-Vidal, Jorge Daniel, et al.
Publicado: (2026)
End-to-End Human Instance Matting
por: Liu, Qinglin, et al.
Publicado: (2024)
por: Liu, Qinglin, et al.
Publicado: (2024)
Ejemplares similares
-
Process Optimization and Deployment for Sensor-Based Human Activity Recognition Based on Deep Learning
por: Liu, Hanyu, et al.
Publicado: (2025) -
CMD-HAR: Cross-Modal Disentanglement for Wearable Human Activity Recognition
por: Yu, Ying, et al.
Publicado: (2025) -
Confidence-driven Gradient Modulation for Multimodal Human Activity Recognition: A Dynamic Contrastive Dual-Path Learning Approach
por: Ji, Panpan, et al.
Publicado: (2025) -
End-to-End Visual Autonomous Parking via Control-Aided Attention
por: Chen, Chao, et al.
Publicado: (2025) -
Guiding Attention in End-to-End Driving Models
por: Porres, Diego, et al.
Publicado: (2024)