MultiTSF: Transformer-based Sensor Fusion for Human-Centric Multi-view and Multi-modal Action Recognition
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Nguyen, Trung Thanh, Kawanishi, Yasutomo, John, Vijay, Komamizu, Takahiro, Ide, Ichiro |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
MultiSensor-Home: A Wide-area Multi-modal Multi-view Dataset for Action Recognition and Transformer-based Sensor Fusion
von: Nguyen, Trung Thanh, et al.
Veröffentlicht: (2025)
von: Nguyen, Trung Thanh, et al.
Veröffentlicht: (2025)
View-aware Cross-modal Distillation for Multi-view Action Recognition
von: Nguyen, Trung Thanh, et al.
Veröffentlicht: (2025)
von: Nguyen, Trung Thanh, et al.
Veröffentlicht: (2025)
Action Selection Learning for Multi-label Multi-view Action Recognition
von: Nguyen, Trung Thanh, et al.
Veröffentlicht: (2024)
von: Nguyen, Trung Thanh, et al.
Veröffentlicht: (2024)
One-Stage Open-Vocabulary Temporal Action Detection Leveraging Temporal Multi-scale and Action Label Features
von: Nguyen, Trung Thanh, et al.
Veröffentlicht: (2024)
von: Nguyen, Trung Thanh, et al.
Veröffentlicht: (2024)
Tracking Small Birds by Detection Candidate Region Filtering and Detection History-aware Association
von: Liu, Tingwei, et al.
Veröffentlicht: (2024)
von: Liu, Tingwei, et al.
Veröffentlicht: (2024)
ForestMamba: Sparse Mamba with Geometry-guided Queries for 3D Forest Point Cloud Segmentation
von: Nguyen, Trung Thanh, et al.
Veröffentlicht: (2026)
von: Nguyen, Trung Thanh, et al.
Veröffentlicht: (2026)
Q-Adapter: Visual Query Adapter for Extracting Textually-related Features in Video Captioning
von: Chen, Junan, et al.
Veröffentlicht: (2025)
von: Chen, Junan, et al.
Veröffentlicht: (2025)
Multi-View Video-Based Learning: Leveraging Weak Labels for Frame-Level Perception
von: John, Vijay, et al.
Veröffentlicht: (2024)
von: John, Vijay, et al.
Veröffentlicht: (2024)
Small Object Detection for Birds with Swin Transformer
von: Huo, Da, et al.
Veröffentlicht: (2025)
von: Huo, Da, et al.
Veröffentlicht: (2025)
Multi-proposal Collaboration and Multi-task Training for Weakly-supervised Video Moment Retrieval
von: Zhang, Bolin, et al.
Veröffentlicht: (2026)
von: Zhang, Bolin, et al.
Veröffentlicht: (2026)
Multi-view Action Recognition via Directed Gromov-Wasserstein Discrepancy
von: Nguyen, Hoang-Quan, et al.
Veröffentlicht: (2024)
von: Nguyen, Hoang-Quan, et al.
Veröffentlicht: (2024)
Multi-view Distillation based on Multi-modal Fusion for Few-shot Action Recognition(CLIP-$\mathrm{M^2}$DF)
von: Guo, Fei, et al.
Veröffentlicht: (2024)
von: Guo, Fei, et al.
Veröffentlicht: (2024)
Action Motifs: Self-Supervised Hierarchical Representation of Human Body Movements
von: Kinoshita, Genki, et al.
Veröffentlicht: (2026)
von: Kinoshita, Genki, et al.
Veröffentlicht: (2026)
Facet-Level Persona Control by Trait-Activated Routing with Contrastive SAE for Role-Playing LLMs
von: Tang, Wenqiu, et al.
Veröffentlicht: (2026)
von: Tang, Wenqiu, et al.
Veröffentlicht: (2026)
Trunk-branch Contrastive Network with Multi-view Deformable Aggregation for Multi-view Action Recognition
von: Yang, Yingyuan, et al.
Veröffentlicht: (2025)
von: Yang, Yingyuan, et al.
Veröffentlicht: (2025)
MultiFuser: Multimodal Fusion Transformer for Enhanced Driver Action Recognition
von: Wang, Ruoyu, et al.
Veröffentlicht: (2024)
von: Wang, Ruoyu, et al.
Veröffentlicht: (2024)
Skarimva: Skeleton-based Action Recognition is a Multi-view Application
von: Bermuth, Daniel, et al.
Veröffentlicht: (2026)
von: Bermuth, Daniel, et al.
Veröffentlicht: (2026)
Human-Centric Transformer for Domain Adaptive Action Recognition
von: Lin, Kun-Yu, et al.
Veröffentlicht: (2024)
von: Lin, Kun-Yu, et al.
Veröffentlicht: (2024)
MPTF-Net: Multi-view Pyramid Transformer Fusion Network for LiDAR-based Place Recognition
von: Li, Shuyuan, et al.
Veröffentlicht: (2026)
von: Li, Shuyuan, et al.
Veröffentlicht: (2026)
Static and Dynamic Graph Alignment Network for Temporal Video Grounding
von: Hu, Zhanjie, et al.
Veröffentlicht: (2026)
von: Hu, Zhanjie, et al.
Veröffentlicht: (2026)
Simultaneous Long-tailed Recognition and Multi-modal Fusion for Highly Imbalanced Multi-modal Data
von: Yoon, Heegeon, et al.
Veröffentlicht: (2026)
von: Yoon, Heegeon, et al.
Veröffentlicht: (2026)
LM-MCVT: A Lightweight Multi-modal Multi-view Convolutional-Vision Transformer Approach for 3D Object Recognition
von: Xiong, Songsong, et al.
Veröffentlicht: (2025)
von: Xiong, Songsong, et al.
Veröffentlicht: (2025)
Actor-agnostic Multi-label Action Recognition with Multi-modal Query
von: Mondal, Anindya, et al.
Veröffentlicht: (2023)
von: Mondal, Anindya, et al.
Veröffentlicht: (2023)
Multi-level and Multi-modal Action Anticipation
von: Kim, Seulgi, et al.
Veröffentlicht: (2025)
von: Kim, Seulgi, et al.
Veröffentlicht: (2025)
Thermal Polarimetric Multi-view Stereo
von: Kushida, Takahiro, et al.
Veröffentlicht: (2025)
von: Kushida, Takahiro, et al.
Veröffentlicht: (2025)
mmWalk: Towards Multi-modal Multi-view Walking Assistance
von: Ying, Kedi, et al.
Veröffentlicht: (2025)
von: Ying, Kedi, et al.
Veröffentlicht: (2025)
Multi-modal and Multi-view Fundus Image Fusion for Retinopathy Diagnosis via Multi-scale Cross-attention and Shifted Window Self-attention
von: Huang, Yonghao, et al.
Veröffentlicht: (2025)
von: Huang, Yonghao, et al.
Veröffentlicht: (2025)
ARN-LSTM: A Multi-Stream Fusion Model for Skeleton-based Action Recognition
von: Wang, Chuanchuan, et al.
Veröffentlicht: (2024)
von: Wang, Chuanchuan, et al.
Veröffentlicht: (2024)
Multi-modal Sensor Fusion for Auto Driving Perception: A Survey
von: Huang, Keli, et al.
Veröffentlicht: (2022)
von: Huang, Keli, et al.
Veröffentlicht: (2022)
GaitMA: Pose-guided Multi-modal Feature Fusion for Gait Recognition
von: Min, Fanxu, et al.
Veröffentlicht: (2024)
von: Min, Fanxu, et al.
Veröffentlicht: (2024)
PE-MVCNet: Multi-view and Cross-modal Fusion Network for Pulmonary Embolism Prediction
von: Guo, Zhaoxin, et al.
Veröffentlicht: (2024)
von: Guo, Zhaoxin, et al.
Veröffentlicht: (2024)
MVAFormer: RGB-based Multi-View Spatio-Temporal Action Recognition with Transformer
von: Yamane, Taiga, et al.
Veröffentlicht: (2025)
von: Yamane, Taiga, et al.
Veröffentlicht: (2025)
Tile Classification Based Viewport Prediction with Multi-modal Fusion Transformer
von: Zhang, Zhihao, et al.
Veröffentlicht: (2023)
von: Zhang, Zhihao, et al.
Veröffentlicht: (2023)
Cross-view Action Recognition Understanding From Exocentric to Egocentric Perspective
von: Truong, Thanh-Dat, et al.
Veröffentlicht: (2023)
von: Truong, Thanh-Dat, et al.
Veröffentlicht: (2023)
Multi-modal Fusion based Q-distribution Prediction for Controlled Nuclear Fusion
von: Wang, Shiao, et al.
Veröffentlicht: (2024)
von: Wang, Shiao, et al.
Veröffentlicht: (2024)
VHAKG: A Multi-modal Knowledge Graph Based on Synchronized Multi-view Videos of Daily Activities
von: Egami, Shusaku, et al.
Veröffentlicht: (2024)
von: Egami, Shusaku, et al.
Veröffentlicht: (2024)
Building a Multi-modal Spatiotemporal Expert for Zero-shot Action Recognition with CLIP
von: Yu, Yating, et al.
Veröffentlicht: (2024)
von: Yu, Yating, et al.
Veröffentlicht: (2024)
Bridging the Gap between Multi-focus and Multi-modal: A Focused Integration Framework for Multi-modal Image Fusion
von: Li, Xilai, et al.
Veröffentlicht: (2023)
von: Li, Xilai, et al.
Veröffentlicht: (2023)
RAMer: Reconstruction-based Adversarial Model for Multi-party Multi-modal Multi-label Emotion Recognition
von: Yang, Xudong, et al.
Veröffentlicht: (2025)
von: Yang, Xudong, et al.
Veröffentlicht: (2025)
Investigating Conceptual Blending of a Diffusion Model for Improving Nonword-to-Image Generation
von: Matsuhira, Chihaya, et al.
Veröffentlicht: (2024)
von: Matsuhira, Chihaya, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
MultiSensor-Home: A Wide-area Multi-modal Multi-view Dataset for Action Recognition and Transformer-based Sensor Fusion
von: Nguyen, Trung Thanh, et al.
Veröffentlicht: (2025) -
View-aware Cross-modal Distillation for Multi-view Action Recognition
von: Nguyen, Trung Thanh, et al.
Veröffentlicht: (2025) -
Action Selection Learning for Multi-label Multi-view Action Recognition
von: Nguyen, Trung Thanh, et al.
Veröffentlicht: (2024) -
One-Stage Open-Vocabulary Temporal Action Detection Leveraging Temporal Multi-scale and Action Label Features
von: Nguyen, Trung Thanh, et al.
Veröffentlicht: (2024) -
Tracking Small Birds by Detection Candidate Region Filtering and Detection History-aware Association
von: Liu, Tingwei, et al.
Veröffentlicht: (2024)