iPay: Integrated Payment Action Recognition via Multimodal Networks and Adaptive Spatial Prior Learning
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Huang, Kaicong, Oh, Weiheng, Guggisberg, Thomas, Ke, Ruimin |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Background Matters Too: A Language-Enhanced Adversarial Framework for Person Re-Identification
von: Huang, Kaicong, et al.
Veröffentlicht: (2025)
von: Huang, Kaicong, et al.
Veröffentlicht: (2025)
Real2Sim: A Physics-driven and Editable Gaussian Splatting Framework for Autonomous Driving Scenes
von: Huang, Kaicong, et al.
Veröffentlicht: (2026)
von: Huang, Kaicong, et al.
Veröffentlicht: (2026)
TransitReID: Transit OD Data Collection with Occlusion-Resistant Dynamic Passenger Re-Identification
von: Huang, Kaicong, et al.
Veröffentlicht: (2025)
von: Huang, Kaicong, et al.
Veröffentlicht: (2025)
Towards Adaptive Fusion of Multimodal Deep Networks for Human Action Recognition
von: Yudistira, Novanto
Veröffentlicht: (2025)
von: Yudistira, Novanto
Veröffentlicht: (2025)
Weakly Supervised Temporal Action Localization via Dual-Prior Collaborative Learning Guided by Multimodal Large Language Models
von: Zhang, Quan, et al.
Veröffentlicht: (2024)
von: Zhang, Quan, et al.
Veröffentlicht: (2024)
SNN-Driven Multimodal Human Action Recognition via Sparse Spatial-Temporal Data Fusion
von: Zheng, Naichuan, et al.
Veröffentlicht: (2025)
von: Zheng, Naichuan, et al.
Veröffentlicht: (2025)
DoGCLR: Dominance-Game Contrastive Learning Network for Skeleton-Based Action Recognition
von: Li, Yanshan, et al.
Veröffentlicht: (2025)
von: Li, Yanshan, et al.
Veröffentlicht: (2025)
MCN-CL: Multimodal Cross-Attention Network and Contrastive Learning for Multimodal Emotion Recognition
von: Li, Feng, et al.
Veröffentlicht: (2025)
von: Li, Feng, et al.
Veröffentlicht: (2025)
Are Spatial-Temporal Graph Convolution Networks for Human Action Recognition Over-Parameterized?
von: Xie, Jianyang, et al.
Veröffentlicht: (2025)
von: Xie, Jianyang, et al.
Veröffentlicht: (2025)
Multimodal Cross-Domain Few-Shot Learning for Egocentric Action Recognition
von: Hatano, Masashi, et al.
Veröffentlicht: (2024)
von: Hatano, Masashi, et al.
Veröffentlicht: (2024)
Guided Interpretable Facial Expression Recognition via Spatial Action Unit Cues
von: Belharbi, Soufiane, et al.
Veröffentlicht: (2024)
von: Belharbi, Soufiane, et al.
Veröffentlicht: (2024)
Joint Image-Instance Spatial-Temporal Attention for Few-shot Action Recognition
von: Qian, Zefeng, et al.
Veröffentlicht: (2025)
von: Qian, Zefeng, et al.
Veröffentlicht: (2025)
Multimodal Prototype-Enhanced Network for Few-Shot Action Recognition
von: Ni, Xinzhe, et al.
Veröffentlicht: (2022)
von: Ni, Xinzhe, et al.
Veröffentlicht: (2022)
Patch as Node: Human-Centric Graph Representation Learning for Multimodal Action Recognition
von: Liang, Zeyu, et al.
Veröffentlicht: (2025)
von: Liang, Zeyu, et al.
Veröffentlicht: (2025)
Unsupervised Spatial-Temporal Feature Enrichment and Fidelity Preservation Network for Skeleton based Action Recognition
von: Li, Chuankun, et al.
Veröffentlicht: (2024)
von: Li, Chuankun, et al.
Veröffentlicht: (2024)
Multi-Track Multimodal Learning on iMiGUE: Micro-Gesture and Emotion Recognition
von: Martirosyan, Arman, et al.
Veröffentlicht: (2025)
von: Martirosyan, Arman, et al.
Veröffentlicht: (2025)
An Effective End-to-End Solution for Multimodal Action Recognition
von: Wang, Songping, et al.
Veröffentlicht: (2025)
von: Wang, Songping, et al.
Veröffentlicht: (2025)
STMT: A Spatial-Temporal Mesh Transformer for MoCap-Based Action Recognition
von: Zhu, Xiaoyu, et al.
Veröffentlicht: (2023)
von: Zhu, Xiaoyu, et al.
Veröffentlicht: (2023)
Human-Centric Transformer for Domain Adaptive Action Recognition
von: Lin, Kun-Yu, et al.
Veröffentlicht: (2024)
von: Lin, Kun-Yu, et al.
Veröffentlicht: (2024)
Sample-level Adaptive Knowledge Distillation for Action Recognition
von: Li, Ping, et al.
Veröffentlicht: (2025)
von: Li, Ping, et al.
Veröffentlicht: (2025)
LPTR-AFLNet: Lightweight Integrated Chinese License Plate Rectification and Recognition Network
von: Xu, Guangzhu, et al.
Veröffentlicht: (2025)
von: Xu, Guangzhu, et al.
Veröffentlicht: (2025)
DINOLight: Robust Ambient Light Normalization with Self-supervised Visual Prior Integration
von: Oh, Youngjin, et al.
Veröffentlicht: (2026)
von: Oh, Youngjin, et al.
Veröffentlicht: (2026)
Video Generation with Learned Action Prior
von: Sarkar, Meenakshi, et al.
Veröffentlicht: (2024)
von: Sarkar, Meenakshi, et al.
Veröffentlicht: (2024)
Leveraging Foundation Models for Multimodal Graph-Based Action Recognition
von: Ziaeetabar, Fatemeh, et al.
Veröffentlicht: (2025)
von: Ziaeetabar, Fatemeh, et al.
Veröffentlicht: (2025)
When Spatial meets Temporal in Action Recognition
von: Chen, Huilin, et al.
Veröffentlicht: (2024)
von: Chen, Huilin, et al.
Veröffentlicht: (2024)
Spatial-Temporal Perception with Causal Inference for Naturalistic Driving Action Recognition
von: Chang, Qing, et al.
Veröffentlicht: (2025)
von: Chang, Qing, et al.
Veröffentlicht: (2025)
Multimodal Video Emotion Recognition with Reliable Reasoning Priors
von: Wang, Zhepeng, et al.
Veröffentlicht: (2025)
von: Wang, Zhepeng, et al.
Veröffentlicht: (2025)
Human Action Recognition (HAR) Using Skeleton-based Spatial Temporal Relative Transformer Network: ST-RTR
von: Mehmood, Faisal, et al.
Veröffentlicht: (2024)
von: Mehmood, Faisal, et al.
Veröffentlicht: (2024)
Multimodal Knowledge Distillation for Egocentric Action Recognition Robust to Missing Modalities
von: Santos-Villafranca, Maria, et al.
Veröffentlicht: (2025)
von: Santos-Villafranca, Maria, et al.
Veröffentlicht: (2025)
MultiFuser: Multimodal Fusion Transformer for Enhanced Driver Action Recognition
von: Wang, Ruoyu, et al.
Veröffentlicht: (2024)
von: Wang, Ruoyu, et al.
Veröffentlicht: (2024)
MA-FSAR: Multimodal Adaptation of CLIP for Few-Shot Action Recognition
von: Xing, Jiazheng, et al.
Veröffentlicht: (2023)
von: Xing, Jiazheng, et al.
Veröffentlicht: (2023)
Efficient Egocentric Action Recognition with Multimodal Data
von: Calzavara, Marco, et al.
Veröffentlicht: (2025)
von: Calzavara, Marco, et al.
Veröffentlicht: (2025)
Advancing Vision Transformer with Enhanced Spatial Priors
von: Fan, Qihang, et al.
Veröffentlicht: (2026)
von: Fan, Qihang, et al.
Veröffentlicht: (2026)
Latent Temporal Discrepancy as Motion Prior: A Loss-Weighting Strategy for Dynamic Fidelity in T2V
von: Wu, Meiqi, et al.
Veröffentlicht: (2026)
von: Wu, Meiqi, et al.
Veröffentlicht: (2026)
FRAMER: Frequency-Aligned Self-Distillation with Adaptive Modulation Leveraging Diffusion Priors for Real-World Image Super-Resolution
von: Choi, Seungho, et al.
Veröffentlicht: (2025)
von: Choi, Seungho, et al.
Veröffentlicht: (2025)
LaDy: Lagrangian-Dynamic Informed Network for Skeleton-based Action Segmentation via Spatial-Temporal Modulation
von: Ji, Haoyu, et al.
Veröffentlicht: (2026)
von: Ji, Haoyu, et al.
Veröffentlicht: (2026)
Signal-SGN: A Spiking Graph Convolutional Network for Skeletal Action Recognition via Learning Temporal-Frequency Dynamics
von: Zheng, Naichuan, et al.
Veröffentlicht: (2024)
von: Zheng, Naichuan, et al.
Veröffentlicht: (2024)
Lane Change Classification and Prediction with Action Recognition Networks
von: Liang, Kai, et al.
Veröffentlicht: (2022)
von: Liang, Kai, et al.
Veröffentlicht: (2022)
Active Generation Network of Human Skeleton for Action Recognition
von: Liu, Long, et al.
Veröffentlicht: (2024)
von: Liu, Long, et al.
Veröffentlicht: (2024)
Action Selection Learning for Multi-label Multi-view Action Recognition
von: Nguyen, Trung Thanh, et al.
Veröffentlicht: (2024)
von: Nguyen, Trung Thanh, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Background Matters Too: A Language-Enhanced Adversarial Framework for Person Re-Identification
von: Huang, Kaicong, et al.
Veröffentlicht: (2025) -
Real2Sim: A Physics-driven and Editable Gaussian Splatting Framework for Autonomous Driving Scenes
von: Huang, Kaicong, et al.
Veröffentlicht: (2026) -
TransitReID: Transit OD Data Collection with Occlusion-Resistant Dynamic Passenger Re-Identification
von: Huang, Kaicong, et al.
Veröffentlicht: (2025) -
Towards Adaptive Fusion of Multimodal Deep Networks for Human Action Recognition
von: Yudistira, Novanto
Veröffentlicht: (2025) -
Weakly Supervised Temporal Action Localization via Dual-Prior Collaborative Learning Guided by Multimodal Large Language Models
von: Zhang, Quan, et al.
Veröffentlicht: (2024)