Leveraging Foundation Models for Multimodal Graph-Based Action Recognition
Fuente:
arXiv
Saved in:
| Main Authors: | Ziaeetabar, Fatemeh, Wörgötter, Florentin |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Why Relational Graphs Will Save the Next Generation of Vision Foundation Models?
by: Ziaeetabar, Fatemeh
Published: (2025)
by: Ziaeetabar, Fatemeh
Published: (2025)
EfficientGFormer: Multimodal Brain Tumor Segmentation via Pruned Graph-Augmented Transformer
by: Ziaeetabar, Fatemeh
Published: (2025)
by: Ziaeetabar, Fatemeh
Published: (2025)
Neuro-Symbolic Manipulation Understanding with Enriched Semantic Event Chains
by: Ziaeetabar, Fatemeh
Published: (2026)
by: Ziaeetabar, Fatemeh
Published: (2026)
Beyond Sequences: A Benchmark for Atomic Hand-Object Interaction Using a Static RNN Encoder
by: Movahed, Yousef Azizi, et al.
Published: (2025)
by: Movahed, Yousef Azizi, et al.
Published: (2025)
DRL-Guided Neural Batch Sampling for Semi-Supervised Pixel-Level Anomaly Detection
by: Noghredeh, Amirhossein Khadivi, et al.
Published: (2025)
by: Noghredeh, Amirhossein Khadivi, et al.
Published: (2025)
Leveraging Foundation Model Automatic Data Augmentation Strategies and Skeletal Points for Hands Action Recognition in Industrial Assembly Lines
by: Wu, Liang, et al.
Published: (2024)
by: Wu, Liang, et al.
Published: (2024)
Leveraging Temporal Contextualization for Video Action Recognition
by: Kim, Minji, et al.
Published: (2024)
by: Kim, Minji, et al.
Published: (2024)
Comparison of marker-less 2D image-based methods for infant pose estimation
by: Jahn, Lennart, et al.
Published: (2024)
by: Jahn, Lennart, et al.
Published: (2024)
Leveraging Generic Foundation Models for Multimodal Surgical Data Analysis
by: Pezold, Simon, et al.
Published: (2025)
by: Pezold, Simon, et al.
Published: (2025)
Patch as Node: Human-Centric Graph Representation Learning for Multimodal Action Recognition
by: Liang, Zeyu, et al.
Published: (2025)
by: Liang, Zeyu, et al.
Published: (2025)
Leveraging CLIP Encoder for Multimodal Emotion Recognition
by: Song, Yehun, et al.
Published: (2025)
by: Song, Yehun, et al.
Published: (2025)
An Improved Graph Pooling Network for Skeleton-Based Action Recognition
by: Wu, Cong, et al.
Published: (2024)
by: Wu, Cong, et al.
Published: (2024)
From Recognition to Prediction: Leveraging Sequence Reasoning for Action Anticipation
by: Liu, Xin, et al.
Published: (2024)
by: Liu, Xin, et al.
Published: (2024)
Advancing Human Action Recognition with Foundation Models trained on Unlabeled Public Videos
by: Qian, Yang, et al.
Published: (2024)
by: Qian, Yang, et al.
Published: (2024)
Leveraging Vision-Language Large Models for Interpretable Video Action Recognition with Semantic Tokenization
by: Peng, Jingwei, et al.
Published: (2025)
by: Peng, Jingwei, et al.
Published: (2025)
Leveraging Medical Foundation Model Features in Graph Neural Network-Based Retrieval of Breast Histopathology Images
by: Saeidi, Nematollah, et al.
Published: (2024)
by: Saeidi, Nematollah, et al.
Published: (2024)
Graph-Based Multimodal and Multi-view Alignment for Keystep Recognition
by: Romero, Julia Lee, et al.
Published: (2025)
by: Romero, Julia Lee, et al.
Published: (2025)
ALow-Cost Real-Time Framework for Industrial Action Recognition Using Foundation Models
by: Wang, Zhicheng, et al.
Published: (2024)
by: Wang, Zhicheng, et al.
Published: (2024)
Foundation Model for Skeleton-Based Human Action Understanding
by: Wang, Hongsong, et al.
Published: (2025)
by: Wang, Hongsong, et al.
Published: (2025)
Leveraging Foundation Models via Knowledge Distillation in Multi-Object Tracking: Distilling DINOv2 Features to FairMOT
by: Faber, Niels G., et al.
Published: (2024)
by: Faber, Niels G., et al.
Published: (2024)
An Effective End-to-End Solution for Multimodal Action Recognition
by: Wang, Songping, et al.
Published: (2025)
by: Wang, Songping, et al.
Published: (2025)
High-Performance Inference Graph Convolutional Networks for Skeleton-Based Action Recognition
by: Wang, Junyi, et al.
Published: (2023)
by: Wang, Junyi, et al.
Published: (2023)
Enhancing Action Recognition by Leveraging the Hierarchical Structure of Actions and Textual Context
by: Benavent-Lledo, Manuel, et al.
Published: (2024)
by: Benavent-Lledo, Manuel, et al.
Published: (2024)
Adversarial Robustness in RGB-Skeleton Action Recognition: Leveraging Attention Modality Reweighter
by: Liu, Chao, et al.
Published: (2024)
by: Liu, Chao, et al.
Published: (2024)
MK-SGN: A Spiking Graph Convolutional Network with Multimodal Fusion and Knowledge Distillation for Skeleton-based Action Recognition
by: Zheng, Naichuan, et al.
Published: (2024)
by: Zheng, Naichuan, et al.
Published: (2024)
Skeleton-Based Action Recognition with Spatial-Structural Graph Convolution
by: Wang, Jingyao, et al.
Published: (2024)
by: Wang, Jingyao, et al.
Published: (2024)
Multimodal Knowledge Distillation for Egocentric Action Recognition Robust to Missing Modalities
by: Santos-Villafranca, Maria, et al.
Published: (2025)
by: Santos-Villafranca, Maria, et al.
Published: (2025)
Towards Adaptive Fusion of Multimodal Deep Networks for Human Action Recognition
by: Yudistira, Novanto
Published: (2025)
by: Yudistira, Novanto
Published: (2025)
Multimodal Cross-Domain Few-Shot Learning for Egocentric Action Recognition
by: Hatano, Masashi, et al.
Published: (2024)
by: Hatano, Masashi, et al.
Published: (2024)
MultiFuser: Multimodal Fusion Transformer for Enhanced Driver Action Recognition
by: Wang, Ruoyu, et al.
Published: (2024)
by: Wang, Ruoyu, et al.
Published: (2024)
MA-FSAR: Multimodal Adaptation of CLIP for Few-Shot Action Recognition
by: Xing, Jiazheng, et al.
Published: (2023)
by: Xing, Jiazheng, et al.
Published: (2023)
Percept, Chat, and then Adapt: Multimodal Knowledge Transfer of Foundation Models for Open-World Video Recognition
by: Chen, Boyu, et al.
Published: (2024)
by: Chen, Boyu, et al.
Published: (2024)
Towards Universal Skeleton-Based Action Recognition
by: Kuang, Jidong, et al.
Published: (2026)
by: Kuang, Jidong, et al.
Published: (2026)
Efficient Egocentric Action Recognition with Multimodal Data
by: Calzavara, Marco, et al.
Published: (2025)
by: Calzavara, Marco, et al.
Published: (2025)
Topological Symmetry Enhanced Graph Convolution for Skeleton-Based Action Recognition
by: Liang, Zeyu, et al.
Published: (2024)
by: Liang, Zeyu, et al.
Published: (2024)
Benchmarking Sensitivity of Continual Graph Learning for Skeleton-Based Action Recognition
by: Wei, Wei, et al.
Published: (2024)
by: Wei, Wei, et al.
Published: (2024)
Diffusion-Based Action Recognition Generalizes to Untrained Domains
by: Guimaraes, Rogerio, et al.
Published: (2025)
by: Guimaraes, Rogerio, et al.
Published: (2025)
G3CN: Gaussian Topology Refinement Gated Graph Convolutional Network for Skeleton-Based Action Recognition
by: Ren, Haiqing, et al.
Published: (2025)
by: Ren, Haiqing, et al.
Published: (2025)
Identifying Ethical Biases in Action Recognition Models
by: Baltaretu, Ana, et al.
Published: (2026)
by: Baltaretu, Ana, et al.
Published: (2026)
Knowledge is Power: Advancing Few-shot Action Recognition with Multimodal Semantics from MLLMs
by: Xing, Jiazheng, et al.
Published: (2026)
by: Xing, Jiazheng, et al.
Published: (2026)
Similar Items
-
Why Relational Graphs Will Save the Next Generation of Vision Foundation Models?
by: Ziaeetabar, Fatemeh
Published: (2025) -
EfficientGFormer: Multimodal Brain Tumor Segmentation via Pruned Graph-Augmented Transformer
by: Ziaeetabar, Fatemeh
Published: (2025) -
Neuro-Symbolic Manipulation Understanding with Enriched Semantic Event Chains
by: Ziaeetabar, Fatemeh
Published: (2026) -
Beyond Sequences: A Benchmark for Atomic Hand-Object Interaction Using a Static RNN Encoder
by: Movahed, Yousef Azizi, et al.
Published: (2025) -
DRL-Guided Neural Batch Sampling for Semi-Supervised Pixel-Level Anomaly Detection
by: Noghredeh, Amirhossein Khadivi, et al.
Published: (2025)