Learning by Aligning 2D Skeleton Sequences and Multi-Modality Fusion
Fuente:
arXiv
Saved in:
| Main Authors: | Tran, Quoc-Huy, Ahmed, Muhammad, Popattia, Murad, Ahmed, M. Hassan, Konin, Andrey, Zia, M. Zeeshan |
|---|---|
| Format: | Preprint |
| Published: |
2023
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Action Segmentation Using 2D Skeleton Heatmaps and Multi-Modality Fusion
by: Hyder, Syed Waleed, et al.
Published: (2023)
by: Hyder, Syed Waleed, et al.
Published: (2023)
Permutation-Aware Action Segmentation via Unsupervised Frame-to-Segment Alignment
by: Tran, Quoc-Huy, et al.
Published: (2023)
by: Tran, Quoc-Huy, et al.
Published: (2023)
Joint Self-Supervised Video Alignment and Action Segmentation
by: Ali, Ali Shah, et al.
Published: (2025)
by: Ali, Ali Shah, et al.
Published: (2025)
Unsupervised Skeleton-Based Action Segmentation via Hierarchical Spatiotemporal Vector Quantization
by: Ahmed, Umer, et al.
Published: (2026)
by: Ahmed, Umer, et al.
Published: (2026)
TemporalVLM: Video LLMs for Temporal Reasoning in Long Videos
by: Fateh, Fawad Javed, et al.
Published: (2024)
by: Fateh, Fawad Javed, et al.
Published: (2024)
Procedure Learning via Regularized Gromov-Wasserstein Optimal Transport
by: Mahmood, Syed Ahmed, et al.
Published: (2025)
by: Mahmood, Syed Ahmed, et al.
Published: (2025)
A Hierarchical Spatiotemporal Action Tokenizer for In-Context Imitation Learning in Robotics
by: Fateh, Fawad Javed, et al.
Published: (2026)
by: Fateh, Fawad Javed, et al.
Published: (2026)
Viper-F1: Fast and Fine-Grained Multimodal Understanding with Cross-Modal State-Space Modulation
by: Trinh, Quoc-Huy
Published: (2025)
by: Trinh, Quoc-Huy
Published: (2025)
Attention-Based Ensemble Learning for Crop Classification Using Landsat 8-9 Fusion
by: Ramzan, Zeeshan, et al.
Published: (2025)
by: Ramzan, Zeeshan, et al.
Published: (2025)
Skeleton-in-Context: Unified Skeleton Sequence Modeling with In-Context Learning
by: Wang, Xinshun, et al.
Published: (2023)
by: Wang, Xinshun, et al.
Published: (2023)
Multi-Modality Co-Learning for Efficient Skeleton-based Action Recognition
by: Liu, Jinfu, et al.
Published: (2024)
by: Liu, Jinfu, et al.
Published: (2024)
Skeleton2vec: A Self-supervised Learning Framework with Contextualized Target Representations for Skeleton Sequence
by: Xu, Ruizhuo, et al.
Published: (2024)
by: Xu, Ruizhuo, et al.
Published: (2024)
3D Reconstruction via Incremental Structure From Motion
by: Zeeshan, Muhammad, et al.
Published: (2025)
by: Zeeshan, Muhammad, et al.
Published: (2025)
Firebolt-VL: Efficient Vision-Language Understanding with Cross-Modality Modulation
by: Trinh, Quoc-Huy, et al.
Published: (2026)
by: Trinh, Quoc-Huy, et al.
Published: (2026)
PulmoFusion: Advancing Pulmonary Health with Efficient Multi-Modal Fusion
by: Sharshar, Ahmed, et al.
Published: (2025)
by: Sharshar, Ahmed, et al.
Published: (2025)
Modality-Specific Enhancement and Complementary Fusion for Semi-Supervised Multi-Modal Brain Tumor Segmentation
by: Chung, Tien-Dat, et al.
Published: (2025)
by: Chung, Tien-Dat, et al.
Published: (2025)
Investigating Zero-Shot Diagnostic Pathology in Vision-Language Models with Efficient Prompt Design
by: Sharma, Vasudev, et al.
Published: (2025)
by: Sharma, Vasudev, et al.
Published: (2025)
Unified Sequence-to-Sequence Learning for Single- and Multi-Modal Visual Object Tracking
by: Chen, Xin, et al.
Published: (2023)
by: Chen, Xin, et al.
Published: (2023)
2D_3D Feature Fusion via Cross-Modal Latent Synthesis and Attention Guided Restoration for Industrial Anomaly Detection
by: Ali, Usman, et al.
Published: (2025)
by: Ali, Usman, et al.
Published: (2025)
Asynchronous Multimodal Video Sequence Fusion via Learning Modality-Exclusive and -Agnostic Representations
by: Yang, Dingkang, et al.
Published: (2024)
by: Yang, Dingkang, et al.
Published: (2024)
Bridging the Skeleton-Text Modality Gap: Diffusion-Powered Modality Alignment for Zero-shot Skeleton-based Action Recognition
by: Do, Jeonghyeok, et al.
Published: (2024)
by: Do, Jeonghyeok, et al.
Published: (2024)
xMOD: Cross-Modal Distillation for 2D/3D Multi-Object Discovery from 2D motion
by: Lahlali, Saad, et al.
Published: (2025)
by: Lahlali, Saad, et al.
Published: (2025)
Bridging Topology and Deep Representation Learning: A TDA-ViT Fusion Model for Four-Class Brain Tumor Classification
by: Ahmed, Faisal
Published: (2026)
by: Ahmed, Faisal
Published: (2026)
STARS: Self-supervised Tuning for 3D Action Recognition in Skeleton Sequences
by: Mehraban, Soroush, et al.
Published: (2024)
by: Mehraban, Soroush, et al.
Published: (2024)
Online 3D Multi-Camera Perception through Robust 2D Tracking and Depth-based Late Aggregation
by: Le, Vu-Minh, et al.
Published: (2025)
by: Le, Vu-Minh, et al.
Published: (2025)
AdaptiveFusion: Adaptive Multi-Modal Multi-View Fusion for 3D Human Body Reconstruction
by: Chen, Anjun, et al.
Published: (2024)
by: Chen, Anjun, et al.
Published: (2024)
A Multi-Scale Spatial Attention-Based Zero-Shot Learning Framework for Low-Light Image Enhancement
by: Aslam, Muhammad Azeem, et al.
Published: (2025)
by: Aslam, Muhammad Azeem, et al.
Published: (2025)
MolFM-Lite: Multi-Modal Molecular Property Prediction with Conformer Ensemble Attention and Cross-Modal Fusion
by: Shah, Syed Omer, et al.
Published: (2026)
by: Shah, Syed Omer, et al.
Published: (2026)
Skeleton-Snippet Contrastive Learning with Multiscale Feature Fusion for Action Localization
by: Cheng, Qiushuo, et al.
Published: (2025)
by: Cheng, Qiushuo, et al.
Published: (2025)
ActiveAnno3D -- An Active Learning Framework for Multi-Modal 3D Object Detection
by: Ghita, Ahmed, et al.
Published: (2024)
by: Ghita, Ahmed, et al.
Published: (2024)
OccCylindrical: Multi-Modal Fusion with Cylindrical Representation for 3D Semantic Occupancy Prediction
by: Ming, Zhenxing, et al.
Published: (2025)
by: Ming, Zhenxing, et al.
Published: (2025)
Modality Invariant Multimodal Learning to Handle Missing Modalities: A Single-Branch Approach
by: Saeed, Muhammad Saad, et al.
Published: (2024)
by: Saeed, Muhammad Saad, et al.
Published: (2024)
Enhanced Multimodal Content Moderation of Children's Videos using Audiovisual Fusion
by: Ahmed, Syed Hammad, et al.
Published: (2024)
by: Ahmed, Syed Hammad, et al.
Published: (2024)
FusionSAM: Visual Multi-Modal Learning with Segment Anything
by: Li, Daixun, et al.
Published: (2024)
by: Li, Daixun, et al.
Published: (2024)
Progressive Multi-Modal Fusion for Robust 3D Object Detection
by: Mohan, Rohit, et al.
Published: (2024)
by: Mohan, Rohit, et al.
Published: (2024)
Equivariant Multi-Modality Image Fusion
by: Zhao, Zixiang, et al.
Published: (2023)
by: Zhao, Zixiang, et al.
Published: (2023)
ARN-LSTM: A Multi-Stream Fusion Model for Skeleton-based Action Recognition
by: Wang, Chuanchuan, et al.
Published: (2024)
by: Wang, Chuanchuan, et al.
Published: (2024)
VRU-CIPI: Crossing Intention Prediction at Intersections for Improving Vulnerable Road Users Safety
by: Abdelrahman, Ahmed S., et al.
Published: (2025)
by: Abdelrahman, Ahmed S., et al.
Published: (2025)
ConvoFusion: Multi-Modal Conversational Diffusion for Co-Speech Gesture Synthesis
by: Mughal, Muhammad Hamza, et al.
Published: (2024)
by: Mughal, Muhammad Hamza, et al.
Published: (2024)
Deep Learning Segmentation and Classification of Red Blood Cells Using a Large Multi-Scanner Dataset
by: Elmanna, Mohamed, et al.
Published: (2024)
by: Elmanna, Mohamed, et al.
Published: (2024)
Similar Items
-
Action Segmentation Using 2D Skeleton Heatmaps and Multi-Modality Fusion
by: Hyder, Syed Waleed, et al.
Published: (2023) -
Permutation-Aware Action Segmentation via Unsupervised Frame-to-Segment Alignment
by: Tran, Quoc-Huy, et al.
Published: (2023) -
Joint Self-Supervised Video Alignment and Action Segmentation
by: Ali, Ali Shah, et al.
Published: (2025) -
Unsupervised Skeleton-Based Action Segmentation via Hierarchical Spatiotemporal Vector Quantization
by: Ahmed, Umer, et al.
Published: (2026) -
TemporalVLM: Video LLMs for Temporal Reasoning in Long Videos
by: Fateh, Fawad Javed, et al.
Published: (2024)